<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: DigitalOcean</title>
    <description>The latest articles on DEV Community by DigitalOcean (digitalocean).</description>
    <link>https://dev.to/digitalocean</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Forganization%2Fprofile_image%2F175%2F369f1227-0eac-4a88-8d3c-08851bf0b117.png</url>
      <title>DEV Community: DigitalOcean</title>
      <link>https://dev.to/digitalocean</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/digitalocean"/>
    <language>en</language>
    <item>
      <title>How to route LLM requests by cost vs. latency</title>
      <dc:creator>DigitalOcean</dc:creator>
      <pubDate>Fri, 02 Oct 2026 18:33:17 +0000</pubDate>
      <link>https://dev.to/digitalocean/how-to-route-llm-requests-by-cost-vs-latency-463j</link>
      <guid>https://dev.to/digitalocean/how-to-route-llm-requests-by-cost-vs-latency-463j</guid>
      <description>&lt;p&gt;Routing LLM requests by cost and latency means sending each request to the cheapest or fastest model that still meets the quality bar, rather than hardcoding a single model for everything. Production traffic isn't uniform: routine lookups, complex troubleshooting, and background jobs have different cost and latency requirements, even within a single app. &lt;/p&gt;

&lt;p&gt;Using an inference router like DigitalOcean Inference Router involves defining request categories, assigning each a selection policy (cost, latency, or a benchmarked best fit), and setting fallback models so an unavailable model doesn't break the request. Cost is the price per token (which can vary by up to 100x across a model catalog); latency is mainly the time to first token (TTFT). The two don't move together, so route per task rather than using a single blended score.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Key takeaways:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Routing is a per-workload policy, not a search for "the best model."
&lt;/li&gt;
&lt;li&gt;Cost and latency usually need separate, explicit priorities.
&lt;/li&gt;
&lt;li&gt;Fallback and cache-aware routing protect the savings that routing is meant to deliver.
&lt;/li&gt;
&lt;li&gt;The DigitalOcean Inference Router implements this as configuration, not custom infrastructure.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  LLM routing strategies
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Static rules&lt;/strong&gt;: Hardcode model per request type. Simple, but brittle.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Policy-based routing&lt;/strong&gt;: A model pool per task with a cost/latency/optimal policy. Most production systems land here.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Dynamic routing&lt;/strong&gt;: A classifier infers the task and policy automatically.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Routing LLM requests: A quick decision framework
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Workload&lt;/th&gt;
&lt;th&gt;Priority&lt;/th&gt;
&lt;th&gt;Policy&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Real-time chat&lt;/td&gt;
&lt;td&gt;Latency&lt;/td&gt;
&lt;td&gt;Speed-optimized (TTFT)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Routine lookups&lt;/td&gt;
&lt;td&gt;Cost&lt;/td&gt;
&lt;td&gt;Cost-optimized&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Complex reasoning&lt;/td&gt;
&lt;td&gt;Quality&lt;/td&gt;
&lt;td&gt;Benchmarked "optimal"&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Batch/background&lt;/td&gt;
&lt;td&gt;Cost&lt;/td&gt;
&lt;td&gt;Cost-optimized&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Fallback models&lt;/strong&gt; keep requests completing when the preferred model is down or rate-limited. &lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cache-aware routing&lt;/strong&gt; matters too: switching models to save a fraction of a cent can break a cached prompt and cost more overall.&lt;/p&gt;

&lt;h2&gt;
  
  
  How the DigitalOcean Inference Router routes LLM requests
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://docs.digitalocean.com/products/inference/how-to/use-inference-router/" rel="noopener noreferrer"&gt;Inference Router&lt;/a&gt; tasks pair a model pool with a selection policy: preset tasks default to a benchmarked &lt;strong&gt;Optimal&lt;/strong&gt; policy; custom tasks choose &lt;strong&gt;Cost Efficiency&lt;/strong&gt;, &lt;strong&gt;Speed Optimization&lt;/strong&gt; (TTFT), or &lt;strong&gt;Manual Ranking&lt;/strong&gt;. &lt;/p&gt;

&lt;p&gt;Fallback models handle unmatched traffic. &lt;code&gt;X-Model-Affinity&lt;/code&gt; preserves cache reuse, and &lt;code&gt;X-Routing-Max-Switch-Spend-Pct&lt;/code&gt; (default 20%) caps the cost of switching mid-session. &lt;/p&gt;

&lt;p&gt;It's a drop-in change. Set &lt;code&gt;"model": "router:your-router-name"&lt;/code&gt;. Routing decisions typically resolve in about 200ms, billed at the serving model's &lt;a href="https://docs.digitalocean.com/products/inference/details/pricing/" rel="noopener noreferrer"&gt;standard rate&lt;/a&gt; with no separate router charge during public preview.&lt;/p&gt;

&lt;h2&gt;
  
  
  LLM routing FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;What is LLM routing?&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
LLM routing is the practice of directing each request to the model best suited to it—by cost, latency, or task fit—instead of sending every request through one model regardless of what it needs. The DigitalOcean Inference Router implements this as a managed feature, so teams get task-aware routing and fallback handling without building it themselves.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How do you &lt;a href="https://www.digitalocean.com/resources/articles/llm-cost-calculation-guide" rel="noopener noreferrer"&gt;model cost-per-token&lt;/a&gt; for LLM inference?&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Multiply input tokens by the model's input rate and output tokens by its output rate, then sum the two, since the rates are usually very different. Because rates vary widely across a model catalog, this is best tracked per task rather than as one blended number. The DigitalOcean Inference Router reports token usage and cost-relevant metrics per model and per task in its Analyze view.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What metrics matter for LLM inference observability?&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Time to first token (TTFT), time per output token (TPOT), and inter-token latency (ITL) are the core metrics for judging responsiveness. A latency-optimized routing policy should be measured against these, not just overall request time. The DigitalOcean Inference Router uses TTFT specifically as the basis for its Speed Optimization selection policy.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What causes cold start latency in GPU inference, and how do you avoid it?&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Cold starts happen when a model has to load onto available GPU capacity before it can serve a request, rather than running on an already-warm instance. Pooled serverless capacity and routing policies that keep related traffic on one model reduce how often this happens. The DigitalOcean Inference Router uses model affinity to keep a session's requests on the same warm model, rather than triggering repeated cold starts across models.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can I route by both cost and latency at the same time?&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Yes—typically by defining separate tasks or policies for different request types rather than one blended score. For example, cost-first for background work and latency-first for real-time chat. The DigitalOcean Inference Router supports this directly. Each task in a router can have its own model pool and its own Cost Efficiency, Speed Optimization, or Manual Ranking policy.&lt;/p&gt;

&lt;h2&gt;
  
  
  References &amp;amp; further reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://docs.digitalocean.com/products/inference/how-to/use-inference-router/" rel="noopener noreferrer"&gt;How to use Inference Router&lt;/a&gt;: The full documentation on tasks, selection policies, fallback models, and the cache-aware router referenced throughout this piece. Worth reading directly before you configure a router for production traffic.
&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://docs.digitalocean.com/products/inference/details/pricing/" rel="noopener noreferrer"&gt;Inference pricing&lt;/a&gt;: Current per-model Serverless Inference rates, since routing decisions are only as good as the pricing data behind them.
&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.digitalocean.com/blog/inference-router-architecture" rel="noopener noreferrer"&gt;How DigitalOcean built Inference Router&lt;/a&gt;: An inside look at the architecture decisions behind cache-aware routing and fallback handling, useful if you want the reasoning behind the mechanics.
&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.digitalocean.com/blog/serverless-inference-deep-dive" rel="noopener noreferrer"&gt;DigitalOcean Serverless Inference: a deep dive&lt;/a&gt;: Background on the serverless layer that Inference Router sits in front of.
&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.digitalocean.com/resources/articles/best-llm-routers" rel="noopener noreferrer"&gt;DigitalOcean's best LLM routers roundup&lt;/a&gt;: A good next read if you're comparing routing options across vendors rather than implementing a policy on a platform you've already chosen.
&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.digitalocean.com/products/inference-engine" rel="noopener noreferrer"&gt;AI Inference Engine overview&lt;/a&gt;: shows where Inference Router fits alongside serverless, batch, and dedicated inference on DigitalOcean's broader platform.&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>startup</category>
      <category>developers</category>
      <category>infrastructure</category>
    </item>
    <item>
      <title>Which inference provider has both serverless and dedicated GPU tiers on the same platform with one API?</title>
      <dc:creator>DigitalOcean</dc:creator>
      <pubDate>Mon, 28 Sep 2026 17:17:40 +0000</pubDate>
      <link>https://dev.to/digitalocean/which-inference-provider-has-both-serverless-and-dedicated-gpu-tiers-on-the-same-platform-with-one-50dc</link>
      <guid>https://dev.to/digitalocean/which-inference-provider-has-both-serverless-and-dedicated-gpu-tiers-on-the-same-platform-with-one-50dc</guid>
      <description>&lt;p&gt;DigitalOcean's Inference Engine puts serverless, dedicated, and batch inference behind a single endpoint, on the same GPU Droplets infrastructure, so a workload can move from pay-per-token prototyping to a reserved-GPU deployment without switching platforms or rewriting its integration. A handful of other inference providers share a single API across serverless and dedicated tiers, too, but they generally limit that unification to inference alone, rather than to the broader cloud the dedicated GPUs run on.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why serverless-only or dedicated-only platforms create a migration tax
&lt;/h2&gt;

&lt;p&gt;Most teams don't start with a clear answer to "How much GPU capacity do we need?". They prototype on a pay-per-token endpoint, and only once traffic gets steady and predictable does it make financial sense to reserve GPUs outright. Many inference providers treat serverless and dedicated as separate products, with distinct SDKs, auth, and billing dashboards, so moving from one to the other means rewriting integration code in the middle of a product's growth curve, which is exactly when a team has the least time to do it. A provider that puts both tiers behind a single API avoids that: you prototype on serverless and repoint the same client to a dedicated endpoint once volume justifies it, often by changing a model or endpoint parameter rather than an integration.&lt;/p&gt;

&lt;h2&gt;
  
  
  How the DigitalOcean Inference Engine unifies serverless, dedicated, and batch inference
&lt;/h2&gt;

&lt;p&gt;The DigitalOcean &lt;a href="https://www.digitalocean.com/products/inference-engine" rel="noopener noreferrer"&gt;Inference Engine&lt;/a&gt; runs three inference patterns over the same GPU Droplets infrastructure, behind a single endpoint. These solutions include real-time &lt;a href="https://docs.digitalocean.com/products/inference/how-to/use-serverless-inference-hub" rel="noopener noreferrer"&gt;Serverless inference&lt;/a&gt; across 70+ open-weight and frontier models, &lt;a href="https://docs.digitalocean.com/products/inference/how-to/use-dedicated-inference" rel="noopener noreferrer"&gt;Dedicated inference&lt;/a&gt; on reserved, single-tenant GPU capacity with support for bringing your own model, and asynchronous Batch Inference for large, non-real-time jobs. The Preview Inference Router adds cost- and latency-based routing on top, and a shared Model Playground and evaluations tooling sit on the same account.&lt;br&gt;
That's what makes DigitalOcean a solid solution for AI-native companies requiring both serverless and dedicated GPU tiers. The unification isn't limited to the inference API itself. It extends to the GPU Droplets the dedicated tier runs on, as well as the rest of the AI-Native Cloud around them. A workload doesn't need to switch vendors, contracts, or billing relationships as it matures from prototype to production.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to check before assuming "one API" elsewhere means the same thing
&lt;/h2&gt;

&lt;p&gt;Several other inference-focused platforms let serverless and dedicated calls share one API too, usually as a way past per-token rate limits or shared-capacity contention once traffic justifies reserving GPUs. But a shared API can sound more unified on a landing page than it is in production. So it's worth confirming directly against each provider's current docs: does provisioning real dedicated capacity happen through the API, or does it route through a sales conversation first? Is the dedicated tier's hardware genuinely part of the same cloud a team already runs on, or does the provider specialize in inference alone, with GPUs sitting apart from the application's data and storage? And does "one API" cover just the inference call—or the account, billing, and observability around it too?&lt;/p&gt;

&lt;h2&gt;
  
  
  What else to verify before standardizing on a provider
&lt;/h2&gt;

&lt;p&gt;Beyond scope, it's worth confirming: whether the model you need is available in both tiers (not every model ships to both), what the cutover mechanism actually looks like (a new endpoint URL versus a full redeploy), how billing reconciles if usage straddles both tiers in the same month, and whether dedicated capacity is truly single-tenant or a reserved slice of shared hardware. For teams already running other infrastructure on a given cloud, it's also worth weighing how much value comes from the inference API alone versus not having to move data, agents, or storage to a second vendor to get dedicated GPU access.&lt;/p&gt;

&lt;h2&gt;
  
  
  References &amp;amp; further reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://www.digitalocean.com/products/inference-engine" rel="noopener noreferrer"&gt;DigitalOcean's Inference Engine&lt;/a&gt; brings serverless, dedicated, and batch inference together in one place.&lt;/li&gt;
&lt;li&gt;Our &lt;a href="https://docs.digitalocean.com/products/inference/how-to/use-dedicated-inference" rel="noopener noreferrer"&gt;Dedicated Inference documentation&lt;/a&gt; covers bring-your-own-model (BYOM) support and scaling settings for teams ready to move off serverless.&lt;/li&gt;
&lt;li&gt;Read our &lt;a href="https://docs.digitalocean.com/products/inference/how-to/use-serverless-inference-hub" rel="noopener noreferrer"&gt;Serverless Inference documentation&lt;/a&gt; for the model catalog and routing options available on the shared endpoint.&lt;/li&gt;
&lt;li&gt;Reading through each inference provider's own docs tends to be the fastest way to confirm whether "serverless and dedicated" really means one API or two separate products.&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>infrastructure</category>
      <category>llm</category>
      <category>startup</category>
    </item>
    <item>
      <title>What It Actually Costs to Serve a 1M-Token Model in Production</title>
      <dc:creator>DigitalOcean</dc:creator>
      <pubDate>Mon, 21 Sep 2026 20:33:36 +0000</pubDate>
      <link>https://dev.to/digitalocean/what-it-actually-costs-to-serve-a-1m-token-model-in-production-4f0k</link>
      <guid>https://dev.to/digitalocean/what-it-actually-costs-to-serve-a-1m-token-model-in-production-4f0k</guid>
      <description>&lt;p&gt;Model providers now advertise context windows large enough to hold a codebase or a stack of contracts in a single request. It's tempting to read that number as a green light, but it isn't. A model that &lt;em&gt;accepts&lt;/em&gt; a million tokens and a system that can &lt;em&gt;serve&lt;/em&gt; a million tokens to real, concurrent users within your latency and cost budget are two different engineering problems. Conflating them is how teams end up with a demo that works and a bill or a &lt;a href="https://www.digitalocean.com/community/conceptual-articles/when-your-vllm-p99" rel="noopener noreferrer"&gt;p99 latency&lt;/a&gt; chart that doesn't.&lt;/p&gt;

&lt;p&gt;You serve a 1M-token model in production by treating context length as a cost and latency variable you manage on purpose, not a spec you flip on. Let’s dive into the memory math, the latency cliff, and the cost curve for serving a model with a million tokens of context, plus the caching and routing techniques that production teams use today to keep all three in check. Plus, a straight answer on when long context is the wrong tool for the job and retrieval should do the work instead.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Key takeaways&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Serving a 1M-token context model in production is a memory, latency, and cost problem layered on top of a model capability. Supporting long context and serving it reliably are not the same claim.
&lt;/li&gt;
&lt;li&gt;Serving long context effectively means stable agent sessions and predictable spend, rather than retries, timeouts, and a bill that scales with the window instead of the question.
&lt;/li&gt;
&lt;li&gt;The real decision isn't whether to "turn on" long context, it's which combination of caching, prefix reuse, and retrieval fits your workload's cost and latency budget.
&lt;/li&gt;
&lt;li&gt;Production-grade 1M-token serving combines general techniques, including KV cache economics and prefix-reuse approaches like SGLang's RadixAttention. Use DigitalOcean Inference Router for cache-aware routing, plus retrieval-augmented generation (RAG) as the alternative and complement.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What "serving" a long-context model actually means
&lt;/h2&gt;

&lt;p&gt;Context is everything you send to a model in a single request, measured in tokens: the prompt, the documents, the conversation history. A 128K-token window is roughly a novel's worth of text, and frontier closed models now advertise up to 1M tokens (roughly, ten novels). That number describes what the model can read in isolation, but not what happens to time-to-first-token, batch size, or your GPU bill once dozens of those requests land on the same cluster at once.&lt;/p&gt;

&lt;p&gt;Support is a model capability, and performance is a serving question. The two come apart the moment real traffic shows up.&lt;/p&gt;

&lt;p&gt;Two mechanics decide whether a long-context request is slow to start or slow to finish:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://www.digitalocean.com/community/tutorials/prefill-decode-disaggregation" rel="noopener noreferrer"&gt;Prefill&lt;/a&gt;: The compute-bound pass where the model reads your entire input before writing anything.
&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.digitalocean.com/community/conceptual-articles/how-kv-caching-slashes-llm-inference-costs-at-scale" rel="noopener noreferrer"&gt;KV cache&lt;/a&gt;: The memory that stores each token's key and value vectors so the model doesn't reread the input on every step.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Dig in to how &lt;a href="https://www.digitalocean.com/community/tutorials/prefill-decode-disaggregation" rel="noopener noreferrer"&gt;prefill/decode disaggregation&lt;/a&gt; works: the two-phase split behind every long-context serving decision below.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why its worth optimizing for long-context
&lt;/h2&gt;

&lt;p&gt;Getting long-context serving right pays off in multiple ways.&amp;nbsp;&lt;/p&gt;

&lt;p&gt;Here's what's actually at stake once you're running this in production:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Stable agent and document workloads&lt;/strong&gt;: A correctly cached long-context session makes it possible for a coding agent or a multi-turn document review to hold state across dozens of turns without resending the same 100K-token context on every message.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Costs you can forecast&lt;/strong&gt;: Once you know whether traffic is cache-friendly or retrieval-shaped, it's easier to accurately predict spend.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No silent accuracy loss&lt;/strong&gt;: Long-context serving done well accounts for the fact that models answer less reliably as input grows. This happens by routing precision lookups to retrieval instead of trusting long-context recall, and by testing accuracy at the context lengths you'll actually run rather than assuming a bigger model or window is automatically more reliable.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Latency you can put in an SLA&lt;/strong&gt;: Understanding the prefill/decode split lets you set a time-to-first-token target you can actually hit, instead of promising an interactive response time a long-context request can't deliver.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Room to scale without a rebuild&lt;/strong&gt;: Architectures built around caching and routing from the start absorb growth in context length and concurrent users without the rewrite that a "just resend everything" approach eventually forces.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What to evaluate before you commit to a long-context approach
&lt;/h2&gt;

&lt;p&gt;Long-context serving isn't a single decision, but rather a set of infrastructure and architecture choices considered together:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Cache-awareness&lt;/strong&gt;: Check whether the serving layer reuses the KV cache across requests that share a common prefix, rather than recomputing it each time. The DigitalOcean &lt;a href="https://www.digitalocean.com/products/inference-engine" rel="noopener noreferrer"&gt;Inference Router&lt;/a&gt;, for example, pins a session to the same model so that a warm cache isn't discarded mid-task when switching to a different, cheaper model.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Batching under long-context load&lt;/strong&gt;: Confirm the system maintains healthy batch sizes as the average context length grows, using techniques like chunked prefill and paged memory allocation, rather than letting a single long request starve everything queued behind it.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Isolation from noisy neighbors&lt;/strong&gt;: Ask whether long-context and interactive traffic run in separate pools. &lt;a href="https://docs.digitalocean.com/products/inference/how-to/use-dedicated-inference/" rel="noopener noreferrer"&gt;Dedicated Inference&lt;/a&gt; on DigitalOcean exists specifically so a large document-processing job doesn't blow up latency for interactive users on the same infrastructure.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pricing that reflects the real cost curve&lt;/strong&gt;: Look for tiered, per-token pricing that acknowledges long-context requests cost more proportionally, plus a prompt-caching rate that's meaningfully cheaper than a full rewrite.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Observability into cache health, not just request counts&lt;/strong&gt;: Make sure you can see KV-cache occupancy, eviction, and preemption counts—not just average latency—since that's what actually explains a p99 problem caused by long-context traffic.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Resend, cache, retrieve, or both: what each costs per session
&lt;/h2&gt;

&lt;p&gt;The cleanest way to understand the cost of each approach is to price a full session rather than a single token.&amp;nbsp;&lt;/p&gt;

&lt;p&gt;These examples take a 10-turn agent conversation carrying a stable 128K-token context—a codebase or a document set—where each turn adds a short question and gets a short answer:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Strategy&lt;/th&gt;
&lt;th&gt;Use Case&lt;/th&gt;
&lt;th&gt;Mechanism&lt;/th&gt;
&lt;th&gt;Cost per 10-turn session (128K context)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Resend the full context every turn&lt;/td&gt;
&lt;td&gt;Short, one-off sessions where simplicity matters more than cost&lt;/td&gt;
&lt;td&gt;Full context re-sent as input tokens on every turn&lt;/td&gt;
&lt;td&gt;$3.93&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cache the context; reuse it&lt;/td&gt;
&lt;td&gt;Stable, reused context — a codebase, a document set, a long agent session&lt;/td&gt;
&lt;td&gt;Prefill cost paid once; later turns reuse the cached KV state&lt;/td&gt;
&lt;td&gt;$0.95&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Retrieve a relevant slice (RAG)&lt;/td&gt;
&lt;td&gt;Large or changing corpora where each query touches a small fraction of the data&lt;/td&gt;
&lt;td&gt;Only the relevant slice enters context; the rest of the corpus stays out&lt;/td&gt;
&lt;td&gt;$0.33&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Figures per DigitalOcean's &lt;a href="https://www.digitalocean.com/community/tutorials/long-context-llm-serving-tradeoffs-2026" rel="noopener noreferrer"&gt;long-context serving tradeoffs analysis&lt;/a&gt;, priced at Claude Sonnet's published rates ($3.00/M input tokens, $3.75/M to write the cache, $0.30/M to read it).&lt;/p&gt;

&lt;p&gt;Caching cuts the bill roughly 4x against resending; retrieval cuts it roughly 12x. At thousands of sessions a day, that architecture choice matters more than any provider's per-token rate.&amp;nbsp;&lt;/p&gt;

&lt;h2&gt;
  
  
  The four tradeoffs that decide whether long-context works in production
&lt;/h2&gt;

&lt;p&gt;A 1M-token context window is a spec, not a guarantee that serving it will work. Memory, latency, throughput, and accuracy each break in their own way as context grows, and "supports 1M tokens" and "serves 1M tokens well" can be very different claims.&lt;/p&gt;

&lt;h3&gt;
  
  
  Memory: the KV cache doesn't scale the way people expect
&lt;/h3&gt;

&lt;p&gt;For every token a model reads, it stores a key and a value vector so it doesn't have to reprocess the input on each new token. That cache grows in direct proportion to context length: &lt;a href="https://www.digitalocean.com/community/tutorials/long-context-inference-production-cost" rel="noopener noreferrer"&gt;the math&lt;/a&gt; works out to roughly 43GB for a single 128K-token request on a 70-billion-parameter model in BF16—close to the model's own memory footprint. Push to a million tokens, and the cache alone exceeds what a single high-end GPU holds, before the model's own weights are even loaded.&lt;/p&gt;

&lt;h3&gt;
  
  
  Latency: why time-to-first-token is the number that breaks
&lt;/h3&gt;

&lt;p&gt;Attention compares every token against every other token, so compute grows with the square of input length: double the context, quadruple the work. That happens entirely during prefill, before a user sees a single word, which is why time-to-first-token ranges from milliseconds with short context to multiple seconds with long context. It also means one long request can hold a GPU for seconds, quietly degrading p95 and p99 latency for every short, unrelated request queued behind it.&lt;/p&gt;

&lt;h3&gt;
  
  
  Throughput: the bandwidth ceiling no amount of compute fixes
&lt;/h3&gt;

&lt;p&gt;Generating each output token requires reading the entire KV cache from GPU memory, and the read speed is limited by hardware bandwidth. Double the context, double the cache, and tokens-per-second roughly halves, with a limit that has nothing to do with how much compute is sitting idle. A cluster can appear underutilized on a FLOPS dashboard and still serve long-context requests slowly because the bottleneck has shifted to the memory bus.&lt;/p&gt;

&lt;h3&gt;
  
  
  Accuracy: why a bigger window doesn't guarantee a better answer
&lt;/h3&gt;

&lt;p&gt;Longer context doesn't reliably mean a better answer. A 2026 measurement study found that accuracy at long context lengths doesn't track model size — a smaller model sometimes outperformed a larger sibling — and that a wrong answer just becomes a retry, which turns an accuracy problem into a latency problem measured in total time-to-correct-answer. Production systems that route by expected accuracy at a given context length, not just by cost, cut that wasted retry time substantially.&lt;/p&gt;

&lt;h2&gt;
  
  
  How production teams are actually solving the long-context problem today
&lt;/h2&gt;

&lt;p&gt;Successful teams aren't fighting the KV cache, they're making it reusable. &lt;a href="https://www.digitalocean.com/resources/articles/what-is-sglang" rel="noopener noreferrer"&gt;SGLang&lt;/a&gt;'s RadixAttention implements prefix caching alongside continuous batching, so a repeated prompt prefix (i.e. a system prompt, a shared reference document, a stable set of few-shot examples) gets computed once and reused across requests instead of recomputed on every call. That keeps throughput comparatively stable even as context length grows into the millions of tokens, which is exactly the workload where naive serving falls apart first.&lt;/p&gt;

&lt;p&gt;You can run this pattern yourself on &lt;a href="https://www.digitalocean.com/products/gpu-droplets" rel="noopener noreferrer"&gt;DigitalOcean GPU Droplets&lt;/a&gt;, or skip the setup work with the &lt;a href="https://www.digitalocean.com/resources/articles/what-is-sglang" rel="noopener noreferrer"&gt;SGLang 1-Click Marketplace app&lt;/a&gt;, which handles the underlying infrastructure for you.&amp;nbsp;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F20hh5y3klu0mii00c2pp.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F20hh5y3klu0mii00c2pp.png" alt="SGLang on DigitalOcean" width="800" height="335"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;DigitalOcean's Inference Router adds cache-aware routing on top of that. It keeps a session pinned to the same model, so a warm KV cache doesn't get discarded mid-task by a switch to a different model chosen for cost reasons.&amp;nbsp;&lt;/p&gt;

&lt;p&gt;If you're self-hosting long-context serving on GPU Droplets rather than going through a managed router, three additional levers matter:&amp;nbsp;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Chunked prefill, which breaks a large prefill pass into pieces so it doesn't stall the rest of the batch.&amp;nbsp;
&lt;/li&gt;
&lt;li&gt;Tuning max batch size and max batched-token limits for your actual mix of context lengths.
&lt;/li&gt;
&lt;li&gt;Keeping long-context traffic in a separate pool from interactive traffic so one doesn't degrade the other.&amp;nbsp;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;When time-to-first-token degrades as context grows, &lt;a href="https://www.digitalocean.com/community/tutorials/debugging-p99-ttft-llm-inference" rel="noopener noreferrer"&gt;DigitalOcean's inference-debugging guidance&lt;/a&gt; points to KV-cache pressure—not the scheduler—as the usual culprit. This is why watching KV-cache occupancy, eviction, and preemption counts diagnoses the problem faster than staring at a scheduler queue.&lt;/p&gt;

&lt;h2&gt;
  
  
  Do you even need long context? Choosing between caching, retrieval, or using both
&lt;/h2&gt;

&lt;p&gt;If a model accepts a million tokens, it's fair to ask why you'd also use a &lt;a href="https://www.digitalocean.com/community/conceptual-articles/a-dive-into-vector-databases" rel="noopener noreferrer"&gt;vector database&lt;/a&gt;, a chunking pipeline, and a retrieval step. The answer is that the price of skipping retrieval includes the quadratic prefill cost, the multi-second time-to-first-token, and the tiered per-token pricing. They’re all paid on every request—even when the answer only needed to use three paragraphs of the input.&lt;/p&gt;

&lt;p&gt;As a practical rule, long context plus caching wins when your working set is stable and reused. For example, many requests against the same document set, or a day-long agent session where the context is the state. RAG wins when your corpus is larger than any context window, changes frequently, or is mostly irrelevant to any single query. Precision lookups favor RAG even when the full corpus would technically fit, since retrieval keeps the model working at the shorter lengths where it's most reliable. Most production systems end up using both: retrieval to decide what enters the context, a long window to hold a generous amount of it, and caching wherever it repeats.&lt;/p&gt;

&lt;h2&gt;
  
  
  Long-context LLM serving FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;How do you serve long-context (1M token) models in production?&lt;/strong&gt;&amp;nbsp;&lt;/p&gt;

&lt;p&gt;Serving a 1M-token model in production means treating context length as a cost and latency variable you manage on purpose, not a spec you flip on, since supporting long context and serving it reliably are not the same claim. It requires managing KV cache memory, the quadratic prefill cost, and the resulting time-to-first-token, then choosing the right mix of caching, prefix reuse, and retrieval for your workload. The DigitalOcean Inference Router applies cache-aware routing to this problem by pinning a session to the same model so a warm KV cache isn't discarded mid-task.&lt;/p&gt;

&lt;p&gt;&amp;nbsp;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What is disaggregated prefill and decode in LLM inference?&lt;/strong&gt;&amp;nbsp;&lt;/p&gt;

&lt;p&gt;Prefill is the compute-bound pass where the model reads your entire input before writing anything, while decode is the token-by-token generation that follows. Attention compares every token against every other token during prefill, so compute grows with the square of input length, meaning one long request can hold a GPU for seconds and degrade latency for every short request queued behind it. Chunked prefill breaks a large prefill pass into pieces so it doesn't stall the rest of the batch.&lt;/p&gt;

&lt;p&gt;&amp;nbsp;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What metrics matter for LLM inference observability (TTFT, TPOT, ITL)?&lt;/strong&gt;&amp;nbsp;&lt;/p&gt;

&lt;p&gt;Time-to-first-token is the number that breaks under long context, ranging from milliseconds with short context to multiple seconds with long context, because it's entirely determined by the prefill pass. Beyond latency averages, observability into KV-cache occupancy, eviction, and preemption counts is what actually explains a p99 problem caused by long-context traffic, not request counts alone. KV-cache pressure, not the scheduler, is the usual culprit when time-to-first-token degrades as context grows.&lt;/p&gt;

&lt;p&gt;&amp;nbsp;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Which inference provider has the best caching to reduce token costs for long autonomous coding sessions?&lt;/strong&gt;&amp;nbsp;&lt;/p&gt;

&lt;p&gt;Caching the context and reusing it lets a coding agent hold state across dozens of turns without resending the same 100K-token context every message. SGLang's RadixAttention implements this as prefix caching alongside continuous batching, computing a repeated prompt prefix once and reusing it across requests. DigitalOcean supports this pattern on GPU Droplets or through the SGLang 1-Click Marketplace app, and the Inference Router adds cache-aware routing so a warm cache isn't discarded by a mid-task model switch.&lt;/p&gt;

&lt;p&gt;&amp;nbsp;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How do you model cost-per-token for LLM inference?&lt;/strong&gt;&amp;nbsp;&lt;/p&gt;

&lt;p&gt;The cleanest way to model cost is pricing a full session rather than a single token: resending the full context every turn costs $3.93 for a 10-turn, 128K-context session, caching and reusing that context brings it to $0.95, and retrieving only the relevant slice (RAG) brings it to $0.33. Caching cuts the bill roughly 4x against resending, while retrieval cuts it roughly 12x, a difference that matters more at scale than any provider's per-token rate. DigitalOcean prices around this same curve, with tiered per-token rates that account for long-context requests costing more than proportionally, plus a prompt-caching rate meaningfully cheaper than a full rewrite.&lt;/p&gt;

&lt;h2&gt;
  
  
  Serve long-context models without managing the cache yourself
&lt;/h2&gt;

&lt;p&gt;The &lt;a href="https://www.digitalocean.com/products/inference-engine" rel="noopener noreferrer"&gt;DigitalOcean Inference Engine&lt;/a&gt; prices per token by model, with prompt-caching rates listed right next to the standard rate. That means you can work out what a long-context session costs before you run it, not after the invoice shows up. The Inference Router automatically applies cache-aware routing, pinning a session to the model whose KV cache is already warm rather than discarding it during a mid-task switch.&amp;nbsp;&lt;/p&gt;

&lt;p&gt;That gets you:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Automatic model-to-request matching, so cost- or latency-sensitive requests route to the right model without the need for writing the routing logic yourself.
&lt;/li&gt;
&lt;li&gt;Dedicated Inference to isolate long-context, latency-sensitive workloads from noisy-neighbor traffic.
&lt;/li&gt;
&lt;li&gt;Serverless-to-dedicated scaling with no migration step, so a workload that outgrows shared capacity doesn't require re-architecting.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://cloud.digitalocean.com/registrations/new" rel="noopener noreferrer"&gt;Get started with DigitalOcean's Inference Engine&lt;/a&gt; →&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.digitalocean.com/customers/hippocratic-ai" rel="noopener noreferrer"&gt;Hippocratic AI&lt;/a&gt; cut prefill latency 2x on long-context clinical sessions running production inference on DigitalOcean's cache-aware routing.&amp;nbsp;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.youtube.com/watch?v=WeAAISnnNVA" rel="noopener noreferrer"&gt;https://www.youtube.com/watch?v=WeAAISnnNVA&lt;/a&gt;&amp;nbsp;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;These customer-reported results reflect Hippocratic AI’s methodology, workload, configuration, and measurement period and may not be representative of results in other environments. Results in customer environments may vary depending on configuration, implementation, and usage. Results and/or savings are not guaranteed.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Any references to third-party companies, trademarks, or logos in this document are for informational purposes only and do not imply any affiliation with, sponsorship by, or endorsement of those third parties.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>infrastructure</category>
      <category>llm</category>
      <category>developers</category>
    </item>
    <item>
      <title>AI token billing continues to cause sticker shock</title>
      <dc:creator>DigitalOcean</dc:creator>
      <pubDate>Fri, 18 Sep 2026 16:25:40 +0000</pubDate>
      <link>https://dev.to/digitalocean/ai-token-billing-continues-to-cause-sticker-shock-4a21</link>
      <guid>https://dev.to/digitalocean/ai-token-billing-continues-to-cause-sticker-shock-4a21</guid>
      <description>&lt;h3&gt;
  
  
  Tech talk: are there inference providers that are transparent about token usage?&amp;nbsp;
&lt;/h3&gt;

&lt;p&gt;A top pain point for companies using AI is the discrepancy between expected and actual invoice rates. Particularly when it comes to tokens. The increase in AI usage and tokenmaxxing has created a paradox where a provider knows exactly how many tokens it generated on your behalf, yet you find out only after the fact. It's become a real cost problem when using reasoning models. Internal step-by-step reasoning counts as billed output, but it's invisible by design. The providers willing to open up their token accounting now are betting that transparency will become table stakes before their competitors catch up.&amp;nbsp;&lt;/p&gt;

&lt;p&gt;The 2025 study, “&lt;a href="https://arxiv.org/html/2508.00912v1" rel="noopener noreferrer"&gt;Predictive Auditing of Hidden Tokens in LLM APIs via Reasoning Length Estimation&lt;/a&gt;,” estimated that hidden reasoning tokens can account for over 90% of a model's total token spend on complex tasks. A short, simple-looking answer can hide 10,000+ reasoning tokens, but the only confirmation of their use is the bill.&lt;/p&gt;

&lt;p&gt;This trend in billing surprises has rightfully drawn industry attention. Especially with token usage projected to grow 24x to 120 quadrillion tokens per month by 2030, according to &lt;a href="https://www.goldmansachs.com/insights/articles/ai-agents-forecast-to-boost-tech-cash-flow-as-usage-soars" rel="noopener noreferrer"&gt;research from Goldman Sachs&lt;/a&gt;. Developers have started noticing the gap between what usage dashboards imply and what invoices actually charge, and the industry's response is arriving on two tracks.&amp;nbsp;&lt;/p&gt;

&lt;p&gt;First, standards bodies are getting involved. The Linux Foundation's proposed &lt;a href="https://www.digitalocean.com/blog/inference-router-architecture" rel="noopener noreferrer"&gt;Tokenomics Foundation&lt;/a&gt; wants open, vendor-neutral measurement standards so token accounting isn't just whatever a provider's internal dashboard says it is.&amp;nbsp;&lt;/p&gt;

&lt;p&gt;Second, providers like Groq and Cerebras are sharing flat per-token pricing. And the newly launched DigitalOcean &lt;a href="https://www.digitalocean.com/blog/inference-router-architecture" rel="noopener noreferrer"&gt;Inference Router&lt;/a&gt; helps your team route tasks to the most cost-effective or optimal model, providing budget control based on your priorities for specific AI tasks.&amp;nbsp;&lt;/p&gt;

&lt;p&gt;Thanks to rapidly scaling AI adoption, developers know these costs aren’t a simple line item to ignore. When evaluating inference options, knowing exactly how your tokens are used and what that can mean for overall billing is essential.&amp;nbsp;&lt;/p&gt;




&lt;h3&gt;
  
  
  Event calendar&amp;nbsp;
&lt;/h3&gt;

&lt;p&gt;&amp;nbsp;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.digitalocean.com/open-intelligence-summit" rel="noopener noreferrer"&gt;&lt;strong&gt;Open Intelligence Summit by DigitalOcean&lt;/strong&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The argument is no longer “Is open source AI cheaper?” It’s now “Who owns intelligence?” A movement is forming around open intelligence and generating the belief that the intelligence powering our AI products and companies should be more open, portable, and controllable.&lt;/p&gt;

&lt;p&gt;Join DigitalOcean, RadixArk, Inferact, and more for deep dives, demos, and discussions about the technologies that keep intelligence portable, interoperable, and in the hands of builders. The two-day event runs at The Midway in San Francisco from October 12-13th.&amp;nbsp;&lt;/p&gt;




&lt;h3&gt;
  
  
  Industry news&amp;nbsp;
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Coding agents are earning real trust in production codebases&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
A &lt;a href="https://arxiv.org/abs/2607.21832" rel="noopener noreferrer"&gt;study&lt;/a&gt; by researchers at Polytechnique Montréal, led by Canada Research Chair Foutse Khomh, analyzed 220,612 PRs across 539 Python repos and found Claude Code's PRs merge 84.3% of the time—well ahead of Codex (73.5%), Cursor (63.9%), Copilot (59.6%), and Devin (43.0%)—with bug rates comparable to or lower than human-written code.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;OpenAI launches GPT-6 Astra&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
OpenAI introduced &lt;a href="https://openai.com/index/gpt-6-astra/" rel="noopener noreferrer"&gt;GPT-6 Astra&lt;/a&gt;, a new frontier model focused on advanced reasoning, computer use, coding, scientific research, and professional workflows. The model achieves major gains across benchmarks, including 99.9% on ARC-AGI-3 and 100% on ExploitBench, while also improving task efficiency and alignment.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Anthropic introduces Claude Fable 5.1 and Mythos 5.1&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
The &lt;a href="https://www.anthropic.com/claude-fable-and-mythos-5-1" rel="noopener noreferrer"&gt;new models&lt;/a&gt; target advanced coding, knowledge work, and scientific research, with Fable 5.1 offering stronger performance at lower cost and Mythos 5.1 providing more permissive safeguards for vetted cybersecurity and life-science users.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Grok Bot comes to Android&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Cursor has launched the &lt;a href="https://forum.cursor.com/t/grok-bot-is-now-live-on-android/170384" rel="noopener noreferrer"&gt;Grok Bot Android app&lt;/a&gt;, letting users assign tasks to AI Bots from their phones, continue conversations across devices, and monitor multiple Bots running in parallel on cloud computers. The app is currently in beta for eligible Cursor and SuperGrok plans.&amp;nbsp;&lt;/p&gt;




&lt;h3&gt;
  
  
  Community tutorials&amp;nbsp;
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://www.digitalocean.com/community/tutorials/why-your-best-model-is-two-models" rel="noopener noreferrer"&gt;&lt;strong&gt;Why Your Best Model Is Two Models: Routing Between Kimi K3 and Claude&lt;/strong&gt;&lt;/a&gt;&lt;br&gt;&lt;br&gt;
Hardcoded routing logic gets messy fast and using a small model like Haiku to classify requests means paying for two calls on every one. DigitalOcean built Plano-Orchestrator, a routing model that scored 87.84% accuracy against GPT-5.1's 86.93% and Claude Sonnet 4.5's 86.11%, deciding where a request goes in about 200 milliseconds. Walk through building a router that pools Kimi K3 and Claude, and see how to check with live traffic data whether it's actually working.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.digitalocean.com/community/tutorials/spot-gpu-droplets-doks" rel="noopener noreferrer"&gt;&lt;strong&gt;Resilient GPU Compute on DigitalOcean Kubernetes: Surviving Spot Interruptions&lt;/strong&gt;&lt;/a&gt;&lt;br&gt;&lt;br&gt;
When DOKS reclaims Spot GPU capacity, it deletes the entire node pool at once, not one node at a time so the usual "autoscaler quietly relaunches it" assumption breaks completely, and nothing automatically brings a reclaimed pool back. Walk through building a pre-provisioned On-Demand fallback pool that costs nothing until a reclaim happens, and run the companion repo's automated interruption simulation to see the failover in action.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.digitalocean.com/community/conceptual-articles/glm-5-3-flash-cost-per-token" rel="noopener noreferrer"&gt;&lt;strong&gt;GLM-5.3-Flash is the cheapest model on DigitalOcean. It's also the most verbose.&lt;/strong&gt;&lt;/a&gt;&lt;br&gt;&lt;br&gt;
GLM-5.3-Flash lists at one twelfth of Qwen3.8-Max's output price, but on unconfigured defaults, it burned 625 output tokens answering what port SSH uses because it silently defaults to maximum reasoning effort, cutting its real-world price advantage down to just 2.3x. See the measured cost breakdown across 2,700 API calls, and get the one-line reasoning_effort fix that made Flash 6.3x cheaper on the same prompts.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.digitalocean.com/community/tutorials/pgvector-at-scale-article" rel="noopener noreferrer"&gt;&lt;strong&gt;When pgvector Starts to Slow Down as Your Vector Table Grows&lt;/strong&gt;&lt;/a&gt;&lt;br&gt;&lt;br&gt;
A pgvector's query latency doesn't creep up gradually as a table grows. It stays flat for a long time, then jumps hard the moment the HNSW graph outgrows available RAM and Postgres starts reading it from disk instead. Get the six-step method for finding your own breaking point, measuring recall, latency, and index build time at the table sizes you'll actually reach.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.digitalocean.com/community/tutorials/does-context-length-affect-inference-cost-linearly" rel="noopener noreferrer"&gt;&lt;strong&gt;Does Context Length Affect Inference Cost Linearly? We Measured Why It Doesn't&lt;/strong&gt;&lt;/a&gt;&lt;br&gt;&lt;br&gt;
Cost calculators price context length linearly, but measured throughput on a single H200 serving Ministral 3 14B fell from 19,089.8 tokens/sec at 2K context to 4,967.2 tokens/sec at 256K: a 3.84x collapse driven by concurrent request capacity dropping from 311 to just 2 as the fixed KV cache pool fills up. See the full cost curve and break-even utilization math showing exactly where a dedicated GPU stops being cheaper than a flat per-token serverless rate.&amp;nbsp;&lt;/p&gt;




&lt;h3&gt;
  
  
  Product updates
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://www.digitalocean.com/community/tutorials/spot-gpu-droplets-doks" rel="noopener noreferrer"&gt;&lt;strong&gt;DigitalOcean Kubernetes (DOKS) now supports Spot GPU Droplets&lt;/strong&gt;&lt;/a&gt;&amp;nbsp;&lt;/p&gt;

&lt;p&gt;Running Spot GPUs on DOKS automates interruption handling with automatic cordoning, draining, and PodDisruptionBudget enforcement. Cluster autoscaling dynamically scales your GPU capacity based on workload demand. Kubernetes capabilities such as labels, taints, and node affinity allow you to configure on-demand node pools as a fallback when Spot capacity is reclaimed. Plus, reclaim events surface as native Kubernetes events. &lt;a href="https://cloud.digitalocean.com/gpus/new?i=403c6d&amp;amp;region=mem1&amp;amp;size=gpu-mi355x8-2304gb-spot&amp;amp;fleetUuid=94ff983a-9cbc-4198-a245-118a1e5e476b" rel="noopener noreferrer"&gt;Try it today&lt;/a&gt;&amp;nbsp;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>startup</category>
    </item>
    <item>
      <title>When is Serverless Inference Cheaper than Your Self Hosted GPU? I Benchmarked gpt-oss-120b on Both</title>
      <dc:creator>Yash Sharma</dc:creator>
      <pubDate>Tue, 23 Jun 2026 14:57:47 +0000</pubDate>
      <link>https://dev.to/digitalocean/when-is-serverless-inference-cheaper-than-your-self-hosted-gpu-i-benchmarked-gpt-oss-120b-on-both-4bla</link>
      <guid>https://dev.to/digitalocean/when-is-serverless-inference-cheaper-than-your-self-hosted-gpu-i-benchmarked-gpt-oss-120b-on-both-4bla</guid>
      <description>&lt;p&gt;If you run LLM inference in production, you eventually will ask yourself, should you rent a GPU and run the model yourself, or do you use a serverless API and pay per token? Everyone has an opinion. Far fewer people show you the actual numbers that decide it.&lt;/p&gt;

&lt;p&gt;So I ran both. I put the same model, &lt;a href="https://www.digitalocean.com/community/tutorials/run-gpt-oss-vllm-amd-gpu-droplet-rocm" rel="noopener noreferrer"&gt;gpt-oss-120b&lt;/a&gt;, on two setups, self-hosted with vLLM on a single &lt;a href="https://www.digitalocean.com/blog/now-available-amd-instinct-mi300x-gpus" rel="noopener noreferrer"&gt;AMD MI300X GPU Droplet&lt;/a&gt;, and on &lt;a href="https://docs.digitalocean.com/products/inference/how-to/si-overview/" rel="noopener noreferrer"&gt;DigitalOcean's Serverless Inference&lt;/a&gt;. Then I measured the cold start, the warm latency, and the cost, and worked out exactly where one becomes the better choice than the other.&lt;/p&gt;

&lt;p&gt;In short, self-hosted GPU is faster and more consistent once it's warm, but it carries a real cold start and you pay for it around the clock. Serverless hides the cold start and costs almost nothing at low volume, but you pay per token. Which one wins comes down to your traffic shape and how much your model actually outputs. Here are the numbers.&lt;/p&gt;

&lt;h2&gt;
  
  
  How long is the cold start on a self-hosted GPU?
&lt;/h2&gt;

&lt;p&gt;When you run a model yourself, the GPU doesn't hold the model permanently. The weights have to be loaded from disk into the GPU's memory and the inference engine has to initialize before it can answer a single request. That startup delay, the gap between "process launched" and "first token out," is the cold start. You pay it every time you start the server fresh, a new deploy, a restart after a crash, or a new node coming up to handle load.&lt;/p&gt;

&lt;p&gt;My setup here was a single AMD MI300X GPU Droplet running gpt-oss-120b with vLLM, with the weights already cached on disk. I started vLLM from cold and timed how long it took before the first token came back.&lt;/p&gt;

&lt;p&gt;It took about &lt;strong&gt;61 seconds&lt;/strong&gt;. That wasn't a one-off, either, across three restarts it landed between 60.8 and 61.4 seconds every time.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjnm67pdeefgjbtehcise.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjnm67pdeefgjbtehcise.png" alt="Cold Start of Model" width="762" height="139"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;One thing worth being precise about, because it's the most common misread, that 61 seconds is &lt;em&gt;not&lt;/em&gt; the time to download the model. The weights were already saved on disk, so this is the cost you pay on a restart or redeploy, not a one-time setup. The startup logs show where the time actually goes:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Phase&lt;/th&gt;
&lt;th&gt;Time&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Load weights from disk into VRAM (~68 GB)&lt;/td&gt;
&lt;td&gt;~24 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;torch.compile&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;~4 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CUDA graph capture&lt;/td&gt;
&lt;td&gt;~11 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Engine init, KV cache, warmup&lt;/td&gt;
&lt;td&gt;~21 s&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;So the 61 seconds &lt;em&gt;includes&lt;/em&gt; compilation and warmup, the entire engine bring-up, right up to serving a request. What it &lt;em&gt;excludes&lt;/em&gt; is the one-time download of the weights from Hugging Face, which you pay once and never again. (These phases come from one representative run and there's some overhead between stages, so they don't sum to exactly 61, but that's where the time lives.)&lt;/p&gt;

&lt;p&gt;This also answers the obvious follow-up, what happens when you scale up? If a new replica boots from an image with the weights baked in, or mounts a shared volume that already has them, it pays this ~61-second load, not a download. A brand-new node with nothing staged would also have to pull the weights first, but well-run setups specifically avoid that, because you don't want every scale event re-downloading 68 GB. So 61 seconds is the realistic number for a properly configured restart or scale-up&lt;/p&gt;

&lt;h2&gt;
  
  
  Warm latency and throughput
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fp5isejarin6bm9q0aass.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fp5isejarin6bm9q0aass.png" alt="Concurrent Benchmarking Of Warm Model" width="645" height="480"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Once the model is loaded, it's a different machine. Warm, the self-hosted MI300X returned the first token in about &lt;strong&gt;322 ms&lt;/strong&gt; and sustained roughly &lt;strong&gt;154 tokens per second&lt;/strong&gt;, and it was remarkably stable, across twenty requests, the spread was about two milliseconds.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkryf08ri5xm8qlnm8vr5.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkryf08ri5xm8qlnm8vr5.png" alt="Benchmarking showcasing VRPAM used" width="800" height="110"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Memory is worth a note, because the raw number looks alarming. The card showed about &lt;strong&gt;173 GB of VRAM in use&lt;/strong&gt;. But the weights themselves are only about 68 GB. vLLM reserves most of the rest up front as KV-cache headroom (roughly 100 GB of it) so it can serve many requests at once. A 120B model doesn't "need" 173 GB; the engine just claims the room ahead of time.&lt;/p&gt;

&lt;p&gt;So the self-hosted trade-off is clean: once it's warm, it's fast, consistent, and entirely yours, but every cold start costs a full minute for our model, and you pay for the &lt;a href="https://www.digitalocean.com/pricing/gpu-droplets" rel="noopener noreferrer"&gt;GPU&lt;/a&gt; whether or not anyone is using it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Does serverless inference have a cold start?
&lt;/h2&gt;

&lt;p&gt;Next I ran the identical test against Serverless Inference. Calling it is straightforward, you create a &lt;a href="https://www.digitalocean.com/community/tutorials/serverless-inference-gradient" rel="noopener noreferrer"&gt;model access key&lt;/a&gt; and hit the &lt;a href="https://www.digitalocean.com/community/tutorials/serverless-inference-openai-sdk" rel="noopener noreferrer"&gt;OpenAI-compatible endpoint&lt;/a&gt;, so the client code is the same and only the base URL changes. I left the endpoint idle first, then measured first-token latency the same way.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqjp5ov87pdm3dyvjqx38.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqjp5ov87pdm3dyvjqx38.png" alt="Concurrent Benchmarking Of Serverless" width="567" height="483"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The first token came back in about &lt;strong&gt;546 ms&lt;/strong&gt;, with a wider spread, anywhere from 446 ms to roughly 1.3 seconds across twenty runs. But there was no spin-up. I ran it twenty times after sitting idle and never caught a cold-start hit.&lt;/p&gt;

&lt;p&gt;Two honest caveats on those numbers. First, the serverless requests travel over the network to DigitalOcean's endpoint, while the GPU test ran locally on the droplet, so some of that extra latency is network distance, not the model being slower. Here's the warm comparison side by side:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Self-hosted MI300X&lt;/th&gt;
&lt;th&gt;Serverless Inference&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Median time to first token&lt;/td&gt;
&lt;td&gt;~322 ms&lt;/td&gt;
&lt;td&gt;~546 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Spread (20 runs)&lt;/td&gt;
&lt;td&gt;~2 ms&lt;/td&gt;
&lt;td&gt;446 ms – 1.3 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Throughput&lt;/td&gt;
&lt;td&gt;~154 tok/s&lt;/td&gt;
&lt;td&gt;(per-token billed)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cold start&lt;/td&gt;
&lt;td&gt;~61 s&lt;/td&gt;
&lt;td&gt;none observed&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;So why no cold start on serverless? It didn't delete the cold start, it absorbed it. DigitalOcean &lt;a href="https://www.digitalocean.com/blog/serverless-inference-deep-dive" rel="noopener noreferrer"&gt;pools GPU capacity across customers&lt;/a&gt;, so the model stayed warm without any effort from me, and I never paid the 61-second hit I took on my own box. The difference is the billing model: serverless isn't charged by the hour, it's charged per token.&lt;/p&gt;

&lt;p&gt;To be fair, serverless &lt;a href="https://www.digitalocean.com/resources/articles/serverless-inference" rel="noopener noreferrer"&gt;isn't immune to cold starts&lt;/a&gt;. If you hit it during a genuinely quiet stretch, you can still catch one. The standard mitigations are sending periodic warm-up requests to keep a worker hot, or designing async-first so a slow first response doesn't matter. In this test I didn't need any of that, it just stayed warm.&lt;/p&gt;

&lt;h2&gt;
  
  
  When is serverless inference cheaper than your own GPU?
&lt;/h2&gt;

&lt;p&gt;This is where the decision actually lives, and it comes down to arithmetic. The GPU is a flat cost, about &lt;a href="https://www.digitalocean.com/pricing/gpu-droplets" rel="noopener noreferrer"&gt;$1.88 an hour&lt;/a&gt; for a single on-demand MI300X, the same whether it serves one request or a million. Serverless is usage-based, gpt-oss-120b is priced at $0.10 per million input tokens and $0.70 per million output tokens, so it costs almost nothing when you're quiet and climbs as you get busier.&lt;/p&gt;

&lt;p&gt;The break-even point is your hourly GPU cost divided by your per-request cost:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;break-even requests/hour = GPU $/hr ÷ [(input_tokens × $0.10/1M) + (output_tokens × $0.70/1M)]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The catch is that the per-request cost depends entirely on how much your model outputs, and that moves the crossover more than you'd expect. I measured it at three response lengths, on the same GPU at the same prices, changing only the output length:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fn2plb496wmer5azexrjs.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fn2plb496wmer5azexrjs.png" alt="Benchmarking" width="586" height="134"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Response type&lt;/th&gt;
&lt;th&gt;Output tokens&lt;/th&gt;
&lt;th&gt;Crossover&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Short (classification / extraction)&lt;/td&gt;
&lt;td&gt;~30&lt;/td&gt;
&lt;td&gt;~18 requests/sec&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Medium (paragraph answer)&lt;/td&gt;
&lt;td&gt;~220&lt;/td&gt;
&lt;td&gt;~3 requests/sec&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Long (code / detailed explanation)&lt;/td&gt;
&lt;td&gt;~1,200&lt;/td&gt;
&lt;td&gt;&amp;lt;1 request/sec&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That's roughly a 25x swing from identical hardware, with nothing changing but response length. Output tokens are the expensive side of the bill, so the chattier your app, the sooner owning a GPU pays for itself.&lt;/p&gt;

&lt;p&gt;In plain terms, if your app sends short, snappy responses, serverless stays cheaper until you're well over a dozen requests per second, nonstop. If it writes long answers, the GPU starts winning below one request per second. (For comparison, &lt;a href="https://docs.digitalocean.com/products/inference/details/pricing/" rel="noopener noreferrer"&gt;Dedicated Inference&lt;/a&gt;, DigitalOcean's managed always-on endpoint, is billed by the GPU-hour like the Droplet but without you managing the it, so its economics sit closer to the self-hosted side of this table than the serverless side.) Drop your own response lengths and GPU rate into the formula and you'll find your exact line.&lt;/p&gt;

&lt;h2&gt;
  
  
  When to use serverless inference, and when not to
&lt;/h2&gt;

&lt;p&gt;No "it depends." Here's the actual call.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;For most teams, serverless is the right default.&lt;/strong&gt; Bursty or spiky traffic, real idle stretches, a dev tool, an internal feature, a side project, anything async where nobody is staring at a spinner on the first request. In all of those, the cold start runs on someone else's pooled capacity, not yours, and you pay nothing while you're quiet. For that kind of traffic, it's almost perfect.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Run the GPU yourself when traffic is steady and high-volume, or when you have a latency SLA you can't miss.&lt;/strong&gt; At that point your traffic rarely stops, so you're not benefiting from serverless's idle savings anyway, and you're using the GPU enough that flat hourly beats per-token. You keep it warm, so the cold start stops mattering. That's not a knock on serverless, it's just the wrong tool for that job.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;If you're in between, start on serverless.&lt;/strong&gt; Watch your token spend, and move to a dedicated GPU the day you cross the line for your response lengths. Don't buy a GPU to solve a problem you don't have yet.&lt;/p&gt;

&lt;h2&gt;
  
  
  Run it yourself
&lt;/h2&gt;

&lt;p&gt;Everything here is reproducible. You can spin up an &lt;a href="https://www.digitalocean.com/community/tutorials/run-gpt-oss-vllm-amd-gpu-droplet-rocm" rel="noopener noreferrer"&gt;AMD GPU Droplet and run gpt-oss-120b on vLLM&lt;/a&gt;, hit the same model on &lt;a href="https://docs.digitalocean.com/products/inference/how-to/si-overview/" rel="noopener noreferrer"&gt;Serverless Inference&lt;/a&gt; with a &lt;a href="https://www.digitalocean.com/community/tutorials/serverless-inference-openai-sdk" rel="noopener noreferrer"&gt;model access key&lt;/a&gt;, and check the &lt;a href="https://docs.digitalocean.com/products/inference/reference/serverless-inference-metrics/" rel="noopener noreferrer"&gt;serverless metrics&lt;/a&gt; and &lt;a href="https://docs.digitalocean.com/products/inference/details/pricing/" rel="noopener noreferrer"&gt;pricing pages&lt;/a&gt; against your own workload.&lt;/p&gt;

&lt;p&gt;Don't take my crossover, run your own. And if you measure a cold start on your own setup, I'd genuinely like to see how the spread looks across different models and hardware.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>The Hidden Cost of Complex AI Platforms: Why Developer Experience Matters</title>
      <dc:creator>Shaoni Mukherjee</dc:creator>
      <pubDate>Thu, 28 May 2026 16:00:00 +0000</pubDate>
      <link>https://dev.to/digitalocean/the-hidden-cost-of-complex-ai-platforms-why-developer-experience-matters-1bb5</link>
      <guid>https://dev.to/digitalocean/the-hidden-cost-of-complex-ai-platforms-why-developer-experience-matters-1bb5</guid>
      <description>&lt;h2&gt;
  
  
  Key Takeaways
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Developer experience is a real cost, not a soft metric&lt;/strong&gt;: Time lost in setup, debugging, and switching tools directly slows down how fast teams can build and iterate.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Most friction comes from fragmented workflows&lt;/strong&gt;: When model hosting, compute, and deployment live in different places, even simple tasks become multi-step processes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Time-to-First-Value (TTFV) is a critical signal:&lt;/strong&gt; The longer it takes to get a working output, the more likely teams are to lose momentum or abandon ideas early.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Scaling introduces a hidden breaking point:&lt;/strong&gt; Moving from a simple API to dedicated infrastructure often forces teams to relearn workflows and rebuild systems.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;This is a systems problem, not a feature gap&lt;/strong&gt;: Many platforms weren’t designed end-to-end, which leads to disconnected experiences as teams grow.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The fastest teams aren’t just using better models&lt;/strong&gt;: They’re working in environments where they can build, test, and scale without constant reconfiguration.&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;The &lt;a href="https://www.digitalocean.com/resources/articles/leading-ai-cloud-providers" rel="noopener noreferrer"&gt;cloud AI platform&lt;/a&gt; ecosystem today looks more powerful than ever, with access to powerful GPUs like NVIDIA H100 and H200, massive libraries of pre-trained models, and full pipelines for &lt;a href="https://www.digitalocean.com/community/tutorials/fine-tuning-llms-on-budget-digitalocean-gpu" rel="noopener noreferrer"&gt;fine-tuning&lt;/a&gt; and inference.&lt;/p&gt;

&lt;p&gt;​​I recently tried deploying a simple inference endpoint for a model. Ideally, it should have taken a few minutes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;provision compute&lt;/li&gt;
&lt;li&gt;load the model&lt;/li&gt;
&lt;li&gt;send a request&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Instead, it took closer to &lt;strong&gt;two hours&lt;/strong&gt; before I got a successful response.&lt;/p&gt;

&lt;p&gt;Not because the model was difficult to run, but because of everything around it:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Figuring out where to start&lt;/li&gt;
&lt;li&gt;No clear documentation&lt;/li&gt;
&lt;li&gt;Generating and configuring the right credentials&lt;/li&gt;
&lt;li&gt;Troubleshooting why the instance wasn’t accessible&lt;/li&gt;
&lt;li&gt;Installing dependencies that weren’t preconfigured&lt;/li&gt;
&lt;li&gt;Retrying after unclear or failed setup steps&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of these steps was particularly complex on its own. But together, they created enough friction to delay even a basic task.&lt;/p&gt;

&lt;p&gt;This pattern shows up often when working with AI platforms today.&lt;/p&gt;

&lt;p&gt;Most discussions focus on visible costs like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Compute pricing&lt;/li&gt;
&lt;li&gt;Storage usage&lt;/li&gt;
&lt;li&gt;API costs&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;But in practice, the higher cost is harder to measure.&lt;/p&gt;

&lt;p&gt;It’s the time spent navigating setup, resolving infrastructure issues, and figuring out how different parts of a platform fit together before any real work begins.&lt;/p&gt;

&lt;h2&gt;
  
  
  The real cost of building AI systems
&lt;/h2&gt;

&lt;p&gt;When teams evaluate AI platforms, the focus usually stays on obvious metrics like compute pricing or model performance. But the actual cost of &lt;a href="https://www.digitalocean.com/community/tutorials/build-ai-agents-the-right-way" rel="noopener noreferrer"&gt;building AI systems&lt;/a&gt; runs much deeper. It shows up in how long it takes to get started, how mentally demanding the platform is, and how much time is lost dealing with infrastructure instead of building products.&lt;/p&gt;

&lt;p&gt;One of the most overlooked factors is &lt;strong&gt;Time-to-First-Value (TTFV)&lt;/strong&gt;, the time it takes to go from signing up on a platform to getting your first meaningful output.&lt;/p&gt;

&lt;p&gt;But when TTFV stretches into hours or even days due to setup issues, unclear steps, or complex configuration, it creates friction right from the start. Developers lose patience, delay experimentation, or abandon the platform altogether. Over time, this directly impacts developer retention and slows down innovation, because fewer ideas make it past the initial stage.&lt;/p&gt;

&lt;h2&gt;
  
  
  Fragmentation: When one platform feels like many
&lt;/h2&gt;

&lt;p&gt;Imagine when a developer tries to log in and finds out multiple logins to separate platforms, which feels not only confusing but also hard to understand. When a single platform feels like multiple disconnected products stitched together.&lt;/p&gt;

&lt;p&gt;On the surface, everything may exist under one umbrella. But once you start using it, the experience tells a different story.&lt;/p&gt;

&lt;h3&gt;
  
  
  Split product surfaces
&lt;/h3&gt;

&lt;p&gt;On platforms like &lt;a href="https://nebius.com/" rel="noopener noreferrer"&gt;Nebius&lt;/a&gt;, you have AI Cloud and Token Factory, which require separate logins; this infrastructure feels like two separate worlds.&lt;/p&gt;

&lt;p&gt;You might provision compute in one place, manage models in another, and handle access or tokens somewhere else entirely. Each part works on its own, but they don’t always feel connected.&lt;/p&gt;

&lt;p&gt;For example, a developer might:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Set up a GPU instance in one interface&lt;/li&gt;
&lt;li&gt;Switch to another section to access models&lt;/li&gt;
&lt;li&gt;Move again to configure authentication or tokens&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Even though it’s technically one platform, it doesn’t feel like a single, cohesive system. This lack of cohesion forces developers to constantly piece together workflows on their own.&lt;/p&gt;

&lt;h3&gt;
  
  
  Confusing navigation
&lt;/h3&gt;

&lt;p&gt;Fragmentation often leads to a simple but frustrating question:&lt;br&gt;
 &lt;strong&gt;“Where do I even start?”&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;When features are spread across different sections or products, developers are left guessing:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Which interface should I use first?&lt;/li&gt;
&lt;li&gt;Where do I run my model?&lt;/li&gt;
&lt;li&gt;Where do I manage credentials or access?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Instead of a clear starting point, the experience becomes exploratory—and not in a good way.&lt;/p&gt;

&lt;p&gt;A common situation is having to jump between different portals just to complete a basic setup. For instance, setting up access in one place and then realizing you need to log into a completely different interface to actually use it.&lt;/p&gt;

&lt;h3&gt;
  
  
  Broken flow
&lt;/h3&gt;

&lt;p&gt;This fragmentation becomes even more apparent when workflows are interrupted.&lt;/p&gt;

&lt;p&gt;Developers may encounter:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Separate logins for different parts of the platform&lt;/li&gt;
&lt;li&gt;Different dashboards that don’t share context&lt;/li&gt;
&lt;li&gt;Disconnected user experiences that don’t carry over progress&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  What fragmentation looks like
&lt;/h3&gt;

&lt;p&gt;A typical workflow, for example, building and deploying an agent, might look simple:&lt;/p&gt;

&lt;p&gt;But instead of happening in a single, continuous flow, each step exists in a different part of the platform.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Compute is managed in one dashboard&lt;/li&gt;
&lt;li&gt;Model configuration happens in another section&lt;/li&gt;
&lt;li&gt;Workflows are defined in a separate interface&lt;/li&gt;
&lt;li&gt;Logs and monitoring are located somewhere else&lt;/li&gt;
&lt;li&gt;Access and credentials are handled independently&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Each step works on its own.&lt;/p&gt;

&lt;h3&gt;
  
  
  The hidden cost
&lt;/h3&gt;

&lt;p&gt;Fragmentation usually doesn’t hurt in the beginning. When a single developer is experimenting, it’s still manageable to move between different sections of a platform and piece things together. The problem starts when the team grows, and the workflow becomes more complex. This typically happens when:&lt;/p&gt;

&lt;p&gt;1) Multiple components like models, agents, and data sources are involved,&lt;/p&gt;

&lt;p&gt;2) More than one developer is working on the system, and&lt;/p&gt;

&lt;p&gt;3) Faster iteration and debugging become important.&lt;/p&gt;

&lt;p&gt;At this stage, constantly switching between interfaces, tools, and dashboards slows everything down because there is no single place to see or manage the full workflow. This issue exists because most platforms are not built as a unified system from the start.&lt;/p&gt;

&lt;p&gt;Fragmentation is not about missing features, but it is about how those features are connected to make it feel like a single system.&lt;/p&gt;

&lt;h2&gt;
  
  
  The anti-developer experience
&lt;/h2&gt;

&lt;p&gt;A common pattern across many AI platforms is asking developers to commit before they’ve had a chance to see real value.&lt;/p&gt;

&lt;p&gt;In some cases, you’re required to add billing details even before running your first model. In others, the free credits are so limited that you can barely complete a meaningful experiment. You might start testing an idea, only to run out of credits halfway through, without fully understanding whether it works.&lt;/p&gt;

&lt;p&gt;This creates psychological friction.&lt;/p&gt;

&lt;p&gt;Instead of freely exploring, developers become cautious. They hesitate to try new models, avoid running multiple experiments, and constantly think about cost rather than creativity. The experience shifts from curiosity to calculation.&lt;/p&gt;

&lt;p&gt;But better-designed platforms take a different approach.&lt;/p&gt;

&lt;p&gt;They give developers enough room to explore properly, sometimes even offering generous free credits, so you can actually spin up resources, run models, and experiment without immediate pressure. You can try things, make mistakes, and learn before worrying about billing.&lt;/p&gt;

&lt;p&gt;Because once developers see something work, they’re far more likely to continue building.&lt;/p&gt;

&lt;h2&gt;
  
  
  The scaling cliff nobody talks about
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://www.digitalocean.com/resources/articles/inference-as-service" rel="noopener noreferrer"&gt;Inference-as-a-service&lt;/a&gt; feels effortless in the beginning. You send a request to an API, get a response, and move on. There is no need to think about infrastructure, scaling, or deployment. This makes it incredibly effective during the early stages, where the focus is on building quickly, experimenting, and testing ideas without friction.&lt;/p&gt;

&lt;p&gt;In this phase, everything works because the system is still small.&lt;/p&gt;

&lt;p&gt;1) The number of requests is low,&lt;/p&gt;

&lt;p&gt;2) Latency is not critical, and&lt;/p&gt;

&lt;p&gt;3) Occasional failures are acceptable.&lt;/p&gt;

&lt;p&gt;The platform handles everything behind the scenes, allowing developers to focus entirely on the product.&lt;/p&gt;

&lt;p&gt;The problem starts when the system begins to grow.&lt;/p&gt;

&lt;p&gt;As usage increases, the same setup is now operating under very different conditions. More users mean more requests, often happening at the same time. Latency is no longer just a technical detail; it becomes part of the user experience. Failures are no longer minor inconveniences; they directly impact reliability.&lt;/p&gt;

&lt;p&gt;This is where cracks begin to appear.&lt;/p&gt;

&lt;h3&gt;
  
  
  A common scaling cliff in inference
&lt;/h3&gt;

&lt;p&gt;A typical early setup looks like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A hosted model endpoint&lt;/li&gt;
&lt;li&gt;Pay-per-request pricing&lt;/li&gt;
&lt;li&gt;No infrastructure management&lt;/li&gt;
&lt;li&gt;Acceptable latency (often in the 300–500 ms range)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;At low to moderate usage, this model works well. Teams can ship quickly, iterate rapidly, and avoid thinking about GPUs or deployment complexity.&lt;/p&gt;

&lt;p&gt;The problem is not at the start, but it emerges when usage becomes &lt;em&gt;predictable and sustained&lt;/em&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Where things start breaking
&lt;/h3&gt;

&lt;p&gt;As request volume grows (for example, into the range of thousands of requests per day), a consistent pattern of issues begins to appear:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Latency variability increases&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Cold starts become more frequent&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://medium.com/javarevisited/mastering-latency-metrics-p90-p95-p99-d5427faea879" rel="noopener noreferrer"&gt;P95&lt;/a&gt; latency spikes unpredictably&lt;/li&gt;
&lt;li&gt;Limited ability to tune performance&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;2. Cost efficiency degrades&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Pay-per-request pricing scales linearly with usage&lt;/li&gt;
&lt;li&gt;No optimization for steady workloads&lt;/li&gt;
&lt;li&gt;The same workload becomes disproportionately expensive&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;3. Lack of capacity guarantees&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;No predictable throughput&lt;/li&gt;
&lt;li&gt;No visibility into resource allocation&lt;/li&gt;
&lt;li&gt;No way to reserve or prioritize compute&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;At this stage, the limitation is not a missing feature but a mismatch between the pricing and deployment model and the workload.&lt;/p&gt;

&lt;h3&gt;
  
  
  The forced transition
&lt;/h3&gt;

&lt;p&gt;The natural next step is moving to dedicated infrastructure.&lt;/p&gt;

&lt;p&gt;In practice, this transition introduces significant complexity:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Selecting GPU types without clear workload mapping&lt;/li&gt;
&lt;li&gt;Configuring deployment environments manually&lt;/li&gt;
&lt;li&gt;Implementing autoscaling policies&lt;/li&gt;
&lt;li&gt;Managing routing, load balancing, and failure handling&lt;/li&gt;
&lt;li&gt;Rebuilding abstractions that were previously handled by the platform&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;What begins as a simple API integration evolves into a full infrastructure problem.&lt;/p&gt;

&lt;h3&gt;
  
  
  The real cost
&lt;/h3&gt;

&lt;p&gt;Teams are forced to shift from:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Product iteration → infrastructure management&lt;/li&gt;
&lt;li&gt;Application logic → deployment tuning&lt;/li&gt;
&lt;li&gt;Fast experimentation → operational maintenance&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This shift directly impacts development velocity.&lt;/p&gt;

&lt;p&gt;In many cases, the bottleneck is no longer model performance or GPU access, but the effort required to operate the system reliably at scale.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why this matters
&lt;/h3&gt;

&lt;p&gt;Inference is often presented as two separate modes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Serverless APIs for getting started&lt;/li&gt;
&lt;li&gt;Dedicated infrastructure for scaling&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;However, the transition between these modes is fragmented.&lt;/p&gt;

&lt;p&gt;This creates a gap where teams:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Overpay for convenience longer than they should&lt;/li&gt;
&lt;li&gt;Delay scaling due to operational complexity&lt;/li&gt;
&lt;li&gt;Or prematurely invest in infrastructure&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The issue is not the availability of tools.&lt;br&gt;
 It is the lack of a smooth, continuous path between them.&lt;/p&gt;

&lt;p&gt;This is a structural problem in the current inference ecosystem — and one that directly impacts how quickly teams can move from prototype to production.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why it feels like a cliff
&lt;/h3&gt;

&lt;p&gt;This shift feels difficult not just because there is more to do, but because the change is abrupt.&lt;/p&gt;

&lt;p&gt;Teams go from a world where everything is abstracted behind a simple API to one where they are responsible for compute, scaling, and reliability. There is no gradual transition between these two states.&lt;/p&gt;

&lt;p&gt;There is no middle layer that offers both simplicity and control.&lt;/p&gt;

&lt;p&gt;That is why it feels like a cliff instead of a smooth progression.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why this happens
&lt;/h3&gt;

&lt;p&gt;This gap exists because platforms are built with different starting points. Inference-focused platforms are designed for simplicity and fast onboarding, so they abstract away infrastructure details. Compute-focused platforms, on the other hand, are built for flexibility and performance, which means they require deeper involvement from the developer.&lt;/p&gt;

&lt;p&gt;Over time, both types of platforms try to expand their capabilities. Inference platforms add more control, and compute platforms add higher-level abstractions. But these additions are layered on top rather than designed as a unified system.&lt;/p&gt;

&lt;p&gt;As a result, the transition between simplicity and control is not seamless.&lt;/p&gt;

&lt;h3&gt;
  
  
  The real impact
&lt;/h3&gt;

&lt;p&gt;This shift usually happens at a critical moment, when the product is gaining traction and needs to scale reliably.&lt;/p&gt;

&lt;p&gt;Instead of focusing on improving the product, teams find themselves dealing with infrastructure, performance issues, and system stability. The pace of development slows down, not because the problem is harder, but because the platform now requires significantly more effort to manage.&lt;/p&gt;

&lt;p&gt;It is what happens when they begin to work at scale, and the platform that once made things easy is no longer enough.&lt;/p&gt;

&lt;h2&gt;
  
  
  What good AI platforms actually look like
&lt;/h2&gt;

&lt;p&gt;After all the friction, the starting problem, platform debugging, understanding the documentation, and platform fragmentation, it is easy to think the problem is missing features, but it's not.&lt;/p&gt;

&lt;p&gt;Most platforms already have the same core capabilities. What actually matters is how much effort it takes to go from an idea to something that works and keep it working as it grows.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Scenario 1&lt;/strong&gt;: Building an AI Agent in an Integrated Workflow&lt;br&gt;
Consider building a simple AI agent or &lt;a href="https://www.digitalocean.com/community/tutorials/build-ai-agent-chatbot-with-gradient-platform" rel="noopener noreferrer"&gt;chatbot&lt;/a&gt; on an integrated platform where models, Knowledge bases, embedding models, and workflows are available in one place.&lt;/p&gt;

&lt;p&gt;A simpler platform will make this process pretty straightforward:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Select the model&lt;/li&gt;
&lt;li&gt;Define the agent logic&lt;/li&gt;
&lt;li&gt;Add appropriate knowledge base&lt;/li&gt;
&lt;li&gt;Add a data source to your knowledge base&lt;/li&gt;
&lt;li&gt;Run a test input&lt;/li&gt;
&lt;li&gt;Make your agent publicly available&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And that's it. What stands out in this setup is not the number of features, but how the flow behaves.&lt;/p&gt;

&lt;p&gt;You don’t need to switch between multiple interfaces to connect components. The model, workflow, and execution are visible in the same place. When you make a change, it reflects immediately without requiring additional setup or restarts.&lt;/p&gt;

&lt;p&gt;If something fails, the issue is tied directly to the step where it happened. You don’t have to search across different dashboards to understand what went wrong.&lt;/p&gt;

&lt;p&gt;The experience feels continuous.&lt;/p&gt;

&lt;p&gt;You start with an idea, implement it, and see the result without getting pulled into infrastructure or configuration issues.&lt;/p&gt;

&lt;p&gt;This is what a unified workflow looks like in practice, not just having all the pieces, but having them work together in a way that reduces effort at every step.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Scenario 2:&lt;/strong&gt; Consider a setup where a team moves from a basic API-based workflow to dedicated inference in order to handle real user traffic more reliably.&lt;/p&gt;

&lt;p&gt;The goal is simple:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Deploy a model with dedicated capacity&lt;/li&gt;
&lt;li&gt;Send requests through a stable endpoint&lt;/li&gt;
&lt;li&gt;Maintain consistent response times&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;What changes in this setup is not the workflow itself, but how predictable it becomes.&lt;/p&gt;

&lt;p&gt;Once the model is deployed on dedicated infrastructure, requests are no longer competing for shared resources. Response times become more consistent, even as usage increases. Instead of worrying about rate limits or sudden slowdowns, the system behaves in a way that is easier to reason about.&lt;/p&gt;

&lt;p&gt;At the same time, the transition does not require rebuilding everything from scratch. The way requests are sent and responses are handled remains familiar. The difference is that there is more control over how the system performs under load.&lt;/p&gt;

&lt;p&gt;If something needs to be adjusted, such as scaling capacity or tuning performance, it can be done without changing the core application logic.&lt;/p&gt;

&lt;p&gt;This is where dedicated inference makes a difference in practice, not by adding complexity, but by making the system more stable as it grows.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;You don’t switch contexts to get basic work done&lt;/strong&gt;
In a well-designed platform, deploying a model, testing it, and monitoring it all happen in one place. You’re not jumping between dashboards, CLI tools, and cloud consoles just to complete a single workflow.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Time-to-First-Value (TTFV) stays consistently low&lt;/strong&gt;
It shouldn’t take hours to figure out how to get a model running. A good platform makes the “first successful response” happen quickly — not just in ideal conditions, but even when you’re unfamiliar with the setup. If you’re spending time debugging environment issues instead of validating outputs, that’s a design failure, not a user error.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The path from prototype to scale doesn’t change shape&lt;/strong&gt;
One of the biggest failure points in current platforms is that the workflow breaks when you scale. A well-designed system keeps the same mental model — the way you deploy and interact with a model at a small scale should still work when traffic increases. You shouldn’t need to relearn everything just to handle more requests.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Infrastructure decisions are abstracted until they actually matter&lt;/strong&gt;
You shouldn’t need to think about GPU types, networking, or provisioning just to test an idea. Good platforms delay these decisions without hiding them completely — they only surface when you have a real reason to care, like optimizing latency or cost.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Failure modes are visible and easy to debug&lt;/strong&gt;
When something breaks, it’s obvious where and why. You’re not digging through multiple systems trying to trace a failed request. Logs, errors, and performance signals are tied directly to the workflow you’re already using, so debugging doesn’t become a separate project.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;The hardest part of building AI systems today isn’t getting access to models or GPUs, but it’s everything that happens around them.&lt;/p&gt;

&lt;p&gt;It’s the time lost moving between tools.&lt;/p&gt;

&lt;p&gt;It’s the friction of stitching together workflows that were never designed to work as one.&lt;br&gt;
It’s the moment when something that worked at a small scale suddenly forces a complete rewrite.&lt;/p&gt;

&lt;p&gt;And most of this doesn’t show up in benchmarks or pricing comparisons. It shows up in delays, workarounds, and abandoned ideas.&lt;/p&gt;

&lt;p&gt;The teams that will win on inference aren’t the ones with the most compute. They’re the ones that can move from idea to working system and then to scale without having to change how they build along the way.&lt;/p&gt;

&lt;p&gt;The real question isn’t which platform has the best features.&lt;/p&gt;

&lt;p&gt;It’s this:&lt;br&gt;
&lt;strong&gt;How many times does your workflow break before you get to something that actually works?&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://medium.com/javarevisited/mastering-latency-metrics-p90-p95-p99-d5427faea879" rel="noopener noreferrer"&gt;Mastering Latency Metrics: P90, P95, P99&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.digitalocean.com/community/tutorials/bulk-inference-content-pipeline-digitalocean-serverless" rel="noopener noreferrer"&gt;How I Built a Content Generation Pipeline Using DigitalOcean Serverless Inference&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.wired.com/story/the-ai-industrys-scaling-obsession-is-headed-for-a-cliff/" rel="noopener noreferrer"&gt;The AI Industry’s Scaling Obsession Is Headed for a Cliff&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>developers</category>
      <category>discuss</category>
      <category>devex</category>
    </item>
    <item>
      <title>How to Deploy Hermes' Self-Improving AI Agent</title>
      <dc:creator>haimantika mitra</dc:creator>
      <pubDate>Tue, 26 May 2026 16:00:00 +0000</pubDate>
      <link>https://dev.to/digitalocean/how-to-deploy-hermes-self-improving-ai-agent-4gm6</link>
      <guid>https://dev.to/digitalocean/how-to-deploy-hermes-self-improving-ai-agent-4gm6</guid>
      <description>&lt;h2&gt;
  
  
  Key takeaways
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;You host Hermes on a DigitalOcean Droplet (Ubuntu 24.04, at least 2 vCPUs and 4 GB RAM), install with the official script, and run &lt;code&gt;hermes setup&lt;/code&gt; for your LLM provider.&lt;/li&gt;
&lt;li&gt;The Telegram bot uses &lt;code&gt;hermes gateway setup&lt;/code&gt; and a persistent gateway service so Hermes stays reachable after you close SSH.&lt;/li&gt;
&lt;li&gt;Skills are Markdown files under &lt;code&gt;~/.hermes/skills/&lt;/code&gt;. MCP servers go in Hermes config and expose tools such as grocery search and cart actions.&lt;/li&gt;
&lt;li&gt;HTTP MCP with OAuth on a headless server uses the URL Hermes prints plus an SSH tunnel from your laptop browser to the Droplet loopback port.&lt;/li&gt;
&lt;li&gt;The grocery walkthrough uses Swiggy Instamart where available; you swap in another MCP URL for your region or stack.&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;2026 is the year of agents doing the task for us rather than just guiding us to. We saw the rise of OpenClaw and now we have &lt;a href="https://hermes-agent.nousresearch.com/docs/" rel="noopener noreferrer"&gt;Hermes agent&lt;/a&gt;. It is an open-source, self-improving AI agent built by &lt;a href="https://nousresearch.com/" rel="noopener noreferrer"&gt;Nous Research&lt;/a&gt;. It has a built-in learning loop that creates skills from experience, improves them during use, nudges itself to persist knowledge, and builds a deepening model of who you are across sessions.&lt;/p&gt;

&lt;p&gt;In this tutorial, you will deploy Hermes on a &lt;a href="https://www.digitalocean.com/products/droplets" rel="noopener noreferrer"&gt;DigitalOcean Droplet&lt;/a&gt;, connect it to Telegram, and extend it with a custom skill. As a practical example, you will see how to build a grocery tracking agent that monitors daily consumption, alerts you when stock is low, and places orders automatically through a grocery delivery service.&lt;/p&gt;

&lt;p&gt;The grocery example uses &lt;a href="https://github.com/Swiggy/swiggy-mcp-server-manifest" rel="noopener noreferrer"&gt;Swiggy Instamart&lt;/a&gt;, which is available in India. But the approach works with any MCP-compatible service, such as a local grocery API, a task manager, a calendar, or a home automation system. The pattern is the same no matter what you connect.&lt;/p&gt;

&lt;h2&gt;
  
  
  Prerequisites
&lt;/h2&gt;

&lt;p&gt;Before you begin, you will need:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A DigitalOcean account. If you do not have one, &lt;a href="https://cloud.digitalocean.com/registrations/new" rel="noopener noreferrer"&gt;sign up here&lt;/a&gt;.
&lt;/li&gt;
&lt;li&gt;A &lt;a href="https://www.digitalocean.com/community/tutorials/initial-server-setup-with-ubuntu" rel="noopener noreferrer"&gt;DigitalOcean Droplet&lt;/a&gt; running Ubuntu 24.04 with at least 2 vCPUs and 4 GB RAM.
&lt;/li&gt;
&lt;li&gt;A Telegram account.
&lt;/li&gt;
&lt;li&gt;API key from an LLM provider. Hermes supports Anthropic, OpenAI, OpenRouter, and others.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What are we building
&lt;/h2&gt;

&lt;p&gt;Hermes is model-agnostic and platform-agnostic. It can connect to any tool that supports the &lt;a href="https://www.digitalocean.com/community/tutorials/control-apps-using-mcp-server" rel="noopener noreferrer"&gt;Model Context Protocol (MCP)&lt;/a&gt;, and it can send and receive messages on Telegram, WhatsApp, Discord, and Slack. You can find more information on their &lt;a href="https://hermes-agent.nousresearch.com/docs/" rel="noopener noreferrer"&gt;documentation&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The architecture in this tutorial looks like this:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fv19jyonedh6059nid2fo.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fv19jyonedh6059nid2fo.png" alt="Architecture" width="800" height="452"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Let’s get started building:&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 1. Create the Droplet
&lt;/h2&gt;

&lt;p&gt;Sign in to your DigitalOcean account and &lt;a href="https://cloud.digitalocean.com/droplets/new" rel="noopener noreferrer"&gt;create a new Droplet&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;On the creation page, select:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Region:&lt;/strong&gt; The one closest to you
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Image:&lt;/strong&gt; Ubuntu 24.04 LTS
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Plan:&lt;/strong&gt; Basic, with at least 2 vCPUs and 4 GB RAM (the &lt;code&gt;s-2vcpu-4gb&lt;/code&gt; size)
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Authentication:&lt;/strong&gt; SSH Key&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you need help adding an SSH key, follow the &lt;a href="https://docs.digitalocean.com/products/droplets/how-to/add-ssh-keys/" rel="noopener noreferrer"&gt;DigitalOcean SSH key guide&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Once the Droplet is created, SSH into it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ssh root@YOUR_DROPLET_IP
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Update the system before installing anything:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;apt update &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; apt upgrade &lt;span class="nt"&gt;-y&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Step 2. Install Hermes Agent
&lt;/h2&gt;

&lt;p&gt;Hermes provides an install script that handles all dependencies including Python, &lt;code&gt;uv&lt;/code&gt;, and the &lt;code&gt;hermes&lt;/code&gt; binary itself.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-fsSL&lt;/span&gt; https://raw.githubusercontent.com/NousResearch/hermes-agent/main/scripts/install.sh | bash
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Once it finishes, reload your shell so the &lt;code&gt;hermes&lt;/code&gt; command is available:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;source&lt;/span&gt; ~/.bashrc
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Run the setup wizard:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;hermes setup
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The wizard will ask you to choose an LLM provider and enter your API key. Select your provider, paste your key, and Hermes will save it to &lt;code&gt;~/.hermes/.env&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Verify the installation:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;hermes &lt;span class="nt"&gt;--version&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Step 3. Connect Hermes to Telegram
&lt;/h2&gt;

&lt;p&gt;Hermes connects to Telegram through a bot. Run the gateway setup:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;hermes gateway setup
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When it asks which platform to use, select Telegram. Hermes will walk you through creating a bot with Telegram's BotFather. Follow the steps in your terminal.&lt;/p&gt;

&lt;p&gt;Once setup is complete, start the gateway:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;hermes gateway start
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;To keep it running after you disconnect from the Droplet, enable it as a system service:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;systemctl &lt;span class="nb"&gt;enable &lt;/span&gt;hermes-gateway
systemctl start hermes-gateway
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Open Telegram, find the bot you just created, and send it a message like "hello." If it responds, Hermes is connected and ready.&lt;/p&gt;

&lt;p&gt;You can now talk to Hermes from anywhere on your phone. Ask it to check the weather, set a reminder, search the web, or run a command on your Droplet. It handles all of this out of the box.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 4. Understand skills and MCP servers
&lt;/h3&gt;

&lt;p&gt;Before building the automation example, let's understand two core Hermes concepts:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Skills&lt;/strong&gt; - These are Markdown files that teach Hermes how to handle specific tasks. You write a skill file describing what the task is, what triggers it, and what steps to follow. Hermes reads all skill files in &lt;code&gt;~/.hermes/skills/&lt;/code&gt; at startup and uses them to handle relevant requests.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;MCP Servers&lt;/strong&gt; - It gives Hermes access to external services. &lt;a href="https://www.digitalocean.com/community/tutorials/model-context-protocol" rel="noopener noreferrer"&gt;MCP (Model Context Protocol)&lt;/a&gt; is an open standard that lets AI agents communicate with APIs in a structured way. If a service publishes an MCP server, Hermes can search its products, manage carts, read calendars, create tasks, and more. You add MCP servers to your Hermes config and Hermes uses them automatically when relevant.&lt;/p&gt;

&lt;p&gt;Together, skills and MCP servers let you build automations that are specific to your life and the tools you use.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 5. Build real-world automation: Grocery tracking
&lt;/h2&gt;

&lt;p&gt;Personally, I am very bad at keeping a tab of groceries and often find myself out of essential stock just when I am about to cook. So I wanted to automate this process for me. I eat similar food everyday and also measure my calories. I used all of this information to automate grocery shopping for me.&lt;/p&gt;

&lt;p&gt;The goal of Hermes was to track my daily grocery consumption, alert me on Telegram when something is running low, and place an order through a grocery delivery service when I say yes.&lt;/p&gt;

&lt;p&gt;This example uses the &lt;a href="https://github.com/Swiggy/swiggy-mcp-server-manifest" rel="noopener noreferrer"&gt;Swiggy Instamart MCP server&lt;/a&gt;, which is available in India. If you are in a different country, you can swap it out for any MCP-compatible grocery or delivery service. The skill logic stays exactly the same.&lt;/p&gt;

&lt;h3&gt;
  
  
  Creating the skill file
&lt;/h3&gt;

&lt;p&gt;Create a directory for the skill:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;mkdir&lt;/span&gt; &lt;span class="nt"&gt;-p&lt;/span&gt; ~/.hermes/skills/grocery
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Create the skill file:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;nano ~/.hermes/skills/grocery/grocery_tracker.md
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Paste the following and edit the &lt;code&gt;groceries:&lt;/code&gt; section to match your actual diet and quantities:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gh"&gt;# Grocery Auto-Order Skill&lt;/span&gt;

&lt;span class="gu"&gt;### My Daily Grocery List&lt;/span&gt;

groceries:
&lt;span class="p"&gt;  -&lt;/span&gt; name: "Milk"
    unit: "ml"
    daily_consumption: 500
    reorder_quantity: 2000
    alert_threshold_days: 3
    search_query: "fresh milk 1 litre"
&lt;span class="p"&gt;
  -&lt;/span&gt; name: "Eggs"
    unit: "pieces"
    daily_consumption: 2
    reorder_quantity: 12
    alert_threshold_days: 3
    search_query: "eggs 12 pack"
&lt;span class="p"&gt;
  -&lt;/span&gt; name: "Oats"
    unit: "gm"
    daily_consumption: 60
    reorder_quantity: 1000
    alert_threshold_days: 7
    search_query: "rolled oats 1kg"

  # Add your own items following the same format

&lt;span class="gu"&gt;### Instructions for Hermes&lt;/span&gt;

&lt;span class="gu"&gt;#### Daily check&lt;/span&gt;
Read ~/.hermes/grocery_inventory.json. Subtract each item's
daily_consumption from its current quantity. For any item where
quantity / daily_consumption is less than or equal to
alert_threshold_days, send a Telegram alert listing the low
items and ask if the user wants to place an order.

&lt;span class="gu"&gt;#### On YES&lt;/span&gt;
Search each low item using its search_query at the user's saved
delivery address. Add reorder_quantity of each to cart. Send a
cart summary via Telegram with prices and total. Wait for CONFIRM.

&lt;span class="gu"&gt;#### On CONFIRM&lt;/span&gt;
Place the order. Update grocery_inventory.json with the restocked
quantities. Send a delivery confirmation message.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Save with &lt;code&gt;Ctrl+O&lt;/code&gt;, then &lt;code&gt;Enter&lt;/code&gt;, then exit with &lt;code&gt;Ctrl+X&lt;/code&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Creating the inventory file
&lt;/h3&gt;

&lt;p&gt;The inventory file tracks how much of each item you have at home right now.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;nano ~/.hermes/grocery_inventory.json
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"last_updated"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"2026-05-07"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"delivery_address"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"YOUR FULL ADDRESS HERE"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"snooze_until"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"items"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"Milk"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"quantity"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;1000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"unit"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"ml"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"Eggs"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"quantity"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;6&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"unit"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"pieces"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"Oats"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"quantity"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;300&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"unit"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"gm"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Replace the quantities with what you actually have at home. Save and exit.&lt;/p&gt;

&lt;h3&gt;
  
  
  Connect to an MCP server
&lt;/h3&gt;

&lt;p&gt;Add the grocery service MCP to your Hermes config:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;hermes config edit
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Scroll to the bottom of the file and add your MCP server. For Swiggy Instamart:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;mcp_servers&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;swiggy-instamart&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;url&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://mcp.swiggy.com/im"&lt;/span&gt;
    &lt;span class="na"&gt;auth&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;oauth&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For any other MCP-compatible service, replace the name and URL with the values from that service's documentation.&lt;/p&gt;

&lt;p&gt;Save and exit, then verify it appears:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;hermes mcp list
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Step 6. Authenticate with your MCP server
&lt;/h2&gt;

&lt;p&gt;Most MCP servers require OAuth authentication. Because your Droplet is headless, the OAuth callback needs to reach your browser through an &lt;a href="https://www.digitalocean.com/community/tutorials/ssh-port-forwarding" rel="noopener noreferrer"&gt;SSH tunnel&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Run the login command on your Droplet:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;hermes mcp login swiggy-instamart
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Hermes will print a URL containing a port number in the redirect URI, like &lt;code&gt;http://127.0.0.1:45123/callback&lt;/code&gt;. Note that port number.&lt;/p&gt;

&lt;p&gt;On your local machine, open a new terminal and run the SSH tunnel using that port:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ssh &lt;span class="nt"&gt;-i&lt;/span&gt; ~/.ssh/id_ed25519 &lt;span class="nt"&gt;-L&lt;/span&gt; 45123:127.0.0.1:45123 root@YOUR_DROPLET_IP
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Keep that terminal open, then immediately open the URL from your Droplet in your browser and complete the login. You have about 30 seconds before the token expires, so have both terminals ready before you start.&lt;/p&gt;

&lt;p&gt;When authentication succeeds, your Droplet terminal will show:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;✓ Authenticated — tools available
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Telling Hermes Your Current Stock
&lt;/h3&gt;

&lt;p&gt;Start Hermes and initialize the inventory:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;hermes
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Type:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Set up my grocery tracker. Ask me for current stock levels for each item in the grocery_tracker skill.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Hermes will ask about each item one by one. Answer with what you currently have at home and it will update &lt;code&gt;grocery_inventory.json&lt;/code&gt; automatically. Once you have given the update, you will receive a message similar to this:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fio61346lvkbkrs9fvvvy.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fio61346lvkbkrs9fvvvy.jpeg" alt="image2" width="708" height="1536"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Setting up the daily alert
&lt;/h3&gt;

&lt;p&gt;Still inside Hermes, set up the cron job:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Add a daily cron at 8am: check my grocery inventory using the grocery_tracker skill, subtract daily consumption, and send me a Telegram message if anything is running low.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Hermes will schedule this and confirm. From now on, every morning it checks your stock and messages you on Telegram if you need to reorder. This is how it looks:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fhcclwz1w6o4zjddrvi1r.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fhcclwz1w6o4zjddrvi1r.jpeg" alt="image3" width="800" height="1482"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 7. Test the full flow
&lt;/h2&gt;

&lt;p&gt;Send a test message to your Telegram bot:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Check my groceries and tell me what's running low.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You should receive a Telegram message like this:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Flblvjle66s7ejl9uvviu.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Flblvjle66s7ejl9uvviu.png" alt="image4" width="800" height="1734"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Reply &lt;strong&gt;YES&lt;/strong&gt;. Hermes searches your connected grocery service, builds a cart, and sends a summary.&lt;/p&gt;

&lt;p&gt;Once everything is set up, you can manage your agent entirely from Telegram:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Message&lt;/th&gt;
&lt;th&gt;What Hermes does&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;Check my groceries&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Shows all items and days of stock remaining&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;I bought 12 eggs&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Updates egg stock in the inventory file&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;Order groceries now&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Skips the check and goes straight to ordering&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;What's running low?&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Lists items near their alert threshold&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;SKIP&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Snoozes today's alert until tomorrow&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;You now have a self-hosted AI agent running on DigitalOcean that works for you around the clock. Hermes connects to Telegram and more messaging platforms so you can reach it from anywhere, and skills plus MCP servers let you extend it to handle almost anything.&lt;/p&gt;

&lt;p&gt;The grocery automation you built in this tutorial is a starting point. The pattern is reusable: write a skill file that describes the task, connect an MCP server that gives Hermes access to the right service, and set a cron job to trigger it automatically. The grocery automation is one example. The same pattern works for any repetitive task in your life. A few ideas to get you started:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Bill reminders-&lt;/strong&gt; Create a skill that tracks recurring payments, calculates due dates, and alerts you three days before each bill is due.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Health tracking-&lt;/strong&gt;  Log your workouts or meals via Telegram and have Hermes summarize your week every Sunday.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Home automation-&lt;/strong&gt; Connect Hermes to a smart home MCP server and have it adjust lights, thermostats, or appliances based on a schedule or your location.&lt;/p&gt;

&lt;h2&gt;
  
  
  Related links
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://github.com/nousresearch/hermes-agent" rel="noopener noreferrer"&gt;Hermes Agent on GitHub&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://modelcontextprotocol.io" rel="noopener noreferrer"&gt;Model Context Protocol documentation&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://docs.digitalocean.com/products/droplets/" rel="noopener noreferrer"&gt;DigitalOcean Droplets documentation&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://docs.digitalocean.com/products/droplets/how-to/add-ssh-keys/" rel="noopener noreferrer"&gt;How to add SSH keys to Droplets&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href="https://core.telegram.org/bots/tutorial" rel="noopener noreferrer"&gt;Telegram BotFather guide&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>hermes</category>
      <category>ai</category>
      <category>agents</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>April 2026 DigitalOcean Tutorials: Inference Optimization and AI Infrastructure</title>
      <dc:creator>DigitalOcean</dc:creator>
      <pubDate>Fri, 22 May 2026 21:55:02 +0000</pubDate>
      <link>https://dev.to/digitalocean/april-2026-digitalocean-tutorials-inference-optimization-and-ai-infrastructure-5fcf</link>
      <guid>https://dev.to/digitalocean/april-2026-digitalocean-tutorials-inference-optimization-and-ai-infrastructure-5fcf</guid>
      <description>&lt;p&gt;Most AI teams hit the same walls once they move past prototyping. The RAG pipeline that worked flawlessly in a demo starts hallucinating under real traffic. Inference costs climb without clear optimization levers. GPU resources sit underutilized while workloads spike elsewhere. &lt;/p&gt;

&lt;p&gt;Most of the time, the root cause traces back to architecture decisions that weren't pressure-tested for production. This month's &lt;a href="https://www.digitalocean.com/community/tutorials" rel="noopener noreferrer"&gt;DigitalOcean tutorials&lt;/a&gt; focus on diagnosing and fixing those failure points across the AI infrastructure stack.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;a href="https://www.digitalocean.com/community/conceptual-articles/why-rag-systems-fail-in-production" rel="noopener noreferrer"&gt;Why RAG Systems Fail in Production&lt;/a&gt;
&lt;/h2&gt;

&lt;p&gt;Why do seemingly solid RAG demos collapse under real-world conditions? This article traces failures back to retrieval quality, latency tradeoffs, and embedding drift. You’ll get a clear picture of how upstream decisions—such as chunking strategy and ranking—directly affect downstream LLM outputs. If your team is building production pipelines, evaluation, monitoring, and retrieval engineering matter just as much as model choice.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fogkanycez0ox4kj1fjmi.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fogkanycez0ox4kj1fjmi.png" alt="RAG System architecture" width="800" height="576"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;a href="https://www.digitalocean.com/community/conceptual-articles/dedicated-vs-serverless-inference-at-scale" rel="noopener noreferrer"&gt;Dedicated vs. Serverless Inference as You Scale&lt;/a&gt;
&lt;/h2&gt;

&lt;p&gt;The choice between serverless and dedicated inference isn't a one-time decision but an evolution driven by how your workload changes over time. Early on, serverless makes sense because traffic is unpredictable and iteration speed matters more than performance optimization. As usage stabilizes, the cracks show up—latency variability frustrates users and per-request pricing gets expensive for always-on systems. Walk-throughs of Modal and Together.ai show where that transition point hits and why delaying it costs you.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;a href="https://www.digitalocean.com/community/conceptual-articles/serverless-fine-tuned-llms" rel="noopener noreferrer"&gt;Fine-Tuned LLMs on Serverless Architecture&lt;/a&gt;
&lt;/h2&gt;

&lt;p&gt;Parameter-efficient methods like LoRA let platforms serve hundreds of fine-tuned model variants from a single GPU by layering small adapter weights on top of a shared frozen base model. This makes serverless, pay-per-token inference possible for custom models without dedicated GPU deployments. The tradeoff is cold starts: idle adapters get evicted from VRAM and need to be reloaded, adding a few hundred milliseconds of latency to the first token. You’ll learn how to minimize that with keep-alive requests, adapter rank tuning, and smarter layer targeting.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;a href="https://www.digitalocean.com/community/tutorials/model-silent-versioning-problem" rel="noopener noreferrer"&gt;The Silent Versioning Problem in AI Inference&lt;/a&gt;
&lt;/h2&gt;

&lt;p&gt;This one is a cautionary tale about what happens when the model behind your endpoint changes and nobody tells you. The serving stack is full of moving parts that can shift independently of the model name, and the result is silent regressions that break prompt tuning and invalidate your evaluations before you even know something moved. It includes a practical buyer's checklist for pressing inference platforms on snapshot pinning, retention commitments, and how they handle disclosure when something in the stack changes.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fai6i5tt8aydh7zx752l8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fai6i5tt8aydh7zx752l8.png" alt="Silent versioning framework" width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;a href="https://www.digitalocean.com/community/conceptual-articles/bottlenecks-llm-inference-optimization" rel="noopener noreferrer"&gt;The Hidden Bottlenecks in LLM Inference and How to Fix Them&lt;/a&gt;
&lt;/h2&gt;

&lt;p&gt;Faster GPUs are not the answer if the rest of your serving stack can't keep up. Spoiler: the bottlenecks are GPU underutilization from rigid batching, memory bandwidth constraints during decode, KV cache fragmentation, and CPU-side overhead from tokenization and prompt assembly. Click through for a deeper look at each one and practical fixes.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;a href="https://www.digitalocean.com/community/tutorials/ai-platform-security-review" rel="noopener noreferrer"&gt;We Built a Private-Document AI App to Test Platform Security. Here Is What We Could Actually Verify&lt;/a&gt;
&lt;/h2&gt;

&lt;p&gt;AI security should always be treated as a first-class concern, not an afterthought. This tutorial puts that to the test by building a private-document chatbot and running the same workflow across six inference platforms: DigitalOcean, Baseten, Nebius, Fireworks AI, Modal, and Together AI. Each platform is evaluated on access controls, data retention defaults, network isolation, audit logging, and shared responsibility clarity. It doubles as a practical framework for figuring out what you can actually verify before sensitive data is in flight.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;a href="https://www.digitalocean.com/community/tutorials/post-inference-querying-mongodb" rel="noopener noreferrer"&gt;Post-Inference Storage and Querying with MongoDB&lt;/a&gt;
&lt;/h2&gt;

&lt;p&gt;Many inference tutorials stop at the model response. This one keeps going. You'll build a FastAPI app that sends images through a vision model, stores the structured predictions in MongoDB, and then exposes endpoints that let you filter by detected labels and confidence scores or run aggregation pipelines across your full dataset. It's a practical blueprint for turning raw model output into something queryable and operational.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;a href="https://www.digitalocean.com/community/tutorials/how-to-build-multi-agent-ai-system-docker-agent-digitalocean" rel="noopener noreferrer"&gt;How to Build a Multi-Agent AI System with Docker and DigitalOcean&lt;/a&gt;
&lt;/h2&gt;

&lt;p&gt;Instead of routing everything through a single model, multi-agent systems let you split a workflow across specialized agents that each handle a different part of the problem and pass results between them. The tradeoff is coordination complexity. This walkthrough covers how to containerize each agent with Docker, manage communication between them, and deploy the full system on DigitalOcean. You'll come away with a working deployment pattern you can adapt to your own orchestration needs.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;a href="https://www.digitalocean.com/community/tutorials/gpu-fleet-optimizer" rel="noopener noreferrer"&gt;Building an AI-Powered GPU Fleet Optimizer with the DigitalOcean AI Platform ADK&lt;/a&gt;
&lt;/h2&gt;

&lt;p&gt;A single idle GPU Droplet left running overnight can add hundreds of dollars to your monthly bill, and standard CPU monitoring won't catch it because it can't see whether the GPU is actually doing work. This tutorial builds an AI-powered agent using the DigitalOcean AI Platform ADK that scrapes NVIDIA DCGM metrics like VRAM usage, engine utilization, and power draw across your fleet in real time. It compares those metrics against configurable thresholds to flag idle resources before they inflate your cloud spend. The repo is designed to be forked and customized to your own workloads, including adding tools that let the agent take action like powering off idle nodes.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>infrastructure</category>
      <category>tutorial</category>
      <category>learning</category>
    </item>
    <item>
      <title>Tutorial: This AI Now Tells You if a Meeting Could Be an Email</title>
      <dc:creator>Andrew Dugan</dc:creator>
      <pubDate>Thu, 21 May 2026 16:00:00 +0000</pubDate>
      <link>https://dev.to/digitalocean/tutorial-this-ai-now-tells-you-if-a-meeting-could-be-an-email-2m3f</link>
      <guid>https://dev.to/digitalocean/tutorial-this-ai-now-tells-you-if-a-meeting-could-be-an-email-2m3f</guid>
      <description>&lt;h2&gt;
  
  
  Key Takeaways
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;DigitalOcean's Inference Router semantically routes prompts to the most appropriate model based on custom instructions. The setup process is 'point-and-click', with no hardcoded "if/else" logic required.&lt;/li&gt;
&lt;li&gt;The router is built directly into the inference pipeline. Users can make inference requests normally, and the router automatically handles the workflow.&lt;/li&gt;
&lt;li&gt;In our workflow, it determines the nature of the task and routes the request to a cheaper, faster model to write an email or a larger, more advanced model to write a meeting agenda. This architecture can scale beyond meetings and can be used for support tickets, code reviews, legal documents, and more.&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;Think back to the last time you received a calendar invite with no agenda, 12 attendees, and a title that says "Quick Sync". We've all either held or attended meetings that "could have been an email" at some point, but what if there was a way to have a gentle nudge built straight into your workflow that only leads us into a meeting when the task requires it. Instead of defaulting to a meeting, one could describe the details of the task that needs to be addressed, and immediately either an email is written for you to send out or a meeting agenda is written ready to attach to your calendar invites. To take it a step further, emails and meeting agendas require different levels of depth and consideration, and ultimately different &lt;a href="https://www.digitalocean.com/resources/articles/large-language-models" rel="noopener noreferrer"&gt;LLMs&lt;/a&gt; to write them.&lt;/p&gt;

&lt;p&gt;We've built exactly this using DigitalOcean's new &lt;a href="https://docs.digitalocean.com/products/inference/how-to/use-inference-router/" rel="noopener noreferrer"&gt;Inference Router&lt;/a&gt;, a policy-driven routing layer that matches each incoming prompt to the right model based on task complexity without hardcoded "if/else" logic required. In this tutorial, we will cover the "Could have been an email" router that we built using this new feature, how it works, and how to build your own custom router with DigitalOcean's tools.&lt;/p&gt;

&lt;h2&gt;
  
  
  How the router works with DigitalOcean
&lt;/h2&gt;

&lt;p&gt;Traditional LLM (large language model) inference involves sending a request to a single model and getting a response. The better or worse the model, the better or worse the response. LLM routers are a layer in between you and a group of models that takes your request, identifies the best model for the request, and has that specific model handle it. Routers can be customized to choose models based on speed, price, specific task, or any other optimization you are looking for. It allows teams to set up a single endpoint for a wide range of needs while getting the best possible price and speed for each request.&lt;/p&gt;

&lt;p&gt;In our case, we built a router with two tasks. The first task we made is &lt;code&gt;write_email&lt;/code&gt;. It is backed by a cheap, fast model (&lt;a href="https://huggingface.co/meta-llama/Llama-3.3-70B-Instruct" rel="noopener noreferrer"&gt;Llama 3.3 Instruct 70B&lt;/a&gt;) for writing a simple email. The second task is &lt;code&gt;write_meeting_agenda&lt;/code&gt;. It is backed by a frontier model (Anthropic &lt;a href="https://www.anthropic.com/news/claude-opus-4-7" rel="noopener noreferrer"&gt;Claude Opus 4.7&lt;/a&gt;) to create a detailed meeting plan to discuss decisions that genuinely require talking to each other. In the request, you describe what you need done, the topic, the stakeholders, and any agenda items, and the router reads that description, matches it against the task definitions, and routes it to whichever model fits. If the request lands on the &lt;code&gt;write_email&lt;/code&gt; task, the router delivers a verdict of "this could be an email" and generates a ready-to-send email draft. If it lands on &lt;code&gt;write_meeting_agenda&lt;/code&gt;, the app confirms the meeting is warranted and produces a structured agenda with talking points and action items. The routing decision itself is the verdict. No additional classification logic is needed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 1. Build the router
&lt;/h2&gt;

&lt;p&gt;The first step to building a router is to log in to your DigitalOcean cloud account, or create an account if you don't have one already. Navigate to the &lt;a href="https://cloud.digitalocean.com/model-studio/router/" rel="noopener noreferrer"&gt;router page&lt;/a&gt; and select "Create Router".&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdoimages.nyc3.cdn.digitaloceanspaces.com%2F010AI-ML%2F2025%2FAndrew%2F19_Meeting_or_Email%2F1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdoimages.nyc3.cdn.digitaloceanspaces.com%2F010AI-ML%2F2025%2FAndrew%2F19_Meeting_or_Email%2F1.png" title="Create a Router in the DigitalOcean Control Panel" alt="The DigitalOcean Create a Router page showing the name and description fields" width="800" height="456"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;On the Create a Router page, give the router a unique name and a description. That description is not just metadata. It serves as a routing prompt, giving the router overall context so it can identify the most appropriate task for each incoming request. From there you define the tasks that make up the router's logic. Each task combines a name, a description, and a model pool with a selection policy. You can either add pre-configured tasks that DigitalOcean has already benchmarked and optimized, or define fully custom tasks that specify exactly which models to use and how to rank them, whether by cost efficiency, speed (&lt;a href="https://www.digitalocean.com/blog/llm-inference-benchmarking" rel="noopener noreferrer"&gt;Time To First Token&lt;/a&gt;), or a manual ranking you control.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdocs.digitalocean.com%2Fscreenshots%2Finference%2Fadd-custom-task.fc4a50918dd6700b40e7f2bdca0e20b358a2f6b7322dc31921cbe8d3f448a21b.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdocs.digitalocean.com%2Fscreenshots%2Finference%2Fadd-custom-task.fc4a50918dd6700b40e7f2bdca0e20b358a2f6b7322dc31921cbe8d3f448a21b.png" title="Add a custom task to the router" alt="The Add Custom Task dialog showing task name, description, and model pool fields" width="800" height="459"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Once your tasks are in place, the last piece is specifying fallback models. Fallback models catch any request that does not cleanly match one of your configured tasks, and they are tried in the priority order you set. This gives the router a safety net so that even if the incoming prompt is ambiguous or outside the scope of your named tasks, a response is still generated rather than failing silently. For our email/meeting router, that means a borderline "is this a meeting or an email?" input never goes unanswered.&lt;/p&gt;

&lt;p&gt;If you prefer automation over the control panel, you can also create the router with a single POST request to &lt;code&gt;https://api.digitalocean.com/v2/gen-ai/models/routers&lt;/code&gt;, passing in the same names, task definitions, selection policies, and fallback models as a JSON body, which is also useful for version-controlling your router alongside your application code.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 2. Build the app
&lt;/h2&gt;

&lt;p&gt;With the router created, integrating it into an application is straightforward because the router is a drop-in replacement for any direct model call. You use the same Chat Completions endpoint (&lt;code&gt;https://inference.do-ai.run/v1/chat/completions&lt;/code&gt;) and the same request shape, but instead of naming a specific model you prefix your router's name with &lt;code&gt;router:&lt;/code&gt; in the &lt;code&gt;model&lt;/code&gt; field. For this app, the field would look like &lt;code&gt;"model": "router:meeting-or-email"&lt;/code&gt;. Authentication works the same way. You generate a Model Access Key from the DigitalOcean Control Panel, export it as &lt;code&gt;MODEL_ACCESS_KEY&lt;/code&gt;, and pass it as a Bearer token in your request header. The user's meeting description, agenda, and attendee list become the message content, and the router takes it from there.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;label&lt;/span&gt; &lt;span class="n"&gt;meeting_or_email&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;py&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;requests&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;meeting_or_email&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;user_input&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;url&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://inference.do-ai.run/v1/chat/completions&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="n"&gt;headers&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Content-Type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;application/json&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Authorization&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Bearer &amp;lt;^&amp;gt;YOUR_MODEL_ACCESS_KEY&amp;lt;^&amp;gt;&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="n"&gt;data&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;model&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;router:meeting-or-email&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;messages&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
            &lt;span class="p"&gt;{&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;system&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
                    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;You are a workplace productivity assistant that evaluates whether a task &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
                    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;requires a live meeting or can be handled asynchronously via email. &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
                    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;If the request involves a straightforward update, announcement, or single-topic &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
                    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;communication with no real-time decision-making needed, write a concise, &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
                    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;professional email draft and state that this could have been an email. &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
                    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;If the request requires discussion, real-time collaboration, debate, or &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
                    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;coordination among multiple stakeholders with competing priorities, produce &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
                    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;a structured meeting agenda with talking points and action items, and confirm &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
                    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;that a meeting is warranted. Always begin your response by clearly stating &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
                    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;your verdict: &lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;This could be an email.&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt; or &lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;This warrants a meeting.&lt;/span&gt;&lt;span class="sh"&gt;'"&lt;/span&gt;
                &lt;span class="p"&gt;),&lt;/span&gt;
            &lt;span class="p"&gt;},&lt;/span&gt;
            &lt;span class="p"&gt;{&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;user_input&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="p"&gt;}&lt;/span&gt;
        &lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;requests&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;url&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;headers&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;response_body&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Model: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;response_body&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;model&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Message: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;response_body&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;choices&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;message&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;model&lt;/code&gt; field in the response body tells you exactly which model the router selected for that request. Requests the router judged as routine land on the cheaper, faster model, while requests it judged as genuinely complex land on the frontier model. The &lt;code&gt;x-model-router-selected-route&lt;/code&gt; response header tells you which task was matched, for example &lt;code&gt;write_email&lt;/code&gt; vs &lt;code&gt;write_meeting_agenda&lt;/code&gt;, or &lt;code&gt;fallback&lt;/code&gt; if none of the tasks matched. The app does not need any if/else logic to decide what kind of meeting it is. It reads the header the router already populated and maps it to a verdict message for the user.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nf"&gt;meeting_or_email&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;I need to plan a large event with multiple stakeholders that will all be involved.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[secondary_label Output]
Model: anthropic-claude-opus-4.7
Message: This warrants a meeting.

Coordinating a large event with multiple stakeholders involves competing priorities, real-time negotiation of responsibilities, and collaborative decision-making that simply cannot be handled efficiently via email threads. Below is a structured agenda to make the meeting productive.

---

## Event Planning Kickoff Meeting
...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You can see above that with a large project the task is routed to Opus 4.7. With a smaller task that just warrants an email, below, the task is routed to Llama3.3.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nf"&gt;meeting_or_email&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;I have some metrics I want to share with my team.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[secondary_label Output]
Model: llama3.3-70b-instruct
Message: This could be an email. 

Here's a draft email you could send to your team:

Subject: Update on Key Metrics

Dear Team,
...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Step 3. Deploy to DigitalOcean App Platform
&lt;/h2&gt;

&lt;p&gt;Before deploying your own router, it is worth spending a few minutes in the Inference Router playground to validate that the router is routing the way you expect. From the &lt;code&gt;My Routers&lt;/code&gt; tab, click the menu next to your router and select a model to compare it against. The Playground opens in a split view where you can type a meeting description and see both the router's response and the comparison model's response side by side. Each result shows the cost difference, end-to-end latency, the specific model the router selected, and the task that was matched for that query. This is a useful check to confirm that your task descriptions are correctly discriminating between routine syncs and complex-coordination requests before any real traffic hits the router.&lt;/p&gt;

&lt;p&gt;Once deployed, the Analyze tab gives you a live view of how the router is performing in production. You can see aggregate metrics across all your routers or drill into a specific one, including total requests, total token usage, model match rate, and fallback rate. Model match rate is the percentage of requests matched to a configured task, and fallback rate is the percentage that fell through to the fallback models instead. For accuracy evaluation, the Router Evaluation tool in the Playground tab lets you upload a labeled dataset and run an LLM-as-a-Judge evaluation that scores responses on completeness, correctness, token usage, and latency. Together these two views give you what you need to iterate on your task descriptions and model pools after launch as you accumulate real meeting data.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;The meeting app we built is a thin wrapper around a genuinely powerful idea. You do not have to choose which model handles a request, you just have to describe the conditions under which each model makes sense and let the router enforce those conditions at runtime. The router does not just save money on tokens. It changes how you think about designing for complexity. Instead of building one prompt that works adequately for everything, you build narrow, well-described task buckets and let semantic matching handle the dispatch.&lt;/p&gt;

&lt;p&gt;The broader lesson here extends well beyond meetings and emails. The same pattern applies anywhere you have a mix of requests hitting a single endpoint. This could include a customer support queue where most tickets are simple FAQs but a few require nuanced reasoning, a code review pipeline where style fixes and architecture feedback warrant very different models, or a legal document classifier where boilerplate and novel clauses should not cost the same to process. Once you have written a router description and a pair of task definitions, you have infrastructure that scales horizontally without adding branching logic to your application code. DigitalOcean's &lt;a href="https://www.digitalocean.com/" rel="noopener noreferrer"&gt;platform&lt;/a&gt; keeps that infrastructure on one bill and one security model, which removes the operational overhead that typically discourages teams from adopting multi-model strategies in the first place.&lt;/p&gt;

&lt;h2&gt;
  
  
  Related links
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://docs.digitalocean.com/products/inference/how-to/use-inference-router/" rel="noopener noreferrer"&gt;How to Use Inference Router&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.digitalocean.com/community/tutorials/how-to-build-parallel-agentic-workflows-with-python" rel="noopener noreferrer"&gt;How to Build Parallel Agentic Workflows with Python&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.digitalocean.com/community/tutorials/mistral-7b-fine-tuning" rel="noopener noreferrer"&gt;Fine-Tune Mistral-7B with LoRA: A Quickstart Guide&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>tutorial</category>
      <category>agentskills</category>
      <category>inference</category>
    </item>
    <item>
      <title>Tutorial: Build a Cost-Aware AI Support Triage API</title>
      <dc:creator>James Skelton</dc:creator>
      <pubDate>Tue, 19 May 2026 23:10:07 +0000</pubDate>
      <link>https://dev.to/digitalocean/tutorial-build-a-cost-aware-ai-support-triage-api-24m5</link>
      <guid>https://dev.to/digitalocean/tutorial-build-a-cost-aware-ai-support-triage-api-24m5</guid>
      <description>&lt;h2&gt;
  
  
  Key takeaways
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;AI applications use a single endpoint to handle multiple complex tasks: classification, urgency scoring, customer-facing drafting, and long-form summarization. &lt;/li&gt;
&lt;li&gt;This does not account for varying cost, latency, and quality requirements. &lt;/li&gt;
&lt;li&gt;Building a FastAPI and using serverless inference infrastructure makes it possible to address these requirements through effective routing.
&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;Most AI applications start with a single model hard-coded into the app. That works well for a prototype, but it breaks down the moment a single endpoint has to handle multiple complex task categories: classification, urgency scoring, customer-facing drafting, and long-form summarization all benefit from different model choices. Those tasks do not share the same cost, latency, or quality requirements.&lt;/p&gt;

&lt;p&gt;Support triage is the cleanest example of this. A user types "how do I reset my password?" and you spend the same per-token rate as you do on a multi-paragraph escalation from an enterprise customer with logs pasted in. You can branch on ticket type in your app code and pick a different model per branch, but now your model selection logic lives inside your handler, your fallback strategy is a try/except, and every pricing change means a redeploy. The consequences include a 70B model classifying one-word tickets, no fallback when that model is slow, and a redeploy every time pricing shifts.&lt;/p&gt;

&lt;p&gt;In this tutorial, we'll use &lt;a href="https://docs.digitalocean.com/products/inference/how-to/use-serverless-inference/" rel="noopener noreferrer"&gt;serverless inference via DigitalOcean's Inference router&lt;/a&gt; to easily and quickly build a &lt;a href="https://fastapi.tiangolo.com/" rel="noopener noreferrer"&gt;FastAPI&lt;/a&gt; support triage endpoint that solves all these problems at once. By the end, you'll route classification, urgency scoring, customer replies, and escalation summaries to the right model for each job — automatically, with built-in fallback, and without a single model name in your application code. You'll have a production-ready API that's 71% cheaper than running everything on a frontier model.&lt;/p&gt;

&lt;h2&gt;
  
  
  What you're building
&lt;/h2&gt;

&lt;p&gt;Let's construct a single endpoint, &lt;code&gt;POST /triage&lt;/code&gt;, that takes a ticket payload and returns:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Classification: the issue category (billing, bug, how-to, account, etc.)&lt;/li&gt;
&lt;li&gt;Urgency + sentiment: a severity score and a read on customer mood&lt;/li&gt;
&lt;li&gt;Drafted reply: a short, customer-facing response&lt;/li&gt;
&lt;li&gt;Escalation summary: a structured brief for a human agent, generated only when the ticket is complex enough to need one&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The architecture moves from this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;App → hardcoded model &lt;span class="o"&gt;(&lt;/span&gt;one model handles every task&lt;span class="o"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;to this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;App → Serverless inference via Inference Router → best-fit model per task
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The router is what makes the second diagram possible without your app knowing anything about which models exist.&lt;/p&gt;

&lt;h2&gt;
  
  
  Serverless Inference and DigitalOcean's Inference Router
&lt;/h2&gt;

&lt;p&gt;The &lt;a href="https://www.digitalocean.com/products/inference-engine" rel="noopener noreferrer"&gt;Inference Router&lt;/a&gt; lets you define tasks and model pools, then routes incoming prompts to the best-fit model based on those task definitions and selection policies. A task is a named job with a description: "&lt;code&gt;classify_ticket&lt;/code&gt;, for example. A model pool is the set of candidate models the router can choose from for that task, governed by a selection policy: lowest cost, lowest latency, a manually set ranking, or a fallback order. You configure all of this once at the router level, and your app calls the router instead of any specific model."&lt;/p&gt;

&lt;p&gt;Serverless inference lets you send API requests to models without having to create an AI agent or worry about managing infrastructure. This allow you to get started quickly without managing any components behind an inference endpoint.&lt;/p&gt;

&lt;p&gt;The API surface is OpenAI-compatible. The base URL is &lt;code&gt;https://inference.do-ai.run/v1/&lt;/code&gt;, and a single model access key covers both foundation models and routers.&lt;/p&gt;

&lt;h2&gt;
  
  
  Project setup
&lt;/h2&gt;

&lt;p&gt;In order to continue, you need Python 3.10+, a &lt;a href="https://cloud.digitalocean.com/login" rel="noopener noreferrer"&gt;DigitalOcean account&lt;/a&gt; with &lt;a href="https://docs.digitalocean.com/products/inference/how-to/use-serverless-inference/" rel="noopener noreferrer"&gt;Serverless Inference&lt;/a&gt; enabled, and a &lt;a href="https://docs.digitalocean.com/products/inference/how-to/model-access-keys/" rel="noopener noreferrer"&gt;model access key&lt;/a&gt;. We have already configured the &lt;a href="https://github.com/Jameshskelton/triage_app" rel="noopener noreferrer"&gt;full project in this repository&lt;/a&gt; for your convenience, but follow along in this next section to build out your own version of the API and learn why we made specific choices for the API.&lt;/p&gt;

&lt;p&gt;The project layout is intentionally small:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;support-triage/
├── main.py
├── sample_tickets.json
├── requirements.txt
└── .env
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;main.py&lt;/code&gt; holds the application code, &lt;code&gt;requirements.txt&lt;/code&gt; the required packages,  &lt;code&gt;sample_tickets.json&lt;/code&gt; is a sample for testing the router, and &lt;code&gt;.env&lt;/code&gt; holds the required secrets, keys, and URL base values.&lt;/p&gt;

&lt;p&gt;To get started, clone the repo onto your machine and install everything by pasting the following into your terminal:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/Jameshskelton/triage_app
&lt;span class="nb"&gt;cd &lt;/span&gt;triage_app
python3 &lt;span class="nt"&gt;-m&lt;/span&gt; venv venv_triage
&lt;span class="nb"&gt;source &lt;/span&gt;venv_triage/bin/activate
pip &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-r&lt;/span&gt; requirements.txt
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The OpenAI SDK works as-is for DigitalOcean's Serverless Inference: you just point base_url and api_key at DigitalOcean instead of OpenAI.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 1: The baseline - direct model calls
&lt;/h3&gt;

&lt;p&gt;Before we touch the router, let's build the version most developers would write first: one model, hardcoded, doing all four jobs. The next few step sections outline the work we did to build the application demo. If you would like to just test the final version, check out our &lt;a href="https://github.com/Jameshskelton/triage_app" rel="noopener noreferrer"&gt;repository&lt;/a&gt; where we stored this project.&lt;/p&gt;

&lt;p&gt;To get started, we created &lt;code&gt;main.py&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;re&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;openai&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;OpenAI&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;fastapi&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;FastAPI&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pydantic&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;BaseModel&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;dotenv&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;load_dotenv&lt;/span&gt;

&lt;span class="nf"&gt;load_dotenv&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;base_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;environ&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;DO_INFERENCE_BASE_URL&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;environ&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;DO_MODEL_ACCESS_KEY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;MODEL&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;llama3.3-70b-instruct&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;  &lt;span class="c1"&gt;# one model for everything
&lt;/span&gt;
&lt;span class="n"&gt;app&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;FastAPI&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;Ticket&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;BaseModel&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;subject&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;
    &lt;span class="n"&gt;body&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;call_model&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;system&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;user&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;resp&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;MODEL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;
            &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;system&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;system&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
            &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;user&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
        &lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;resp&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;choices&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;strip&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="nd"&gt;@app.post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/triage&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;triage&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ticket&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;Ticket&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;text&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Subject: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;ticket&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;subject&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="se"&gt;\n\n&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;ticket&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;body&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

    &lt;span class="n"&gt;category&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;call_model&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Classify this support ticket into one of: billing, bug, how-to, account, other. Reply with one word.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;urgency&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;call_model&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Score urgency from 1 (low) to 5 (critical) and note sentiment. Reply as &lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;score: N, sentiment: X&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;reply&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;call_model&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Write a short, professional reply to this customer. Maximum 4 sentences.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;summary&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;call_model&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Summarize this ticket for a human agent. Include the problem, what&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;s been tried, and recommended next steps.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;category&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;category&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;urgency&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;urgency&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;reply&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;reply&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;escalation_summary&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;summary&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If we have set up our &lt;code&gt;.env&lt;/code&gt; file correctly with the right API keys and values, we can run it using the code below:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;uvicorn main:app &lt;span class="nt"&gt;--reload&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Let’s test it with two tickets (one trivial, one complex), and audit the results.&lt;/p&gt;

&lt;p&gt;Example input 1:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-X&lt;/span&gt; POST localhost:8000/triage &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
  "subject": "Password reset",
  "body": "How do I reset my password?"
}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Example output 1:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="o"&gt;{&lt;/span&gt;
  &lt;span class="s2"&gt;"category"&lt;/span&gt;: &lt;span class="s2"&gt;"account"&lt;/span&gt;,
  &lt;span class="s2"&gt;"urgency"&lt;/span&gt;: &lt;span class="s2"&gt;"score: 1, sentiment: neutral"&lt;/span&gt;,
  &lt;span class="s2"&gt;"reply"&lt;/span&gt;: &lt;span class="s2"&gt;"You can reset your password by selecting the Forgot password link on the sign-in page and following the email instructions. If you do not receive the reset email, check your spam folder or contact support for help."&lt;/span&gt;,
  &lt;span class="s2"&gt;"escalation_summary"&lt;/span&gt;: &lt;span class="s2"&gt;"The customer is asking how to reset their password. No signs of account compromise, outage, or escalation risk. Recommended next step: provide standard password reset instructions."&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Example input 2:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-X&lt;/span&gt; POST localhost:8000/triage &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
  "subject": "Production outage on enterprise account",
  "body": "Our team has been unable to access the dashboard since 09:14 UTC. We have ~200 internal users blocked. Attached are logs showing 502s from the API gateway..."
}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This gives us something like the corresponding example output 2:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="o"&gt;{&lt;/span&gt;
  &lt;span class="s2"&gt;"category"&lt;/span&gt;: &lt;span class="s2"&gt;"bug"&lt;/span&gt;,
  &lt;span class="s2"&gt;"urgency"&lt;/span&gt;: &lt;span class="s2"&gt;"score: 5, sentiment: frustrated"&lt;/span&gt;,
  &lt;span class="s2"&gt;"reply"&lt;/span&gt;: &lt;span class="s2"&gt;"Thank you for reporting this. We understand that a production dashboard outage affecting around 200 users is urgent, and we are escalating this to our engineering team immediately. Please continue to share any relevant logs or timestamps while we investigate."&lt;/span&gt;,
  &lt;span class="s2"&gt;"escalation_summary"&lt;/span&gt;: &lt;span class="s2"&gt;"Enterprise customer reports a production dashboard outage beginning at 09:14 UTC. Approximately 200 internal users are blocked. Logs indicate 502 responses from the API gateway. Recommended next steps: escalate to engineering, inspect gateway and upstream service health, correlate errors around 09:14 UTC, and provide the customer with frequent status updates."&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Both responses are useful. That is exactly why this baseline is tempting.&lt;/p&gt;

&lt;p&gt;But look at what just happened: the same 70B model handled everything. The model classified "How do I reset my password?" into a simple category, scored urgency, drafted a short reply, and wrote an escalation summary that the ticket did not really need. Then it handled the enterprise outage, where the larger model actually makes sense.&lt;/p&gt;

&lt;p&gt;That is the problem. The trivial ticket and the production outage have very different cost, latency, and quality requirements, but the app treats them the same. You are paying overkill rates for simple work, there is no fallback if the model is slow or unavailable, and any model-selection change means editing application code and redeploying. Let's fix that.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 2: Configure the Inference Router
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdoimages.nyc3.cdn.digitaloceanspaces.com%2F010AI-ML%2F2026%2FJames%2Finference%2520router.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdoimages.nyc3.cdn.digitaloceanspaces.com%2F010AI-ML%2F2026%2FJames%2Finference%2520router.gif" alt="image" width="" height=""&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;In the &lt;a href="https://cloud.digitalocean.com" rel="noopener noreferrer"&gt;DigitalOcean control panel&lt;/a&gt;, navigate to the Inference Router using the left-hand sidebar. Then, create a new Inference Router. Name your Router appropriately, and give it a descriptive description of what it will do. For example, we named ours &lt;code&gt;triage-router&lt;/code&gt;, and described it as “Demo Triage API for DO tutorial”.&lt;/p&gt;

&lt;p&gt;The router then needs its four tasks, each with a description and a model pool with a selection policy. Each of these is outlined below. If you want to copy them to recreate this experiment, copy and paste the values within to the Router tasks individually. This will make probabilistically similar results to what we have.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fskqr7zhd8cvix23kr48e.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fskqr7zhd8cvix23kr48e.png" alt="image" width="800" height="1443"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Task name&lt;/th&gt;
&lt;th&gt;Description (fed to the router)&lt;/th&gt;
&lt;th&gt;Model pool strategy&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;classify_ticket&lt;/td&gt;
&lt;td&gt;Categorize short support messages into issue types (billing, bug, how-to, account).&lt;/td&gt;
&lt;td&gt;Lowest cost&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;urgency_detection&lt;/td&gt;
&lt;td&gt;Detect severity, sentiment, and escalation risk in a single pass.&lt;/td&gt;
&lt;td&gt;Lowest latency&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;draft_customer_reply&lt;/td&gt;
&lt;td&gt;Generate a short, professional customer-facing reply.&lt;/td&gt;
&lt;td&gt;Manual ranking&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;escalate_complex_issue&lt;/td&gt;
&lt;td&gt;Summarize complex tickets into structured briefs for a human agent.&lt;/td&gt;
&lt;td&gt;Manual ranking&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fkoio5hpfuv1l6enql9fj.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fkoio5hpfuv1l6enql9fj.png" alt="image" width="800" height="1312"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;When we are creating the description, selecting the router prioritization policy, and selecting the model, we need to consider the exact task we want completed to optimize our results. Here are a few things worth noting as you configure these:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Task descriptions matter. The router uses them to match incoming requests to the right task. Be specific about what the task does, what kind of input it expects, and the format of the output.&lt;/li&gt;
&lt;li&gt;Put at least two models in every pool. A pool of one is a single point of failure. Even your "lowest cost" pool should have a fallback in case the primary is unavailable.&lt;/li&gt;
&lt;li&gt;The selection policy is enforced inside the pool, not across pools. "Lowest cost" means "the cheapest model in this pool that's currently healthy," not "the cheapest model on the platform."&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Once the router is saved, you'll get a router ID. That's what your app will call.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 3: Refactor the app to use the router
&lt;/h3&gt;

&lt;p&gt;Now the satisfying part. Replace the hardcoded MODEL constant with the router ID, and pass the task name through the request. Below is an example of what you could do to make it work, though not exactly what we did in our final release.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;ROUTER&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;your-router-id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;  &lt;span class="c1"&gt;# from the DigitalOcean control panel
&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;parse_urgency&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;urgency_text&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Extract the integer score from &lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;score: N, sentiment: X&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;. Defaults to 3 if unparseable.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="n"&gt;match&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;re&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;search&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;r&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;score:\s*(\d)&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;urgency_text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;re&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;IGNORECASE&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;int&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;match&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;group&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;match&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;call_router&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;task&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;system&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;user&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;resp&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;ROUTER&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;
            &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;system&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;system&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
            &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;user&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
        &lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="n"&gt;extra_body&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;task&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;task&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;  &lt;span class="c1"&gt;# router uses this to pick the pool
&lt;/span&gt;    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;resp&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;choices&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;strip&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;served_by&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;resp&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;  &lt;span class="c1"&gt;# the model the router actually picked
&lt;/span&gt;    &lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="nd"&gt;@app.post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/triage&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;triage&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ticket&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;Ticket&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;text&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Subject: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;ticket&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;subject&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="se"&gt;\n\n&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;ticket&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;body&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

    &lt;span class="n"&gt;category&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;call_router&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;classify_ticket&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Classify this support ticket into one of: billing, bug, how-to, account, other. Reply with one word.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;urgency&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;call_router&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;urgency_detection&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Score urgency from 1 (low) to 5 (critical) and note sentiment. Reply as &lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;score: N, sentiment: X&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;reply&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;call_router&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;draft_customer_reply&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Write a short, professional reply to this customer. Maximum 4 sentences.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="c1"&gt;# Only escalate when urgency warrants a human brief
&lt;/span&gt;    &lt;span class="n"&gt;urgency_score&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;parse_urgency&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;urgency&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
    &lt;span class="n"&gt;summary&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;urgency_score&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;summary&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;call_router&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;escalate_complex_issue&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Summarize this ticket for a human agent. Include the problem, what&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;s been tried, and recommended next steps.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;category&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;category&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;urgency&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;urgency&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;urgency_score&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;urgency_score&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;reply&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;reply&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;escalation_summary&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;summary&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;summary&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;routing&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;classify_ticket&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;category&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;served_by&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;urgency_detection&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;urgency&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;served_by&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;draft_customer_reply&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;reply&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;served_by&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;escalate_complex_issue&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;summary&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;served_by&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;summary&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's the whole change. It’s already done for you in the GitHub version, so there’s no need to manually do it yourself.&lt;/p&gt;

&lt;p&gt;With this, there are no model names anywhere in the app. The router decides which model handles each task, using the policies you configured. If you want to swap the underlying model for draft_customer_reply next month, you do it in the router, not in this file.&lt;/p&gt;

&lt;p&gt;The app triages one ticket by breaking it into smaller AI jobs instead of asking one model to do everything at once. When you call POST /triage, main.py builds the ticket text, then sends separate router calls for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;classify_ticket: decides the ticket category, like billing, bug, how-to, account, or other.&lt;/li&gt;
&lt;li&gt;urgency_detection: scores severity from 1 to 5 and detects sentiment; the code uses the score to decide whether to escalate.&lt;/li&gt;
&lt;li&gt;draft_customer_reply: writes a short customer-facing response.&lt;/li&gt;
&lt;li&gt;escalate_complex_issue: Tickets scoring 4 or 5 on urgency trigger the escalation summary; lower scores skip it entirely, which is where most of the cost savings live.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The key thing: the app always calls your DigitalOcean router ID from .env as the model, and the router decides which underlying model should handle each prompt.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 4: Run mixed tickets through the router
&lt;/h3&gt;

&lt;p&gt;With the router wired in, let's test it. The interesting behavior shows up when you feed the endpoint a mix of simple and complex examples. Here's a small batch of simple to complex examples in sample_tickets.json:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="o"&gt;[&lt;/span&gt;
  &lt;span class="o"&gt;{&lt;/span&gt;&lt;span class="s2"&gt;"subject"&lt;/span&gt;: &lt;span class="s2"&gt;"Password reset"&lt;/span&gt;, &lt;span class="s2"&gt;"body"&lt;/span&gt;: &lt;span class="s2"&gt;"How do I reset my password?"&lt;/span&gt;&lt;span class="o"&gt;}&lt;/span&gt;,
  &lt;span class="o"&gt;{&lt;/span&gt;&lt;span class="s2"&gt;"subject"&lt;/span&gt;: &lt;span class="s2"&gt;"Invoice question"&lt;/span&gt;, &lt;span class="s2"&gt;"body"&lt;/span&gt;: &lt;span class="s2"&gt;"Why was I charged twice on invoice INV-3382?"&lt;/span&gt;&lt;span class="o"&gt;}&lt;/span&gt;,
  &lt;span class="o"&gt;{&lt;/span&gt;&lt;span class="s2"&gt;"subject"&lt;/span&gt;: &lt;span class="s2"&gt;"This is ridiculous"&lt;/span&gt;, &lt;span class="s2"&gt;"body"&lt;/span&gt;: &lt;span class="s2"&gt;"Third time this week your dashboard has gone down during our standup. We're seriously evaluating alternatives."&lt;/span&gt;&lt;span class="o"&gt;}&lt;/span&gt;,
  &lt;span class="o"&gt;{&lt;/span&gt;&lt;span class="s2"&gt;"subject"&lt;/span&gt;: &lt;span class="s2"&gt;"Dashboard weird"&lt;/span&gt;, &lt;span class="s2"&gt;"body"&lt;/span&gt;: &lt;span class="s2"&gt;"the dashboard is weird since yesterday"&lt;/span&gt;&lt;span class="o"&gt;}&lt;/span&gt;,
  &lt;span class="o"&gt;{&lt;/span&gt;&lt;span class="s2"&gt;"subject"&lt;/span&gt;: &lt;span class="s2"&gt;"Production outage"&lt;/span&gt;, &lt;span class="s2"&gt;"body"&lt;/span&gt;: &lt;span class="s2"&gt;"Our team has been unable to access the dashboard since 09:14 UTC. ~200 internal users blocked. Logs attached show 502s from the API gateway, traced to..."&lt;/span&gt;&lt;span class="o"&gt;}&lt;/span&gt;,
  &lt;span class="o"&gt;{&lt;/span&gt;&lt;span class="s2"&gt;"subject"&lt;/span&gt;: &lt;span class="s2"&gt;"Feature request + complaint"&lt;/span&gt;, &lt;span class="s2"&gt;"body"&lt;/span&gt;: &lt;span class="s2"&gt;"Can you add bulk export? Also the existing export is too slow and crashes on &amp;gt;10k rows."&lt;/span&gt;&lt;span class="o"&gt;}&lt;/span&gt;,
  &lt;span class="o"&gt;{&lt;/span&gt;&lt;span class="s2"&gt;"subject"&lt;/span&gt;: &lt;span class="s2"&gt;"API auth"&lt;/span&gt;, &lt;span class="s2"&gt;"body"&lt;/span&gt;: &lt;span class="s2"&gt;"Getting 401s after rotating my key. Following the docs at /auth/rotate but the new key returns invalid."&lt;/span&gt;&lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="o"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In order to test them in sequence, we have provided &lt;code&gt;run_batch.py&lt;/code&gt; to facilitate this test. You can run it yourself with the following command:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python3 run_batch.py sample_tickets.json &lt;span class="nt"&gt;--json&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Loop through them and you'll see the routing do its job. The one-line "how do I reset my password?" hits the lowest-cost pool for classification and a small, fast model for urgency. The angry churn-risk message gets flagged high-urgency quickly, but the drafted reply comes from the higher-quality pool because that response is going to a real customer. The production outage gets routed to the higher-quality pool for the escalation summary, because that summary is what a human engineer is going to read at 09:15 UTC.&lt;/p&gt;

&lt;p&gt;Because call_router surfaces resp.model as served_by, every response now tells you exactly which model handled each task. Here's what the production outage ticket returns:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="o"&gt;{&lt;/span&gt;
  &lt;span class="s2"&gt;"category"&lt;/span&gt;: &lt;span class="s2"&gt;"bug"&lt;/span&gt;,
  &lt;span class="s2"&gt;"urgency"&lt;/span&gt;: &lt;span class="s2"&gt;"score: 5, sentiment: frustrated"&lt;/span&gt;,
  &lt;span class="s2"&gt;"urgency_score"&lt;/span&gt;: 5,
  &lt;span class="s2"&gt;"reply"&lt;/span&gt;: &lt;span class="s2"&gt;"Thank you for reporting this..."&lt;/span&gt;,
  &lt;span class="s2"&gt;"escalation_summary"&lt;/span&gt;: &lt;span class="s2"&gt;"Enterprise customer reports a production dashboard outage..."&lt;/span&gt;,
  &lt;span class="s2"&gt;"routing"&lt;/span&gt;: &lt;span class="o"&gt;{&lt;/span&gt;
    &lt;span class="s2"&gt;"classify_ticket"&lt;/span&gt;: &lt;span class="s2"&gt;"openai-gpt-5-nano"&lt;/span&gt;,
    &lt;span class="s2"&gt;"urgency_detection"&lt;/span&gt;: &lt;span class="s2"&gt;"anthropic-claude-haiku-4.5"&lt;/span&gt;,
    &lt;span class="s2"&gt;"draft_customer_reply"&lt;/span&gt;: &lt;span class="s2"&gt;"anthropic-claude-sonnet-4.6"&lt;/span&gt;,
    &lt;span class="s2"&gt;"escalate_complex_issue"&lt;/span&gt;: &lt;span class="s2"&gt;"anthropic-claude-opus-4.7"&lt;/span&gt;
  &lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;One request, four different models, zero model names in your application code. The cheap classifier handled the one-word category decision, Haiku scored urgency in a single fast pass, Sonnet drafted the customer-facing reply, and Opus produced the brief your on-call engineer reads. Run the password-reset ticket and the routing.escalate_complex_issue field comes back as null — the urgency score didn't clear the threshold, and that null is real money saved.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this actually saves you
&lt;/h2&gt;

&lt;p&gt;Let's put numbers on it. Assume an average ticket is 300 input tokens, with output tokens varying by task (40 for classification, 30 for urgency, 150 for a reply, 250 for an escalation summary). In our 7-ticket sample, 2-3 score high enough to escalate; we use 20% as a steady-state estimate.&lt;/p&gt;

&lt;p&gt;Using DigitalOcean's published serverless inference rates:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Task&lt;/th&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Per-ticket cost&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;classify_ticket&lt;/td&gt;
&lt;td&gt;GPT-5 Nano&lt;/td&gt;
&lt;td&gt;$0.000031&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;urgency_detection&lt;/td&gt;
&lt;td&gt;Claude Haiku 4.5&lt;/td&gt;
&lt;td&gt;$0.000450&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;draft_customer_reply&lt;/td&gt;
&lt;td&gt;Claude Sonnet 4.6&lt;/td&gt;
&lt;td&gt;$0.003150&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;escalate_complex_issue (fires ~20% of tickets)&lt;/td&gt;
&lt;td&gt;Claude Opus 4.7&lt;/td&gt;
&lt;td&gt;$0.007750&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;At 100,000 tickets/month, three strategies compared:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Strategy&lt;/th&gt;
&lt;th&gt;Monthly cost&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Hardcoded Llama 3.3 70B for everything&lt;/td&gt;
&lt;td&gt;$109&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Router (cost-aware)&lt;/td&gt;
&lt;td&gt;$518&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hardcoded Claude Opus 4.7 for everything&lt;/td&gt;
&lt;td&gt;$1,775&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The honest result: the router isn't the cheapest option. Hardcoded Llama 70B is. But Llama 70B writing your enterprise outage reply is the cost. You're only saving money by treating a churn-risk ticket the same as a password reset.&lt;/p&gt;

&lt;p&gt;The fair comparison is against the realistic alternative: once you decide Llama's customer-facing replies aren't good enough, the choice is Opus-for-everything or the router. The router is 71% cheaper than all-Opus while only routing the expensive Opus 4.7 model to the tickets that actually need it.&lt;/p&gt;

&lt;p&gt;Run this math on your own ticket mix before committing. The ratio of trivial-to-complex tickets is the biggest lever: a queue that's 80% password resets saves far more than one that's 80% escalations.&lt;/p&gt;

&lt;h2&gt;
  
  
  Production checklist
&lt;/h2&gt;

&lt;p&gt;Before you put this in front of real tickets:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Log task type, latency, token usage, and selected model on every call. You can't tune what you can't see, and the router's value is invisible without per-task metrics.&lt;/li&gt;
&lt;li&gt;Build a small eval set per task. Maybe 20 tickets per task with known-good outputs. Run it before changing pool composition. The whole point of the router is that you can swap models without code changes, but you still want to know whether the swap was an improvement.&lt;/li&gt;
&lt;li&gt;Keep at least one fallback in every pool. A pool of one defeats half the reason to use a router.&lt;/li&gt;
&lt;li&gt;Use direct model calls for controlled benchmarks. When you're measuring a specific model's behavior, you don't want the router making your benchmark non-deterministic.&lt;/li&gt;
&lt;li&gt;Revisit routing rules quarterly. Model pricing and quality shift. The pool that was "lowest cost" six months ago might not be today.&lt;/li&gt;
&lt;li&gt;Treat task descriptions as production config. Version them, review changes, don't edit them in the UI without a record.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Closing thoughts
&lt;/h2&gt;

&lt;p&gt;The app you ended up with isn't bigger than the one you started with: it's actually smaller, because the model selection logic moved out of the code and into the router. The router is doing the work that used to be a match statement: matching tasks to models, falling back when something's unavailable, and giving you a single place to change strategy. Serverless inference via DigitalOcean's Inference Router enables your app more flexibility and efficiency without any of the hassle of a hardcoded setup.&lt;/p&gt;

&lt;p&gt;From here, a few natural next steps: stream the &lt;code&gt;draft_customer_reply&lt;/code&gt; task back to the client so agents can start reading before generation finishes; wire the escalation summaries into your real ticketing system; or stand up a second router for an unrelated workflow and reuse the same access key.&lt;/p&gt;

&lt;p&gt;The full sample code is available in the companion repo, and the router configuration takes about five minutes in the &lt;a href="https://cloud.digitalocean.com" rel="noopener noreferrer"&gt;DigitalOcean control panel&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>tutorial</category>
      <category>api</category>
      <category>inference</category>
    </item>
    <item>
      <title>Python Decorators: From Basics to Real-World Use Cases</title>
      <dc:creator>DigitalOcean</dc:creator>
      <pubDate>Tue, 12 May 2026 21:04:07 +0000</pubDate>
      <link>https://dev.to/digitalocean/python-decorators-from-basics-to-real-world-use-cases-n5f</link>
      <guid>https://dev.to/digitalocean/python-decorators-from-basics-to-real-world-use-cases-n5f</guid>
      <description>&lt;p&gt;&lt;em&gt;This article was originally written by Shaoni Mukherjee (AI Technical Writer)&lt;/em&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Key takeaways
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Python decorators allow additional functionality to be added to functions without changing the original function code.&lt;/li&gt;
&lt;li&gt;Decorators help reduce repeated code and improve code reusability.&lt;/li&gt;
&lt;li&gt;The &lt;code&gt;@decorator_name&lt;/code&gt; syntax is a cleaner way of wrapping functions.&lt;/li&gt;
&lt;li&gt;Decorators are commonly used for logging, authentication, caching, validation, and performance monitoring.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;*args&lt;/code&gt; and &lt;code&gt;**kwargs&lt;/code&gt; make decorators flexible enough to work with different function arguments.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;functools.wraps&lt;/code&gt; helps preserve the original function metadata and should be considered a best practice.&lt;/li&gt;
&lt;li&gt;Multiple decorators can be chained together to add multiple layers of functionality.&lt;/li&gt;
&lt;li&gt;Frameworks like Flask and Django rely heavily on decorators for routing, authentication, and request handling.&lt;/li&gt;
&lt;li&gt;Decorators should be kept simple and focused to maintain readability and easier debugging.&lt;/li&gt;
&lt;li&gt;Understanding decorators is important for writing cleaner and more maintainable Python applications.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Introduction
&lt;/h2&gt;

&lt;p&gt;While building real-world &lt;a href="https://www.digitalocean.com/community/tutorials/python-tutorial" rel="noopener noreferrer"&gt;Python&lt;/a&gt; applications, a common challenge is the repetition of certain logic codes, such as logging, authentication, validation, time, or performance monitoring across multiple functions. For instance, API endpoints often require user authentication checks, and performance-critical functions may need execution time tracking.&lt;/p&gt;

&lt;p&gt;Adding the same logic code within each function often leads to cluttered code, reduced readability, and increased maintenance effort. Decorators address this problem by creating the separation of such cross-cutting concerns into reusable components that can be applied to functions in a clean and consistent manner. In frameworks like &lt;a href="https://www.digitalocean.com/community/tutorials/how-to-create-your-first-web-application-using-flask-and-python-3" rel="noopener noreferrer"&gt;Flask&lt;/a&gt;, the &lt;code&gt;@app.route("/")&lt;/code&gt; decorator links a URL to a function without requiring explicit routing logic, while in &lt;a href="https://www.digitalocean.com/solutions/django-hosting" rel="noopener noreferrer"&gt;Django&lt;/a&gt;, decorators such as &lt;code&gt;@login_required&lt;/code&gt; enforce access control by restricting views to authenticated users. This approach promotes modularity, improves code clarity, and simplifies the overall structure of applications.&lt;/p&gt;

&lt;h3&gt;
  
  
  What are Python decorators?
&lt;/h3&gt;

&lt;p&gt;Decorators are basically a wrapper around a function to modify it for better use. The function remains the same, but the decorator adds an extra something to the function.&lt;/p&gt;

&lt;h4&gt;
  
  
  The core idea
&lt;/h4&gt;

&lt;p&gt;Say you have a simple function:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;greet&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Hello, world!&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now imagine you want to print a line before and after every function you write, without modifying each one. A decorator lets you do exactly that:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;my_decorator&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;func&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;wrapper&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
        &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;--- Before ---&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="nf"&gt;func&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;           &lt;span class="c1"&gt;# calls the original function
&lt;/span&gt;        &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;--- After ---&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;wrapper&lt;/span&gt;

&lt;span class="nd"&gt;@my_decorator&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;greet&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Hello, world!&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="nf"&gt;greet&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Output:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="o"&gt;---&lt;/span&gt; &lt;span class="n"&gt;Before&lt;/span&gt; &lt;span class="o"&gt;---&lt;/span&gt;
&lt;span class="n"&gt;Hello&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;world&lt;/span&gt;&lt;span class="err"&gt;!&lt;/span&gt;
&lt;span class="o"&gt;---&lt;/span&gt; &lt;span class="n"&gt;After&lt;/span&gt; &lt;span class="o"&gt;---&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;@my_decorator&lt;/code&gt; line is just shorthand for &lt;code&gt;greet = my_decorator(greet)&lt;/code&gt;. Python replaces your function with the wrapped version automatically.&lt;br&gt;
To understand the concept better, let us take a real-world example of timing a function:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;time&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;timer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;func&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;wrapper&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;        &lt;span class="c1"&gt;# *args lets it work with ANY function
&lt;/span&gt;        &lt;span class="n"&gt;start&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;time&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
        &lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;func&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;end&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;time&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
        &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;func&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;__name__&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; took &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;end&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;start&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; seconds&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;wrapper&lt;/span&gt;

&lt;span class="nd"&gt;@timer&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;slow_task&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sleep&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Task done!&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="nf"&gt;slow_task&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Output:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;Task&lt;/span&gt; &lt;span class="n"&gt;done&lt;/span&gt;&lt;span class="err"&gt;!&lt;/span&gt;
&lt;span class="n"&gt;slow_task&lt;/span&gt; &lt;span class="n"&gt;took&lt;/span&gt; &lt;span class="mf"&gt;1.0012&lt;/span&gt; &lt;span class="n"&gt;seconds&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Why decorators matter (especially in real projects)
&lt;/h4&gt;

&lt;p&gt;They're everywhere in Python. Common use cases include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;@staticmethod&lt;/code&gt; / &lt;code&gt;@classmethod&lt;/code&gt;&lt;/strong&gt; — built into Python for class methods&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;@app.route('/home')&lt;/code&gt;&lt;/strong&gt; — Flask/Django use them to define web routes&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;@login_required&lt;/code&gt;&lt;/strong&gt; — Django uses this to protect pages behind authentication&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Logging, caching, retrying failed requests&lt;/strong&gt; — all cleanly handled with decorators&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A decorator takes a function, adds behavior around it, and returns a new function without touching the original code.&lt;/p&gt;

&lt;h2&gt;
  
  
  How decorators work internally
&lt;/h2&gt;

&lt;p&gt;To understand decorators better, we will first need to understand a few core Python concepts:&lt;/p&gt;

&lt;h3&gt;
  
  
  Foundation: Functions are objects in Python
&lt;/h3&gt;

&lt;p&gt;In Python, functions aren't special, but they're just objects like integers or strings.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;say_hello&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Hello!&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# Pass a function as an argument
&lt;/span&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;run_it&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;func&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="nf"&gt;func&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="nf"&gt;run_it&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;say_hello&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;   &lt;span class="c1"&gt;# prints: Hello!
&lt;/span&gt;
&lt;span class="c1"&gt;# Assign a function to a variable
&lt;/span&gt;&lt;span class="n"&gt;my_func&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;say_hello&lt;/span&gt;
&lt;span class="nf"&gt;my_func&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;           &lt;span class="c1"&gt;# prints: Hello!
&lt;/span&gt;
&lt;span class="c1"&gt;# Return a function from another function
&lt;/span&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;get_greeter&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;say_hi&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
        &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Hi!&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;say_hi&lt;/span&gt;   &lt;span class="c1"&gt;# returning the function, not calling it
&lt;/span&gt;
&lt;span class="n"&gt;greeter&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;get_greeter&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="nf"&gt;greeter&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;           &lt;span class="c1"&gt;# prints: Hi!
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is the entire foundation that decorators are built on.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why are decorators needed?
&lt;/h3&gt;

&lt;p&gt;Imagine there are many functions in a project, and each function needs logging.&lt;/p&gt;

&lt;h4&gt;
  
  
  Without decorators:
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;add&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Function started&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Function ended&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;multiply&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Function started&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Function ended&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Problem:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Repeated code&lt;/li&gt;
&lt;li&gt;Hard to maintain in large projects&lt;/li&gt;
&lt;li&gt;If logging changes, every function must be updated&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Decorators solve this problem by reusing common functionality.&lt;/p&gt;

&lt;h4&gt;
  
  
  With decorators:
&lt;/h4&gt;

&lt;p&gt;Using decorators, the repeated code (&lt;code&gt;"Function started"&lt;/code&gt; and &lt;code&gt;"Function ended"&lt;/code&gt;) can be moved into a single reusable decorator.&lt;br&gt;
Instead of writing the same lines inside every function, the decorator handles it automatically.&lt;/p&gt;
&lt;h4&gt;
  
  
  Step 1: Create the decorator
&lt;/h4&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;log_function&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;func&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;wrapper&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Function started&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

        &lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;func&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

        &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Function ended&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;wrapper&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;h4&gt;
  
  
  &lt;strong&gt;Step 2: Apply the Decorator&lt;/strong&gt;
&lt;/h4&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nd"&gt;@log_function&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;add&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt;


&lt;span class="nd"&gt;@log_function&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;multiply&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;h4&gt;
  
  
  &lt;strong&gt;Calling the Functions&lt;/strong&gt;
&lt;/h4&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;add&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;multiply&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Output:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;Function&lt;/span&gt; &lt;span class="n"&gt;started&lt;/span&gt;
&lt;span class="n"&gt;Function&lt;/span&gt; &lt;span class="n"&gt;ended&lt;/span&gt;
&lt;span class="mi"&gt;5&lt;/span&gt;

&lt;span class="n"&gt;Function&lt;/span&gt; &lt;span class="n"&gt;started&lt;/span&gt;
&lt;span class="n"&gt;Function&lt;/span&gt; &lt;span class="n"&gt;ended&lt;/span&gt;
&lt;span class="mi"&gt;20&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  What changed?
&lt;/h4&gt;

&lt;p&gt;The functions now only contain their main logic:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The extra behavior (logging) is handled by the decorator separately.&lt;/p&gt;

&lt;h3&gt;
  
  
  Visual understanding
&lt;/h3&gt;

&lt;p&gt;When this runs:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nf"&gt;add&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Python internally does this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;add&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;log_function&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;add&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;So the actual flow becomes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nf"&gt;wrapper&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="err"&gt;├──&lt;/span&gt; &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Function started&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="err"&gt;├──&lt;/span&gt; &lt;span class="n"&gt;call&lt;/span&gt; &lt;span class="n"&gt;original&lt;/span&gt; &lt;span class="nf"&gt;add&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="err"&gt;├──&lt;/span&gt; &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Function ended&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="err"&gt;└──&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Better Version Using *args and **kwargs
&lt;/h2&gt;

&lt;p&gt;The previous decorator only works for functions with two arguments.&lt;br&gt;
A more reusable decorator looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;log_function&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;func&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;wrapper&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Function started&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

        &lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;func&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

        &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Function ended&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;wrapper&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now it works with:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;any number of arguments&lt;/li&gt;
&lt;li&gt;positional arguments&lt;/li&gt;
&lt;li&gt;keyword arguments&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Why this is powerful
&lt;/h2&gt;

&lt;p&gt;Imagine 100 functions needing logging.&lt;/p&gt;

&lt;p&gt;Without decorators:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;repeated code everywhere&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;With decorators:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;write logging once&lt;/li&gt;
&lt;li&gt;reuse everywhere&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is one of the biggest reasons decorators are widely used in real-world Python projects and frameworks like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://flask.palletsprojects.com/en/stable/" rel="noopener noreferrer"&gt;Flask&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.djangoproject.com/" rel="noopener noreferrer"&gt;Django&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://fastapi.tiangolo.com/" rel="noopener noreferrer"&gt;FastAPI&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.digitalocean.com/community/tutorials/pytorch-101-advanced" rel="noopener noreferrer"&gt;PyTorch&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.digitalocean.com/community/tutorials/introduction-to-tensorflow-build-ai-across-domains" rel="noopener noreferrer"&gt;TensorFlow&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Common practical examples of Python decorators
&lt;/h2&gt;

&lt;p&gt;A few of the most common practical examples are listed here, from solo projects to production systems.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Timing and performance measurement
&lt;/h3&gt;

&lt;p&gt;Useful when profiling slow functions or benchmarking code.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;time&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;functools&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;wraps&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;timer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;func&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="nd"&gt;@wraps&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;func&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;wrapper&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;start&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;perf_counter&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
        &lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;func&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;end&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;perf_counter&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
        &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;func&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;__name__&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; ran in &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;end&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;start&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;s&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;wrapper&lt;/span&gt;

&lt;span class="nd"&gt;@timer&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;process_data&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;n&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;total&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;n&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;total&lt;/span&gt;

&lt;span class="nf"&gt;process_data&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1_000_000&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="c1"&gt;# process_data ran in 0.0312s
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;perf_counter()&lt;/code&gt; is preferred over &lt;code&gt;time.time()&lt;/code&gt; for short measurements, and it's higher resolution and is not affected by system clock adjustments.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Logging
&lt;/h3&gt;

&lt;p&gt;Instead of adding print statements everywhere, a logging decorator handles it in one place.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;logging&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;functools&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;wraps&lt;/span&gt;

&lt;span class="n"&gt;logging&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;basicConfig&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;level&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;logging&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;INFO&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;log_calls&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;func&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="nd"&gt;@wraps&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;func&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;wrapper&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;logging&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;info&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Calling &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;func&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;__name__&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; | args=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; kwargs=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;func&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;logging&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;info&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;func&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;__name__&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; returned &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;wrapper&lt;/span&gt;

&lt;span class="nd"&gt;@log_calls&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;multiply&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt;

&lt;span class="nf"&gt;multiply&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="c1"&gt;# INFO: Calling multiply | args=(4, 5) kwargs={}
# INFO: multiply returned 20
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In production, you'd swap &lt;code&gt;logging.info&lt;/code&gt; for a structured logger like &lt;code&gt;structlog&lt;/code&gt; or a cloud logging sink.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Retry on failure
&lt;/h3&gt;

&lt;p&gt;Critical for network calls, API requests, or anything that can fail transiently.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;time&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;functools&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;wraps&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;retry&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;times&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;delay&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;decorator&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;func&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="nd"&gt;@wraps&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;func&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;wrapper&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
            &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;attempt&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;times&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
                &lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;func&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
                &lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="nb"&gt;Exception&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Attempt &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;attempt&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; failed: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
                    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;attempt&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="n"&gt;times&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                        &lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sleep&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;delay&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;Exception&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;func&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;__name__&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; failed after &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;times&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; attempts&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;wrapper&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;decorator&lt;/span&gt;

&lt;span class="nd"&gt;@retry&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;times&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;delay&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;fetch_data&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;url&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;requests&lt;/span&gt;
    &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;requests&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;url&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;timeout&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;raise_for_status&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="nf"&gt;fetch_data&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://api.example.com/data&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="c1"&gt;# Attempt 1 failed: Connection timeout
# Attempt 2 failed: Connection timeout
# Attempt 3 failed: Connection timeout
# Exception: fetch_data failed after 3 attempts
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Notice this is a &lt;strong&gt;decorator factory&lt;/strong&gt; — &lt;code&gt;retry(times=3)&lt;/code&gt; returns the actual decorator. This is how you pass arguments to decorators.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Caching memoization
&lt;/h3&gt;

&lt;p&gt;Avoids recomputing expensive results by storing previous outputs.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;functools&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;wraps&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;memoize&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;func&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;cache&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{}&lt;/span&gt;
    &lt;span class="nd"&gt;@wraps&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;func&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;wrapper&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;args&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;cache&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;cache&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;func&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Cache miss — computing for &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;else&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Cache hit for &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;cache&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;wrapper&lt;/span&gt;

&lt;span class="nd"&gt;@memoize&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;fibonacci&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;n&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;n&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;n&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;fibonacci&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;n&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="nf"&gt;fibonacci&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;n&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="nf"&gt;fibonacci&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;6&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="c1"&gt;# Cache miss — computing for (6,)
# Cache miss — computing for (5,)
# ...
&lt;/span&gt;&lt;span class="nf"&gt;fibonacci&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;6&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="c1"&gt;# Cache hit for (6,)   ← instantly returns stored result
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Python actually ships a production-grade version of this built in:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;functools&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;lru_cache&lt;/span&gt;

&lt;span class="nd"&gt;@lru_cache&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;maxsize&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;128&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;fibonacci&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;n&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;n&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;n&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;fibonacci&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;n&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="nf"&gt;fibonacci&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;n&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;lru_cache&lt;/code&gt; (Least Recently Used) is thread-safe and evicts old entries when the cache is full — use it over a hand-rolled version in real projects.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Access control authorization
&lt;/h3&gt;

&lt;p&gt;A staple in web frameworks like Flask and Django.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;functools&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;wraps&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;require_role&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;role&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;decorator&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;func&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="nd"&gt;@wraps&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;func&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;wrapper&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;user&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;user&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="n"&gt;role&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;PermissionError&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Access denied. Required role: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;role&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;func&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;user&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;wrapper&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;decorator&lt;/span&gt;

&lt;span class="nd"&gt;@require_role&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;admin&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;delete_user&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;user&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;user_id&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Deleting user &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;user_id&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;admin&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Shaoni&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;admin&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="n"&gt;guest&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Guest&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;viewer&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="nf"&gt;delete_user&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;admin&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;42&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;    &lt;span class="c1"&gt;# Deleting user 42
&lt;/span&gt;&lt;span class="nf"&gt;delete_user&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;guest&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;42&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;    &lt;span class="c1"&gt;# PermissionError: Access denied. Required role: admin
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Django's &lt;code&gt;@login_required&lt;/code&gt; and &lt;code&gt;@permission_required&lt;/code&gt; follow this exact pattern internally.&lt;/p&gt;

&lt;h3&gt;
  
  
  6. Input validation
&lt;/h3&gt;

&lt;p&gt;Validate arguments before they even reach your function's logic.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;functools&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;wraps&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;validate_positive&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;arg_positions&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;decorator&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;func&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="nd"&gt;@wraps&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;func&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;wrapper&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
            &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;arg_positions&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                    &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;ValueError&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
                        &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Argument at position &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; must be positive, got &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
                    &lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;func&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;wrapper&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;decorator&lt;/span&gt;

&lt;span class="nd"&gt;@validate_positive&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;calculate_area&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;width&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;height&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;width&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;height&lt;/span&gt;

&lt;span class="nf"&gt;calculate_area&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;    &lt;span class="c1"&gt;# 50
&lt;/span&gt;&lt;span class="nf"&gt;calculate_area&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;   &lt;span class="c1"&gt;# ValueError: Argument at position 0 must be positive
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  &lt;strong&gt;7. Rate Limiting&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;Preventing a function from being called too frequently is very common in API clients.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;time&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;functools&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;wraps&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;rate_limit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;calls_per_second&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;min_interval&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mf"&gt;1.0&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="n"&gt;calls_per_second&lt;/span&gt;
    &lt;span class="n"&gt;last_called&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mf"&gt;0.0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;   &lt;span class="c1"&gt;# mutable container to hold state in closure
&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;decorator&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;func&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="nd"&gt;@wraps&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;func&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;wrapper&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
            &lt;span class="n"&gt;elapsed&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;time&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;last_called&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
            &lt;span class="n"&gt;wait&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;min_interval&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;elapsed&lt;/span&gt;
            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;wait&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Rate limit: waiting &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;wait&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;s&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
                &lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sleep&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;wait&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="n"&gt;last_called&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;time&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;func&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;wrapper&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;decorator&lt;/span&gt;

&lt;span class="nd"&gt;@rate_limit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;calls_per_second&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;call_api&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;endpoint&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Calling &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;endpoint&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="nf"&gt;call_api&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/users&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;call_api&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/posts&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;    &lt;span class="c1"&gt;# Rate limit: waiting 0.49s
&lt;/span&gt;&lt;span class="nf"&gt;call_api&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/comments&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="c1"&gt;# Rate limit: waiting 0.49s
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Quick reference
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Decorator&lt;/th&gt;
&lt;th&gt;Use Case&lt;/th&gt;
&lt;th&gt;Real-world Equivalent&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;@timer&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Measure execution time&lt;/td&gt;
&lt;td&gt;Profiling, benchmarking&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;@log_calls&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Audit function calls&lt;/td&gt;
&lt;td&gt;Observability, debugging&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;@retry&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Handle transient failures&lt;/td&gt;
&lt;td&gt;API clients, DB connections&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;@lru_cache&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Cache expensive results&lt;/td&gt;
&lt;td&gt;ML inference, DB queries&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;@require_role&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Guard endpoints by role&lt;/td&gt;
&lt;td&gt;Django, Flask auth&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;@validate_positive&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Sanitize inputs early&lt;/td&gt;
&lt;td&gt;Data pipelines, APIs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;@rate_limit&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Throttle call frequency&lt;/td&gt;
&lt;td&gt;External API clients&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Real-world use cases in frameworks
&lt;/h2&gt;

&lt;p&gt;Decorators are heavily used in modern Python frameworks because they provide a clean and reusable way to add functionality to applications without modifying the core business logic.&lt;br&gt;
Frameworks such as Flask and Django use decorators for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Routing&lt;/li&gt;
&lt;li&gt;Authentication&lt;/li&gt;
&lt;li&gt;Authorization&lt;/li&gt;
&lt;li&gt;Caching&lt;/li&gt;
&lt;li&gt;Request validation&lt;/li&gt;
&lt;li&gt;Restricting HTTP methods&lt;/li&gt;
&lt;li&gt;Logging&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These decorators make applications cleaner, easier to maintain, and more readable.&lt;/p&gt;
&lt;h3&gt;
  
  
  Flask routing decorator
&lt;/h3&gt;

&lt;p&gt;One of the most common examples of decorators appears in Flask routing.&lt;br&gt;
Using Flask:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;flask&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Flask&lt;/span&gt;

&lt;span class="n"&gt;app&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Flask&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;__name__&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="nd"&gt;@app.route&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;home&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
   &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Homepage&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Here:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nd"&gt;@app.route&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;is a decorator.&lt;br&gt;
It tells Flask:&lt;br&gt;
“When a user visits &lt;code&gt;/&lt;/code&gt;, execute the &lt;code&gt;home()&lt;/code&gt; function.”&lt;/p&gt;
&lt;h3&gt;
  
  
  Flask authentication decorator
&lt;/h3&gt;

&lt;p&gt;Decorators are also commonly used for authentication.&lt;br&gt;
Example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nd"&gt;@app.route&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/dashboard&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nd"&gt;@login_required&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;dashboard&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
   &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Dashboard&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Here:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nd"&gt;@login_required&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;checks whether the user is logged in before allowing access to the dashboard.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why this is useful
&lt;/h3&gt;

&lt;p&gt;Without decorators, authentication checks would need to be repeated inside every protected function.&lt;br&gt;
Example without decorator:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;dashboard&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
   &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;logged_in&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
       &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Please log in&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
   &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Dashboard&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Using decorators:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;avoids repeated code&lt;/li&gt;
&lt;li&gt;keeps route definitions clean&lt;/li&gt;
&lt;li&gt;centralizes authentication logic&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This becomes extremely useful in large applications with many protected routes.&lt;/p&gt;

&lt;h3&gt;
  
  
  Django authentication decorator
&lt;/h3&gt;

&lt;p&gt;Django also uses decorators extensively.&lt;br&gt;
Example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;django.contrib.auth.decorators&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;login_required&lt;/span&gt;
&lt;span class="nd"&gt;@login_required&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;dashboard&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
   &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nc"&gt;HttpResponse&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Welcome&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;@login_required&lt;/code&gt; decorator ensures:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;only authenticated users can access the view&lt;/li&gt;
&lt;li&gt;unauthorized users are redirected to the login page&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Benefits
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Reusable security checks&lt;/li&gt;
&lt;li&gt;Cleaner view functions&lt;/li&gt;
&lt;li&gt;Better maintainability&lt;/li&gt;
&lt;li&gt;Centralized authentication handling&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Django HTTP method restriction
&lt;/h3&gt;

&lt;p&gt;Django provides decorators to restrict HTTP request methods.&lt;/p&gt;

&lt;p&gt;Example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;django.views.decorators.http&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;require_POST&lt;/span&gt;
&lt;span class="nd"&gt;@require_POST&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;submit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
   &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nc"&gt;HttpResponse&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Submitted&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The decorator:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nd"&gt;@require_POST&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;ensures the function only accepts POST requests.&lt;br&gt;
If a GET request is sent, Django automatically returns an error.&lt;/p&gt;
&lt;h3&gt;
  
  
  Why this matters
&lt;/h3&gt;

&lt;p&gt;This helps:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;enforce API rules&lt;/li&gt;
&lt;li&gt;improve security&lt;/li&gt;
&lt;li&gt;prevent invalid request types&lt;/li&gt;
&lt;li&gt;simplify validation logic&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Without decorators, manual checks would be needed inside every function.&lt;/p&gt;
&lt;h3&gt;
  
  
  Django caching decorator
&lt;/h3&gt;

&lt;p&gt;Decorators are also used for performance optimization.&lt;/p&gt;

&lt;p&gt;Example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;django.views.decorators.cache&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;cache_page&lt;/span&gt;
&lt;span class="nd"&gt;@cache_page&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;60&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;my_view&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
   &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nc"&gt;HttpResponse&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Cached&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Here:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nd"&gt;@cache_page&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;60&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;stores the response for 60 seconds.&lt;br&gt;
If another user requests the same page during that time:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Django serves the cached version&lt;/li&gt;
&lt;li&gt;the function does not run again&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;
  
  
  Advanced decorator concepts
&lt;/h2&gt;

&lt;p&gt;Once the basic concepts are understood, the next step is to learn how decorators are implemented in production-grade Python applications. Advanced decorator patterns solve practical problems such as preserving function metadata, creating configurable decorators, and combining multiple decorators together.&lt;/p&gt;

&lt;p&gt;These concepts are widely used in frameworks, libraries, and enterprise-level Python applications.&lt;/p&gt;
&lt;h3&gt;
  
  
  Preserving function metadata with &lt;code&gt;functools.wraps&lt;/code&gt;
&lt;/h3&gt;

&lt;p&gt;One common issue with decorators is that they replace the original function with the wrapper function. As a result, important metadata such as the function name, documentation string, annotations, and debugging information may be lost.&lt;/p&gt;

&lt;p&gt;Consider the following decorator:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;decorator&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;func&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;

   &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;wrapper&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
       &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;func&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

   &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;wrapper&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Using it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nd"&gt;@decorator&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;greet&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
   &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;This function greets the user&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
   &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Hello&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now checking the function name:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;greet&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;__name__&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Output:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;wrapper&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Instead of returning &lt;code&gt;"greet"&lt;/code&gt;, Python returns &lt;code&gt;"wrapper"&lt;/code&gt; because the original metadata has been overridden by the wrapper function.&lt;br&gt;
This creates problems for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;debugging&lt;/li&gt;
&lt;li&gt;logging&lt;/li&gt;
&lt;li&gt;API documentation&lt;/li&gt;
&lt;li&gt;introspection&lt;/li&gt;
&lt;li&gt;testing frameworks&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;To solve this problem, Python provides &lt;code&gt;functools.wraps&lt;/code&gt;.&lt;/p&gt;
&lt;h2&gt;
  
  
  Using &lt;code&gt;functools.wraps&lt;/code&gt;
&lt;/h2&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;functools&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;wraps&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;decorator&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;func&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;

   &lt;span class="nd"&gt;@wraps&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;func&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
   &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;wrapper&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
       &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;func&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

   &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;wrapper&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Using it again:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nd"&gt;@decorator&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;greet&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
   &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;This function greets the user&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
   &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Hello&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;greet&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;__name__&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Output:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;greet&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;@wraps(func)&lt;/code&gt; decorator copies the original function metadata into the wrapper function. This is considered a best practice when writing decorators in production applications.&lt;/p&gt;

&lt;h3&gt;
  
  
  Decorators with arguments
&lt;/h3&gt;

&lt;p&gt;In many real-world scenarios, decorators need configuration values. This requires creating decorators that accept arguments.&lt;br&gt;
A decorator with arguments introduces an additional level of nesting.&lt;br&gt;
Example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;repeat&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;n&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;

   &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;decorator&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;func&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;

       &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;wrapper&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;

           &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;n&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
               &lt;span class="nf"&gt;func&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

       &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;wrapper&lt;/span&gt;

   &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;decorator&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Using it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nd"&gt;@repeat&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;greet&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
   &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Hello&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Calling:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nf"&gt;greet&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Output:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;Hello&lt;/span&gt;
&lt;span class="n"&gt;Hello&lt;/span&gt;
&lt;span class="n"&gt;Hello&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Understanding the structure
&lt;/h2&gt;

&lt;p&gt;This example contains three functions:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nf"&gt;repeat&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;        &lt;span class="err"&gt;→&lt;/span&gt; &lt;span class="n"&gt;accepts&lt;/span&gt; &lt;span class="n"&gt;decorator&lt;/span&gt; &lt;span class="n"&gt;arguments&lt;/span&gt;
&lt;span class="nf"&gt;decorator&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;     &lt;span class="err"&gt;→&lt;/span&gt; &lt;span class="n"&gt;accepts&lt;/span&gt; &lt;span class="n"&gt;the&lt;/span&gt; &lt;span class="n"&gt;original&lt;/span&gt; &lt;span class="n"&gt;function&lt;/span&gt;
&lt;span class="nf"&gt;wrapper&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;       &lt;span class="err"&gt;→&lt;/span&gt; &lt;span class="n"&gt;executes&lt;/span&gt; &lt;span class="n"&gt;additional&lt;/span&gt; &lt;span class="n"&gt;logic&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The execution flow becomes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;greet&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;repeat&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;)(&lt;/span&gt;&lt;span class="n"&gt;greet&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This pattern is heavily used in:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;retry mechanisms&lt;/li&gt;
&lt;li&gt;caching systems&lt;/li&gt;
&lt;li&gt;rate limiting&lt;/li&gt;
&lt;li&gt;authorization frameworks&lt;/li&gt;
&lt;li&gt;logging systems&lt;/li&gt;
&lt;li&gt;timeout handling&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For example, a retry decorator may accept the number of retries:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nd"&gt;@retry&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A caching decorator may accept an expiration time:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nd"&gt;@cache&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;expire&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;60&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Decorator arguments make decorators significantly more flexible and reusable.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;Chaining Multiple Decorators&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;Python allows multiple decorators to be applied to the same function.&lt;/p&gt;

&lt;p&gt;Example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nd"&gt;@decorator_one&lt;/span&gt;
&lt;span class="nd"&gt;@decorator_two&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;func&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
   &lt;span class="k"&gt;pass&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is internally interpreted as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;func&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;decorator_one&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;decorator_two&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;func&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The execution order is important.&lt;/p&gt;

&lt;p&gt;Python applies decorators from bottom to top:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;code&gt;decorator_two&lt;/code&gt; wraps the function first&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;decorator_one&lt;/code&gt; wraps the result next&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  Example of chained decorators
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;decorator_one&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;func&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;

   &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;wrapper&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
       &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Decorator One - Before&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

       &lt;span class="nf"&gt;func&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

       &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Decorator One - After&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

   &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;wrapper&lt;/span&gt;


&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;decorator_two&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;func&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;

   &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;wrapper&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
       &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Decorator Two - Before&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

       &lt;span class="nf"&gt;func&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

       &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Decorator Two - After&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

   &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;wrapper&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Applying both decorators:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nd"&gt;@decorator_one&lt;/span&gt;
&lt;span class="nd"&gt;@decorator_two&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;greet&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
   &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Hello&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Calling:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nf"&gt;greet&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Output:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;Decorator&lt;/span&gt; &lt;span class="n"&gt;One&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;Before&lt;/span&gt;
&lt;span class="n"&gt;Decorator&lt;/span&gt; &lt;span class="n"&gt;Two&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;Before&lt;/span&gt;
&lt;span class="n"&gt;Hello&lt;/span&gt;
&lt;span class="n"&gt;Decorator&lt;/span&gt; &lt;span class="n"&gt;Two&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;After&lt;/span&gt;
&lt;span class="n"&gt;Decorator&lt;/span&gt; &lt;span class="n"&gt;One&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;After&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Understanding the execution flow
&lt;/h3&gt;

&lt;p&gt;The function call stack becomes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nf"&gt;decorator_one&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
   &lt;span class="nf"&gt;decorator_two&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
       &lt;span class="n"&gt;greet&lt;/span&gt;
   &lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This creates nested execution layers where each decorator adds behavior before and after the wrapped function. Decorator chaining is extensively used in frameworks. For example, a web route may simultaneously use:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;authentication&lt;/li&gt;
&lt;li&gt;caching&lt;/li&gt;
&lt;li&gt;rate limiting&lt;/li&gt;
&lt;li&gt;logging&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nd"&gt;@app.route&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/dashboard&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nd"&gt;@login_required&lt;/span&gt;
&lt;span class="nd"&gt;@cache_page&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;60&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;dashboard&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
   &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Dashboard&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each decorator contributes a separate layer of functionality while keeping the core business logic clean and isolated.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;Python decorators provide a clean and powerful way to add extra functionality to functions without modifying the original code. They help reduce code duplication, improve reusability, and make applications easier to maintain.&lt;/p&gt;

&lt;p&gt;From simple logging examples to advanced use cases in frameworks like Flask and Django, decorators play an important role in modern Python development. Understanding how decorators work helps in writing cleaner, more scalable, and more professional Python code.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>python</category>
      <category>tutorial</category>
      <category>learning</category>
    </item>
    <item>
      <title>NVIDIA B300 Blackwell Ultra: A Technical Deep Dive</title>
      <dc:creator>DigitalOcean</dc:creator>
      <pubDate>Thu, 07 May 2026 23:53:39 +0000</pubDate>
      <link>https://dev.to/digitalocean/nvidia-b300-blackwell-ultra-a-technical-deep-dive-5c6i</link>
      <guid>https://dev.to/digitalocean/nvidia-b300-blackwell-ultra-a-technical-deep-dive-5c6i</guid>
      <description>&lt;p&gt;The NVIDIA B300 (Blackwell Ultra) is NVIDIA's latest data center GPU, built for AI training and inference. In this deep dive, we break down the full architecture, from its dual-die design and 5th-generation tensor cores to NVFP4 precision and NVLink 5 scaling.        &lt;/p&gt;

&lt;p&gt;What we cover:&lt;br&gt;
  &lt;a href="https://www.youtube.com/watch?v=Kf_3n_pxa0I" rel="noopener noreferrer"&gt;00:00&lt;/a&gt; - Introduction&lt;br&gt;
  &lt;a href="https://www.youtube.com/watch?v=Kf_3n_pxa0I&amp;amp;t=56s" rel="noopener noreferrer"&gt;00:56&lt;/a&gt; - Why the B300 exists&lt;br&gt;
  &lt;a href="https://www.youtube.com/watch?v=Kf_3n_pxa0I&amp;amp;t=145s" rel="noopener noreferrer"&gt;02:25&lt;/a&gt; - B300 vs B200 vs H100 — the numbers&lt;br&gt;
  &lt;a href="https://www.youtube.com/watch?v=Kf_3n_pxa0I&amp;amp;t=239s" rel="noopener noreferrer"&gt;03:59&lt;/a&gt; - Dual-reticle design &amp;amp; NV-HBI interconnect&lt;br&gt;
  &lt;a href="https://www.youtube.com/watch?v=Kf_3n_pxa0I&amp;amp;t=304s" rel="noopener noreferrer"&gt;05:04&lt;/a&gt; - 5th-gen tensor cores &amp;amp; NVFP4&lt;br&gt;
  &lt;a href="https://www.youtube.com/watch?v=Kf_3n_pxa0I&amp;amp;t=476s" rel="noopener noreferrer"&gt;07:56&lt;/a&gt; - 288GB HBM3e memory breakdown&lt;br&gt;
  &lt;a href="https://www.youtube.com/watch?v=Kf_3n_pxa0I&amp;amp;t=544s" rel="noopener noreferrer"&gt;09:04&lt;/a&gt; - Multi-GPU &amp;amp; NVLink 5 architecture&lt;br&gt;
  &lt;a href="https://youtu.be/watch?v=Kf_3n_pxa0I&amp;amp;t=578s" rel="noopener noreferrer"&gt;10:38&lt;/a&gt; - Performance &amp;amp; efficiency summary&lt;/p&gt;

</description>
      <category>nvidia</category>
      <category>ai</category>
      <category>gpu</category>
      <category>hardware</category>
    </item>
  </channel>
</rss>
