DEV Community

Cover image for How to route LLM requests by cost vs. latency
DigitalOcean for DigitalOcean

Posted on

How to route LLM requests by cost vs. latency

Routing LLM requests by cost and latency means sending each request to the cheapest or fastest model that still meets the quality bar, rather than hardcoding a single model for everything. Production traffic isn't uniform: routine lookups, complex troubleshooting, and background jobs have different cost and latency requirements, even within a single app.

Using an inference router like DigitalOcean Inference Router involves defining request categories, assigning each a selection policy (cost, latency, or a benchmarked best fit), and setting fallback models so an unavailable model doesn't break the request. Cost is the price per token (which can vary by up to 100x across a model catalog); latency is mainly the time to first token (TTFT). The two don't move together, so route per task rather than using a single blended score.

Key takeaways:

  • Routing is a per-workload policy, not a search for "the best model."
  • Cost and latency usually need separate, explicit priorities.
  • Fallback and cache-aware routing protect the savings that routing is meant to deliver.
  • The DigitalOcean Inference Router implements this as configuration, not custom infrastructure.

LLM routing strategies

  • Static rules: Hardcode model per request type. Simple, but brittle.
  • Policy-based routing: A model pool per task with a cost/latency/optimal policy. Most production systems land here.
  • Dynamic routing: A classifier infers the task and policy automatically.

Routing LLM requests: A quick decision framework

Workload Priority Policy
Real-time chat Latency Speed-optimized (TTFT)
Routine lookups Cost Cost-optimized
Complex reasoning Quality Benchmarked "optimal"
Batch/background Cost Cost-optimized

Fallback models keep requests completing when the preferred model is down or rate-limited.

Cache-aware routing matters too: switching models to save a fraction of a cent can break a cached prompt and cost more overall.

How the DigitalOcean Inference Router routes LLM requests

Inference Router tasks pair a model pool with a selection policy: preset tasks default to a benchmarked Optimal policy; custom tasks choose Cost Efficiency, Speed Optimization (TTFT), or Manual Ranking.

Fallback models handle unmatched traffic. X-Model-Affinity preserves cache reuse, and X-Routing-Max-Switch-Spend-Pct (default 20%) caps the cost of switching mid-session.

It's a drop-in change. Set "model": "router:your-router-name". Routing decisions typically resolve in about 200ms, billed at the serving model's standard rate with no separate router charge during public preview.

LLM routing FAQ

What is LLM routing?

LLM routing is the practice of directing each request to the model best suited to it—by cost, latency, or task fit—instead of sending every request through one model regardless of what it needs. The DigitalOcean Inference Router implements this as a managed feature, so teams get task-aware routing and fallback handling without building it themselves.

How do you model cost-per-token for LLM inference?

Multiply input tokens by the model's input rate and output tokens by its output rate, then sum the two, since the rates are usually very different. Because rates vary widely across a model catalog, this is best tracked per task rather than as one blended number. The DigitalOcean Inference Router reports token usage and cost-relevant metrics per model and per task in its Analyze view.

What metrics matter for LLM inference observability?

Time to first token (TTFT), time per output token (TPOT), and inter-token latency (ITL) are the core metrics for judging responsiveness. A latency-optimized routing policy should be measured against these, not just overall request time. The DigitalOcean Inference Router uses TTFT specifically as the basis for its Speed Optimization selection policy.

What causes cold start latency in GPU inference, and how do you avoid it?

Cold starts happen when a model has to load onto available GPU capacity before it can serve a request, rather than running on an already-warm instance. Pooled serverless capacity and routing policies that keep related traffic on one model reduce how often this happens. The DigitalOcean Inference Router uses model affinity to keep a session's requests on the same warm model, rather than triggering repeated cold starts across models.

Can I route by both cost and latency at the same time?

Yes—typically by defining separate tasks or policies for different request types rather than one blended score. For example, cost-first for background work and latency-first for real-time chat. The DigitalOcean Inference Router supports this directly. Each task in a router can have its own model pool and its own Cost Efficiency, Speed Optimization, or Manual Ranking policy.

References & further reading

  • How to use Inference Router: The full documentation on tasks, selection policies, fallback models, and the cache-aware router referenced throughout this piece. Worth reading directly before you configure a router for production traffic.
  • Inference pricing: Current per-model Serverless Inference rates, since routing decisions are only as good as the pricing data behind them.
  • How DigitalOcean built Inference Router: An inside look at the architecture decisions behind cache-aware routing and fallback handling, useful if you want the reasoning behind the mechanics.
  • DigitalOcean Serverless Inference: a deep dive: Background on the serverless layer that Inference Router sits in front of.
  • DigitalOcean's best LLM routers roundup: A good next read if you're comparing routing options across vendors rather than implementing a policy on a platform you've already chosen.
  • AI Inference Engine overview: shows where Inference Router fits alongside serverless, batch, and dedicated inference on DigitalOcean's broader platform.

Top comments (0)