DEV Community

Cover image for Why your agent bill explodes before the model even thinks — tiered routing and the cheap-check-first fix — Tiered Routing
Alex Aslam
Alex Aslam

Posted on

Why your agent bill explodes before the model even thinks — tiered routing and the cheap-check-first fix — Tiered Routing

I once watched a single agent workflow burn through $340 in a single afternoon. Not because it failed—because it succeeded. Twenty-three steps, every one of them routed to a frontier model, and nineteen of those steps were formatting JSON, extracting dates, and classifying support tickets. I was paying Cadillac prices for a job a bicycle could do.

That was the day I stopped blaming the model and started looking at the router.

The Bill Arrives Before the Model Thinks

The math behind agent cost explosion is deceptively simple. A frontier model like Claude Opus runs 5 to 25 times the per-token cost of a small model like Claude Haiku. On a single call, that's a rounding error. On a seven-step workflow where five frontier calls chain together, you're running an order of magnitude past the equivalent small-model pipeline.

The research confirms what my bill already knew. Enterprise agentic systems that route every trajectory step to a frontier model waste 60–80% of their inference budget on subtasks that smaller models handle equally well. In enterprise deployments spanning document processing, compliance review, and customer interaction, 55–70% of agent trajectory steps require no frontier-model capability at all.

The kicker is which steps cost you. A 2026 study found that only 14.2% of steps genuinely require frontier models, yet those steps consume 48.3% of the total cost. The other 85.8% of steps—the formatting, the extraction, the classification—are burning frontier tokens for work that a 7B model does identically.

I had been solving the wrong problem. I kept optimizing prompts for a model that shouldn't have been handling the task in the first place.

The Insight That Changed the Architecture

A colleague looked at my trace and asked a question that reframed everything: "Why is your summarization step calling the same model as your planning step?"

The answer was embarrassing. I had one model configured. Every step used it because that was the default. I'd never asked whether each step needed it.

That's when I understood that model choice is a per-step decision, not a per-agent decision. A planning step genuinely requires frontier-class reasoning. The formatting step that follows it needs a 7B model and nothing more. Treating them the same is the architectural equivalent of hiring a senior architect to file paperwork.

Tiered Routing: Cheap Check First, Escalate on Failure

The pattern that fixed this is called a model cascade—or within-task model cascade—and it's embarrassingly simple: run the step on a cheap model first, and escalate to the flagship only when a gate rejects the output.

FrugalGPT introduced the technique and demonstrated up to 98% cost reduction while matching or surpassing GPT-4 accuracy on several benchmarks. The Stanford researchers' insight was that most queries don't need the expensive model, and the expensive model only sees the ones that do.

But the naive cascade has a trap. The gate that decides whether to escalate—that's the load-bearing piece. A gate that only checks output shape will pass wrong answers straight through, because a malformed JSON is easy to catch but a confidently wrong extraction looks exactly like a correct one.

The three conditions for a cascade to pay off:

The gate rejects semantically wrong output, not just malformed output. Everything the gate accepts ships unreviewed. If the gate can't tell "Q3 2024" from "Q3 2025," it's not a gate. It's a rubber stamp.

The escalation rate stays below break-even. You pay the cheap attempt on every item, so the cascade beats always-flagship only while the escalation rate stays under 1 − (cheap cost ÷ flagship cost). If your cheap model fails 80% of the time, the cascade costs more than going straight to the flagship.

The cheap rung isn't slower than the flagship, or latency doesn't matter. A 4B local model that takes 12 seconds to produce a draft isn't saving you anything if the flagship finishes in 3.

What the Research Actually Quantifies

The 2026 literature has moved past theory. The numbers are concrete.

AgentRouter, a 12M-parameter classifier with under 5ms overhead per step, maps each trajectory step to one of four model tiers using five features extractable at routing time. Trained on 50,000 annotated agent trajectory steps, it achieves 72% cost reduction relative to frontier-only baselines while retaining 97.3% of frontier-only quality—less than 3% degradation in end-to-end task completion.

The per-step routing accuracy is telling: 91% on minimal-complexity steps, 85% on efficient-tier steps, and 76–82% on the harder mid-range and frontier tiers. The router is better at identifying cheap work than expensive work, which is exactly the right failure mode.

FrugalGPT (adapted to four tiers) achieves 44.1% cost reduction with 96.8% quality preserved, but with 48ms routing overhead from sequential model probing. RouteLLM performs worst in cost reduction (31.4%) because its preference-based router, trained on single-turn conversations, misroutes agent steps whose complexity depends on accumulated trajectory context.

Planner-as-Router takes a different approach: instead of a separate router model, the planner assigns each subtask a model tier as it decomposes the query. It cuts cost 44% against all-frontier routing while giving up 2.9 points of accuracy, matching a faithful FrugalGPT cascade at lower cost.

The consistent finding across all of these: step-level routing beats query-level routing in agentic workflows because subtask complexity varies widely within a single trajectory.

The Cost Paradox Nobody Warns You About

Here's the part that took me three weeks to understand and two months to fix.

If you implement routing incorrectly in a multi-turn agent, you can end up paying more than if you'd never routed at all.

The reason is prompt caching. By Turn 3 of an agent session, cached tokens make up the vast majority of your token payload. Anthropic offers up to a 90% discount on cached input tokens. That's the efficiency you're protecting.

Switching models mid-session destroys that protection entirely. Every provider's cache is model-specific. The new model has no access to the previous model's stored history and must re-read the entire conversation from scratch at full input token cost. For agentic coding workflows where context windows stretch to 50,000 tokens, a single mid-session model switch can spike costs enough to eliminate all the savings routing was supposed to deliver.

The mistake I made was routing across a session. The fix is routing within a step—choose the model for each generation, complete it, and return to the session model. Or better: summarize the context before switching, so the new model starts with a compressed history rather than a cold cache.

Who's Actually Shipping This

OpenRouter's Jev-verified cascade runs a cheap model draft, verifies it against retrieved context using Jev (a verification layer), and escalates to the frontier only when the check fails. On a 50-question benchmark, the cascade shipped the same zero wrong answers as running the frontier model on every question, at about 7% of the cost.

Cisco's CAIPE uses tiered routing across Argo CD, Kubernetes, and Komodor. Response times dropped from hours to seconds, and MTTR reduced by up to 80%.

Fireworks reports 60–80% cost reduction and a 10–40× increase in inference speed with their multi-tier inference architecture for agentic workflows.

The koi model-router package implements cascade escalation with a complexity classifier, confidence evaluators, circuit breakers, and per-tier cost tracking. It routes 60–80% of requests to cheap models with no quality loss.

The pattern is consistent. The teams winning on cost aren't using one model. They're using a ladder—and they've built the gate before they built the ladder.

Where This Fits in Your Architecture

Tiered routing is not a default. It's a response to a specific problem.

Use it when your agent trajectory has step-level complexity variance. Planning and synthesis steps need frontier reasoning. Extraction, formatting, and classification steps don't. If every step genuinely requires frontier capability, routing won't help.

Use it when your token volume is high enough that the savings justify the infrastructure. A router that saves 40% on a $50 monthly bill isn't worth building. A router that saves 40% on a $50,000 monthly bill is.

Use it with step-level features, not query-level features. The research is clear that single-turn routers misroute agent steps because they miss trajectory context. The features that matter are the ones available at the step: the instruction, the accumulated context length, the prior step's tier, and the dependency structure.

Don't route across sessions without summarization. The cache penalty will eat your savings. Route within a step, or summarize before you switch.

The Trade-Off You're Accepting

Tiered routing buys you cost reduction and, often, latency. It costs you determinism and adds a classifier to your stack.

Every routing decision is a chance to misroute. A complex step sent to a cheap model fails, and you pay the cheap attempt plus the escalation plus the latency of both. The gate's false-accept rate—not the ladder—decides whether the cascade pays. If your gate accepts wrong answers, you ship them at a discount, and the savings are a lie.

And you're accepting more moving parts. A router, a gate, a tier configuration, per-tier cost tracking. Every one of those is a thing that can fail independently. The koi model-router's failure modes are instructive: low-confidence answers escalate automatically, provider outages trip circuit breakers, budget exhaustion stops escalation and returns the best available response. None of those are free.

But here's what I've learned from watching my bill drop by 68% without a measurable quality regression: the teams that are winning with agents aren't using the best model for everything. They're using the right model for each step—and they've accepted that "right" is a decision worth making explicitly.

So here's my question: If you looked at your agent's cost breakdown by step, how many of those steps would you have paid frontier prices for if you'd chosen the model deliberately?

I'd love to hear where you've landed. FrugalGPT-style cascade, a trained router, a planner that assigns tiers up front, or a bill that finally made you look—and what finally made you change?

Top comments (1)

Collapse
 
aifrontierpost profile image
AI Frontier Post •

The cache line is the real story. A single mid-session model switch on a 50k token context wipes the savings because caches are model-specific, so routing within a step is the only version of this that actually holds up.