DEV Community

Cover image for 200 agents, 2,011,438 tool calls: who's paying for your AI?
TokenLat
TokenLat

Posted on AI-assisted

200 agents, 2,011,438 tool calls: who's paying for your AI?

A team I worked with shipped 200 internal agents last quarter. The monthly bill landed at 2,011,438 tool calls. Finance asked the usual thing: "which model is eating our budget?"

The intuitive answer — "the frontier model, obviously" — was wrong. And why it was wrong says something useful about how agentic systems actually spend money.

The counterintuitive part

Most token spend in an agent loop isn't in the hard reasoning step you're picturing. It's in the boring middle:

  • the 60-line retrieval that reformats a doc the model already saw,
  • the JSON reshaping between two tools,
  • the "which tool next?" classifier that fires before every action,
  • the summary pass that condenses a transcript nobody reads.

None of those need a frontier model. They need a model — a cheap one, routed to that specific sub-task. But the agent was wired with one "default" model for the whole loop, so every sub-task paid frontier prices.

We measured it: roughly 80% of the bill was sub-tasks that cleared the quality bar at ~1/20th the cost on a smaller model.

This is a routing problem, not a prompting problem

You don't fix this with a cleverer system prompt. You fix it by making the sub-task boundary visible to a router.

Instead of agent.llm = frontier_model, each sub-task declares what it needs:

subtasks:
  retrieve_context:
    min_capability: embedding-lookup
    route_to: small-retriever      # cheap, no reasoning needed
  classify_next_tool:
    min_capability: light-classification
    route_to: mini-classifier      # tiny, fast
  synthesize_answer:
    min_capability: frontier-reasoning
    route_to: frontier-model       # only here do we pay up
Enter fullscreen mode Exit fullscreen mode

The router picks the cheapest model that clears min_capability. The agent's behavior doesn't change. The bill does.

flowchart LR
    A[Sub-task emitted] --> R{Router}
    R -->|embedding-lookup| S[small-retriever]
    R -->|light-classification| M[mini-classifier]
    R -->|frontier-reasoning| F[frontier-model]
    S --> L[per-step ledger]
    M --> L
    F --> L
    L --> C{PDPA-aligned?}
    C -->|yes| OK[serve in SG-hosted region]
    C -->|no| B[route to compliant region]

Per-step cost attribution is the missing half

Routing only helps if you can see it. Most teams have one line item: "LLM." That hides the problem — finance sees a number, not a decision.

Attach cost to every sub-task, not every request:

def route(subtask):
    model = cheapest_that_clears(subtask.min_capability)
    cost = price_per_1k[model] * estimate_tokens(subtask)
    ledger.log(subtask.name, model, cost)   # per-step, not per-request
    return model
Enter fullscreen mode Exit fullscreen mode

Now the bill reads by capability tier: "retrieval $X, classification $Y, synthesis $Z." When synthesis balloons, you know which sub-task to attack — not which model to blame.

Cost and compliance converge on the same layer

Here's what surprised the team: the router was also where data-residency rules had to live. Some sub-tasks touched regulated customer data, so they had to stay in a PDPA-aligned, SG-hosted region even when a cheaper model sat elsewhere.

So the same routing layer that cut the bill also enforced jurisdiction. Cost optimization and compliance stopped being separate meetings — they became one config block.

That's the real unlock: a routing layer isn't infrastructure you tolerate. It's where your cost and your constraints get expressed as code.

The question I'll leave you with

If you mapped every sub-task in your agent loop to the cheapest model that clears its quality bar, how much of your current bill would survive?

No tool to sell, no link to click. Just this: most agent bills are really routing bills in disguise.

Top comments (3)

Collapse
 
arhancanli profile image
Arhan Canli •

The 80% at ~1/20th figure gives a ceiling you can check: if 80% of the bill moves to a model at 5% of the price, the bill becomes 0.20 + 0.80 x 0.05 = 0.24 of the original, so about 76% off, and only if routing is free and nothing escalates back to the frontier model. The part I'd want to see is how "cleared the quality bar" was measured per sub-task. A classifier step that is right 97% of the time on 300 labelled calls has a 95% interval of roughly 94.4% to 98.4%, and in an agent loop a wrong "which tool next?" does not cost one bad answer, it costs every step after it. So the useful number per tier is the error rate on that sub-task plus the retry or fallback cost it triggers, logged next to the price in the same per-step ledger you describe. Did you replay a sample of the cheap-tier outputs against the frontier model's before moving traffic, or was the bar judged on end-task success?

Collapse
 
tokenlat profile image
TokenLat •

Good catch — and your ceiling math is exactly the optimistic bound: 0.20 + 0.80×0.05 = 0.24, so ~76% off, and it silently assumes routing is free and that nothing escalates back to the frontier. Both are where real deployments bleed.

On measurement: the bar was set per sub-task, not on end-task success. Each tier got a held-out labelled set, and before traffic moved we replayed a sample of the cheap-tier outputs against the frontier and scored them on that sub-task's rubric — not just "did the final answer look right." You're right that a 97%-accurate classifier on 300 calls is a wide interval, and in an agent loop the cost compounds, so the ledger logs the per-tier error rate next to its price and the fallback cost it triggers.

The assumption I'd trust least is the "nothing escalates back" one; in practice a small, monitored fraction does, and that's the term that eats the savings. If you've got a sub-task where you suspect the cheap tier is silently wrong, that's the highest-leverage place to look.

Collapse
 
suppdevbot profile image
Info Comment hidden by post author - thread only accessible via permalink
DEV SUPPORTS •

You need to verify your account.

Enter fullscreen mode Exit fullscreen mode

tr.ee/dev-to

Some comments have been hidden by the post's author - find out more