DEV Community

Cover image for Cost per token is the wrong number for coding agents. Here's what we measured instead.
Yugesh Jha
Yugesh Jha

Posted on

Cost per token is the wrong number for coding agents. Here's what we measured instead.

We build AstraCode, an AI code editor whose agent has to prove its own work. Model calls are the largest line on our bill, so in September we stopped guessing and measured: eleven models, the same set of real coding tasks, each run more than once, scored on pass rate and on cost per finished task.

Step count beats token price

The cheapest model per token is not the cheapest model per task. On one task a model took 130 agent steps; another finished in 32. Every step re-sends context, so the step count multiplies everything else.

It is not as simple as "expensive models wander more" either. Some of the priciest models took fewer steps than the cheap one. We wrote that rule down, then had to cross it out.

Takeaway: log steps per task alongside tokens. A model that is 3x cheaper per token and takes 4x the steps is a more expensive model.

Routing by prompt doesn't work

The obvious optimisation is a router: read the request, send easy ones to a small model and hard ones to a big one. We tried three versions of it.

  • A per-step router came out 0.4% worse than not routing at all.
  • A classifier that predicted difficulty from the prompt reached a rank correlation of about 0.75 and still lost.
  • Even an oracle with perfect hindsight picked the small model on every task, so there was no headroom to win.

The reason is variance. The same small model passed one task 67% of the time on the identical prompt. Whether a run succeeds is decided during the attempt, not by the prompt, so no prompt classifier can see it.

Detection does work

What does work is checking the result:

  1. Run the small model.
  2. Check its work. Run the tests, then undo the change and confirm those tests fail without it, so a test that passes either way doesn't count as proof.
  3. Only if a check fails, hand the task to a stronger model.

Pass rate went from 0.64 to about 1.00, at roughly $0.20 per task against $0.03 for the small model alone.

These are our tasks and our harness, so treat them as one data point, not a ranking. Every model we tested is good at something.

What we changed

AstraCode no longer asks you to pick a model. AstraOne starts with the cheaper model and escalates when a check fails. The checks are the same ones you'd want anyway: tests that are shown to fail without the change, a diff you review hunk by hunk, and a checkpoint before every turn.

If you're building on LLMs, two habits are worth more than any model choice:

  • measure cost per finished task, not per token;
  • run every task several times before you trust a number.

AstraCode is free to start, no card: https://astracode.io

Top comments (0)