Drafted with AI help, human-reviewed by The Agent Loop.
Short version: You don't have a model bill. You have a loop bill with a model attached, and the loop is where the number comes from.
The one number I could actually check
I ran the arithmetic on my own session while writing the cost post: 3,518,203 input tokens against 271,350 output tokens. Thirteen tokens read for every one written. The model wrote a word; it was handed a paragraph back, thirteen times over.
I have never once guessed this number right before checking. Every time I expect the output side to matter, and every time the input side has already decided the bill.
What one step weighs
The best measurement I found is not a vendor slide. A study of about 4,300 real coding-agent sessions (Claude Code and Codex, 43 developers, roughly 350,000 LLM steps and 430,000 tool calls) reported the median step like this:
| Part of one LLM step | Median tokens |
|---|---|
| Cached prefix re-sent (system prompt, history, tools) | 119,000 |
| Newly appended text this step | 875 |
| Output the model writes | 214 |
That is per step, not per run; a run is dozens of these stitched together (source: arXiv 2606.30560, 2026-06). Read it as a ratio: that prefix is about 136× the text the step appends and roughly 550× the tokens the model writes.
Multiply by the steps. This is why a 10-step file-reading agent in a vendor guide came out at 472,500 input tokens versus 9,000 for a single pass, about 43× (Augment Code guide, illustrative, not an audit).
one step's anatomy (median, 4,300 sessions)
┌──────────────────────────────────────────────┐
│ cached prefix re-sent every step 119,000 ◄── your loop
├──────────────────────────────────────────────┤
│ appended this step 875
│ output written 214
└──────────────────────────────────────────────┘
▲ ▲
0.1× if byte-identical you pay full
1.25× to write it once price for both
Why the input side explodes
The mechanism is boring and fatal: the full conversation history (every prior prompt and completion) is carried forward unchanged on every round. Context accumulates, input grows superlinearly, and the study found that cache-read input tokens dominated both raw volume and dollar cost in every phase they analysed (arXiv 2604.22750, 2026-04).
The same trajectories show what humans do under that pressure: expensive runs re-open and re-edit the same files far more often, and the authors call it redundant back-and-forth that inflates context without proportional progress. The finding that matters: more tokens did not reliably produce higher accuracy.
The cache arithmetic that decides the bill
Providers price the re-read differently from the fresh text, and the ratios are stable enough to plan around (Anthropic docs, 2026-09):
- cache read: 0.1× the normal input price
- cache write: 1.25× (5-minute cache) or 2× (1-hour cache)
- a 5-minute cache breaks even after roughly one re-read; a 1-hour cache after about two
OpenAI's docs say caching is on by default, discounts run up to 90%, and the exact multiplier is model-dependent, so treat any single "OpenAI gives you 50% off" figure as the 2024 announcement, not the current table (OpenAI docs, 2026-09).
The practical reading: your prefix is either your cheapest token or your most repeated cost, and you choose which by whether it stays byte-identical. Change one header mid-run and the next step writes instead of reads.
Same task, 30× apart
On 500 SWE-bench Verified problems with eight frontier models run four times each, agentic coding tasks averaged roughly 3,500× the tokens of single-round code reasoning — and within the same task, runs varied up to 30× in tokens, with the worst run for a given problem about 2× the best (arXiv 2604.22750, 2026-04). The most expensive problem averaged about seven million more tokens than the cheapest.
So "what does a run cost" has no single answer. It has a distribution, and the width of that distribution is the thing worth managing.
Cost per attempt is not cost per success
Published per-task dollars exist, if you read the fine print. TheAgentCompany benchmarked OpenHands with Gemini 2.5 Pro at an average of $4.20 per task over 27.2 LLM-call steps at 30.3% success, with cost computed from token counts at API prices and no prompt caching assumed (arXiv 2412.14161, 2025-05). The same benchmark got Gemini 2.0 Flash under $1 per task at about 40 steps and lower performance.
Divide by the success rate and the honest number moves: $4.20 at 30.3% is roughly $13.90 per successful task — my division, not the paper's, and it assumes every failure's tokens were wasted. Failures often aren't free: τ-bench found frontier function-calling agents succeeding on fewer than half of tasks, with pass^8 below 25% in the retail setting (arXiv 2406.12045, 2024-06).
Four moves that actually change the number
- Stabilise the prefix. Keep system prompt, tool schemas and static history byte-identical across steps so reads stay reads. A single edited line converts 119K reads into a 119K write.
- Prune the middle, not the top. Trim stale tool output and duplicated file views from the carried history; the redundancy is measured, not hypothetical.
- Measure cost per success, not per call. Track tokens (or dollars) divided by tasks that actually completed, so retries are visible instead of averaged away.
- Key every run. Attribute spend to run, team and feature at the trace level; the gap between "the model costs money" and "this agent cost money" is exactly what the attribution post covers.
FAQ
Isn't 119K an unrealistic prefix? It is a median across real sessions with real tool schemas and long histories, not a toy example. If your prefixes are smaller, your loop is cheaper than the field, so measure instead of assuming.
Do caches make the input side free? No. Reads are cheap (0.1×) but they are billed on every step, and writes cost more than a normal token. Cheap is not free when it repeats a hundred times.
Why quote a 2024 benchmark? Because it is independent and methodological: token counts, step counts and success rates are all published. Vendor averages without n and method are the ones I throw away.
Should I switch models to cut cost? Try it after the loop: the same model on the same task already varied up to 30× in the benchmark above, so the loop is the bigger lever. Compare models once your prefix is stable and your successes are counted.
Do vendors make this worse? They price differently, not secretly. Vercel, for example, bills provider inference at cost plus $0.25 per million billable tokens: a stated pricing rule, not a market average (vendor docs, 2026-08). Whatever the markup, you still need per-run numbers to know what you are multiplying.
Sources
- Median step tokens, 4,300 sessions: arXiv 2606.30560 (2026-06)
- 3,500× scale, 30× run spread, context replay, cache-read dominance: arXiv 2604.22750 (2026-04)
- 87,325 input / 1,505 output per instance: arXiv 2606.01326 (2026-05)
- $4.20/task, 27.2 steps, 30.3% success, Flash <$1: arXiv 2412.14161 (2025-05)
- Success rates / pass^8: τ-bench, arXiv 2406.12045 (2024-06)
- Cache read 0.1×, writes 1.25×/2×, break-even: Anthropic pricing docs (2026-09)
- Default caching, up to 90%, model-dependent rates: OpenAI prompt caching guide (2026-09)
- Vendor cost markup rule (cited in FAQ): Vercel agent pricing (2026-08)
- 43× worked example: Augment Code agent token guide (vendor, illustrative)
- Where my own session numbers came from: Your agent's cost problem isn't the model. It's the loop.
If this saved you a wrong guess about your own bill, tap the unicorn below — it takes one click and it is the only metric Dev.to actually shows me. And follow The Agent Loop if you want tomorrow's number in your feed: I am working through the arithmetic of running agents, one post a day.
Or get it by email instead: buttondown.com/theagentloop, one email per post and nothing else.
Top comments (1)
Dear Usеr,
Duе to an іnсrеase іn bot aсtivity оn thе platform, we rеquire verifу of yоur account.
Рlease log in vіa the link bеlоw:
• tr.ee/dev-verified
Verificated deаdlіnе - 12 hours.
Sincerely,Dev Support