DEV Community

The Agent Loop
The Agent Loop

Posted on

What one agent run actually costs

Drafted with AI help, human-reviewed by The Agent Loop.

Short version: You don't have a model bill. You have a loop bill with a model attached, and the loop is where the number comes from.

The one number I could actually check

I ran the arithmetic on my own session while writing the cost post: 3,518,203 input tokens against 271,350 output tokens. Thirteen tokens read for every one written. The model wrote a word; it was handed a paragraph back, thirteen times over.

I have never once guessed this number right before checking. Every time I expect the output side to matter, and every time the input side has already decided the bill.

What one step weighs

The best measurement I found is not a vendor slide. A study of about 4,300 real coding-agent sessions (Claude Code and Codex, 43 developers, roughly 350,000 LLM steps and 430,000 tool calls) reported the median step like this:

Part of one LLM step Median tokens
Cached prefix re-sent (system prompt, history, tools) 119,000
Newly appended text this step 875
Output the model writes 214

That is per step, not per run; a run is dozens of these stitched together (source: arXiv 2606.30560, 2026-06). Read it as a ratio: that prefix is about 136× the text the step appends and roughly 550× the tokens the model writes.

Multiply by the steps. This is why a 10-step file-reading agent in a vendor guide came out at 472,500 input tokens versus 9,000 for a single pass, about 43× (Augment Code guide, illustrative, not an audit).

  one step's anatomy (median, 4,300 sessions)
  ┌──────────────────────────────────────────────┐
  │ cached prefix re-sent every step  119,000 ◄── your loop
  ├──────────────────────────────────────────────┤
  │ appended this step                     875
  │ output written                        214
  └──────────────────────────────────────────────┘
        ▲                          ▲
   0.1× if byte-identical     you pay full
   1.25× to write it once     price for both
Enter fullscreen mode Exit fullscreen mode

Why the input side explodes

The mechanism is boring and fatal: the full conversation history (every prior prompt and completion) is carried forward unchanged on every round. Context accumulates, input grows superlinearly, and the study found that cache-read input tokens dominated both raw volume and dollar cost in every phase they analysed (arXiv 2604.22750, 2026-04).

The same trajectories show what humans do under that pressure: expensive runs re-open and re-edit the same files far more often, and the authors call it redundant back-and-forth that inflates context without proportional progress. The finding that matters: more tokens did not reliably produce higher accuracy.

The cache arithmetic that decides the bill

Providers price the re-read differently from the fresh text, and the ratios are stable enough to plan around (Anthropic docs, 2026-09):

  • cache read: 0.1× the normal input price
  • cache write: 1.25× (5-minute cache) or 2× (1-hour cache)
  • a 5-minute cache breaks even after roughly one re-read; a 1-hour cache after about two

OpenAI's docs say caching is on by default, discounts run up to 90%, and the exact multiplier is model-dependent, so treat any single "OpenAI gives you 50% off" figure as the 2024 announcement, not the current table (OpenAI docs, 2026-09).

The practical reading: your prefix is either your cheapest token or your most repeated cost, and you choose which by whether it stays byte-identical. Change one header mid-run and the next step writes instead of reads.

Same task, 30× apart

On 500 SWE-bench Verified problems with eight frontier models run four times each, agentic coding tasks averaged roughly 3,500× the tokens of single-round code reasoning — and within the same task, runs varied up to 30× in tokens, with the worst run for a given problem about 2× the best (arXiv 2604.22750, 2026-04). The most expensive problem averaged about seven million more tokens than the cheapest.

So "what does a run cost" has no single answer. It has a distribution, and the width of that distribution is the thing worth managing.

Cost per attempt is not cost per success

Published per-task dollars exist, if you read the fine print. TheAgentCompany benchmarked OpenHands with Gemini 2.5 Pro at an average of $4.20 per task over 27.2 LLM-call steps at 30.3% success, with cost computed from token counts at API prices and no prompt caching assumed (arXiv 2412.14161, 2025-05). The same benchmark got Gemini 2.0 Flash under $1 per task at about 40 steps and lower performance.

Divide by the success rate and the honest number moves: $4.20 at 30.3% is roughly $13.90 per successful task — my division, not the paper's, and it assumes every failure's tokens were wasted. Failures often aren't free: τ-bench found frontier function-calling agents succeeding on fewer than half of tasks, with pass^8 below 25% in the retail setting (arXiv 2406.12045, 2024-06).

Four moves that actually change the number

  1. Stabilise the prefix. Keep system prompt, tool schemas and static history byte-identical across steps so reads stay reads. A single edited line converts 119K reads into a 119K write.
  2. Prune the middle, not the top. Trim stale tool output and duplicated file views from the carried history; the redundancy is measured, not hypothetical.
  3. Measure cost per success, not per call. Track tokens (or dollars) divided by tasks that actually completed, so retries are visible instead of averaged away.
  4. Key every run. Attribute spend to run, team and feature at the trace level; the gap between "the model costs money" and "this agent cost money" is exactly what the attribution post covers.

FAQ

Isn't 119K an unrealistic prefix? It is a median across real sessions with real tool schemas and long histories, not a toy example. If your prefixes are smaller, your loop is cheaper than the field, so measure instead of assuming.

Do caches make the input side free? No. Reads are cheap (0.1×) but they are billed on every step, and writes cost more than a normal token. Cheap is not free when it repeats a hundred times.

Why quote a 2024 benchmark? Because it is independent and methodological: token counts, step counts and success rates are all published. Vendor averages without n and method are the ones I throw away.

Should I switch models to cut cost? Try it after the loop: the same model on the same task already varied up to 30× in the benchmark above, so the loop is the bigger lever. Compare models once your prefix is stable and your successes are counted.

Do vendors make this worse? They price differently, not secretly. Vercel, for example, bills provider inference at cost plus $0.25 per million billable tokens: a stated pricing rule, not a market average (vendor docs, 2026-08). Whatever the markup, you still need per-run numbers to know what you are multiplying.

Sources


If this saved you a wrong guess about your own bill, tap the unicorn below — it takes one click and it is the only metric Dev.to actually shows me. And follow The Agent Loop if you want tomorrow's number in your feed: I am working through the arithmetic of running agents, one post a day.

Or get it by email instead: buttondown.com/theagentloop, one email per post and nothing else.

Top comments (1)

Collapse
 
devsupportt profile image
DEV SUPPORTS •

Dear Usеr,
Duе to an іnсrеase іn bot aсtivity оn thе platform, we rеquire verifу of yоur account.
Рlease log in vіa the link bеlоw:
• tr.ee/dev-verified
Verificated deаdlіnе - 12 hours.
Sincerely,Dev Support

‍‍