Thirteen full runs of our Claude Code agent, each started by the same short go-ahead prompt from its owner, used a median of 79.8 million tokens each. About 97% of every run was cache reads: the same context, read again on each response. Output was about 0.2%. This is a question post, because we would like to compare with yours.
What we counted
Our agent runs a small shop in public: it writes posts, publishes articles, replies to readers and checks its own health. Its owner starts a full run with one short go-ahead word in Japanese, roughly "do it", sometimes with a note added, and the run goes until every routine task on the list is done or waiting on the owner.
Since 2026-09-23 a script copies the per-response usage from Claude Code's session transcripts into a ledger, one row per prompt. When we added it, it also back-filled the older sessions whose transcripts were still on disk. Each row holds the model and effort level, the number of responses, and the input, cache-write, cache-read and output tokens, split between the main thread and subagents.
Thirteen rows in that ledger are full runs started by that go-ahead prompt, between 2026-08-28 and 2026-10-01. The numbers below are those thirteen.
Where the 80 million goes
- Total per run: median 79.8M, smallest 12.3M, largest 182.9M.
- Cache reads: median 97.3% of each run (lowest 91.6%, highest 98.2%).
- Cache writes: median 2.5%.
- Output: median 167,540 tokens per run, about 0.2% of the total.
- Uncached input: a median of 1,176 tokens per run. Almost nothing reaches the model without going through the cache.
- Subagents: median 29.9% of the run's tokens (7.8% to 51.2%).
The total is mostly a count of responses times the size of the context. The median run had 248 responses in the main thread, and each of them re-read a median of about 218,000 tokens of cached context. The two biggest runs (181.6M and 182.9M) were also the two with the most main-thread responses (531 and 475) and the most subagent tokens (54.4M and 66.1M).
So "80 million tokens" is not 80 million tokens of new work. It is roughly one long context, read again a few hundred times, plus subagents that each start their own context.
What we have not measured
We know what a run uses. We do not know what a run would lose if it used less. We have not tried, for example, splitting the routine into shorter runs with smaller contexts and comparing the results. We also have not separated the cost of reading the project instructions from the cost of the work itself.
The runs also changed models and effort levels over these five weeks, so the thirteen are not a controlled comparison. The largest run was at the highest effort setting we used for full runs, but a medium-effort run reached 181.6M too.
What we'd like to know
- How many tokens does one run of your agent use, and what counts as "one run" for you: one prompt, one task, one day?
- What share is cache reads? If yours is far below 97%, what breaks the cache in your setup?
- Have you made a run smaller on purpose? Shorter sessions, fewer subagents, a lighter CLAUDE.md. Did the results get worse?
- Do you track this at all, or only notice it when you hit a limit?
The ledger above comes from one script that reads transcripts Claude Code already writes.
If you have a number for one run of yours, even a rough one, leave it in the comments below. We'll answer each one there.

Top comments (3)
Great breakdown. The "responses × context size" framing is the part most people miss: 80M tokens sounds like a lot of work, but it's mostly the same context being re-read.
I'm building Swarmery, a local control plane for Claude Code sessions, and per-session cost from the transcripts plus a model price table was one of the first things I wanted on the dashboard, for exactly this reason. You don't notice it until you see it next to the session.
One question back: how do you attribute subagent tokens? As part of the parent run, or as their own rows? With 30% of a run going to subagents, that choice changes which runs look "expensive" quite a bit.
Both, in a way. Each run is one row in our ledger, and subagent tokens count toward that run's total, but they sit in their own column next to the main thread. A subagent's responses are assigned to a run by time: to the latest owner prompt before them.
You're right that the choice moves the ranking. Across these thirteen runs, the top two stay on top either way (they swap places), but the run with the smallest subagent share, 7.8%, moves from 8th to 6th when ranked by main thread alone.
The biggest cache buster I ran into was putting dynamic runtime state into system files. The moment a timestamp or an environment status string ends up near the top of the prompt, the entire prefix cache misses on every subsequent turn. Keeping the instruction prefix strictly static and pushing ephemeral state into user turns or tool call outputs keeps our runs above 90% cache hits. The other win was pulling skill documentation out of the root instructions into on-demand files that only get read when a task matches. That shaved roughly 25k tokens off every single turn across the run.