A recent post here by @ji_ai measured something startling: Claude Code subagents were 48% of their bill while producing 0.9% of the output. The culprit: every Task() call re-sends a fixed preamble of roughly 51K tokens — system prompt, full tool schemas, CLAUDE.md, skill listings — before the subagent even reads your prompt.
I wanted that number for my own usage, not just one anecdote. So I built subagent-tax: a zero-dependency Python CLI that scans your local ~/.claude/projects transcripts, counts every subagent invocation, and converts it into tokens and dollars — ranked by what's actually cuttable.
What it does
pip install subagent-tax
subagent-tax
Task (subagent) calls : 132 across 18 session(s)
Preamble model: tokens re-sent per Task call [heuristic]
system prompt 20,000 tok
built-in tool schemas 8,000 tok
MCP tool schemas 16,000 tok
CLAUDE.md 3,000 tok
skill listings 4,000 tok
TOTAL per call 51,000 tok
Estimated waste: 132 calls x 51,000 tok = 6,732,000 tokens ~= $20.20
Cuttable contributions (tokens per Task call):
1. system prompt 20,000 tok/call
2. MCP server: playwright 10,000 tok/call
...
Top trim suggestion:
Slim down (custom system prompt) system prompt: saves ~20,000 tokens
per Task call (~$7.92 at 132 observed calls)
The interesting part isn't the total — it's the ranking. The tool tells you which MCP server, which skill, or which chunk of CLAUDE.md costs the most per call, with a concrete trim suggestion ("disable MCP server X, save ~N tokens per call").
Calibrate it to your setup
The 51K default comes from the published measurement, but your preamble is yours. Feed it real data:
# measure your actual CLAUDE.md and skills instead of heuristics
subagent-tax --claude-md ./CLAUDE.md --skills-dir ~/.claude/skills
# import per-server schema sizes from mcp-tax (my earlier tool)
mcp-tax audit --json > /tmp/mcp.json
subagent-tax --mcp-tax-report /tmp/mcp.json
# or override any component directly
subagent-tax --set system_prompt=15000 --price-input 3.00 --format json
Where it fits: mcp-tax vs subagent-tax
I built mcp-tax earlier to audit the total size of MCP server schemas. subagent-tax is the other half: the repeat cost — how many tokens get re-sent on every single subagent call because the preamble is fixed. They compose: mcp-tax measures it once, subagent-tax multiplies it by your call count.
Honest limitations
No exact token counts without real API payloads — text is estimated at ~4 chars/token, and dollar amounts use a pricing input you should verify against current published rates. It only sees local JSONL transcripts, and "waste" is a simplification: some preamble is working context, not pure waste. The README documents all of this.
Links: GitHub · pip install subagent-tax · MIT licensed, stdlib-only, zero dependencies.
If you run it, I'd genuinely like to know: what's your most expensive preamble component?
Top comments (0)