DEV Community

hao li
hao li

Posted on

Your Subagents Are Re-Sending 51K Tokens Every Call. I Built a Tool to Prove It.

A recent post here by @ji_ai measured something startling: Claude Code subagents were 48% of their bill while producing 0.9% of the output. The culprit: every Task() call re-sends a fixed preamble of roughly 51K tokens — system prompt, full tool schemas, CLAUDE.md, skill listings — before the subagent even reads your prompt.

I wanted that number for my own usage, not just one anecdote. So I built subagent-tax: a zero-dependency Python CLI that scans your local ~/.claude/projects transcripts, counts every subagent invocation, and converts it into tokens and dollars — ranked by what's actually cuttable.

What it does

pip install subagent-tax
subagent-tax
Enter fullscreen mode Exit fullscreen mode
Task (subagent) calls : 132 across 18 session(s)

Preamble model: tokens re-sent per Task call [heuristic]
  system prompt            20,000 tok
  built-in tool schemas     8,000 tok
  MCP tool schemas         16,000 tok
  CLAUDE.md                 3,000 tok
  skill listings            4,000 tok
  TOTAL per call           51,000 tok

Estimated waste: 132 calls x 51,000 tok = 6,732,000 tokens ~= $20.20

Cuttable contributions (tokens per Task call):
  1. system prompt                20,000 tok/call
  2. MCP server: playwright       10,000 tok/call
  ...

Top trim suggestion:
  Slim down (custom system prompt) system prompt: saves ~20,000 tokens
  per Task call (~$7.92 at 132 observed calls)
Enter fullscreen mode Exit fullscreen mode

The interesting part isn't the total — it's the ranking. The tool tells you which MCP server, which skill, or which chunk of CLAUDE.md costs the most per call, with a concrete trim suggestion ("disable MCP server X, save ~N tokens per call").

Calibrate it to your setup

The 51K default comes from the published measurement, but your preamble is yours. Feed it real data:

# measure your actual CLAUDE.md and skills instead of heuristics
subagent-tax --claude-md ./CLAUDE.md --skills-dir ~/.claude/skills

# import per-server schema sizes from mcp-tax (my earlier tool)
mcp-tax audit --json > /tmp/mcp.json
subagent-tax --mcp-tax-report /tmp/mcp.json

# or override any component directly
subagent-tax --set system_prompt=15000 --price-input 3.00 --format json
Enter fullscreen mode Exit fullscreen mode

Where it fits: mcp-tax vs subagent-tax

I built mcp-tax earlier to audit the total size of MCP server schemas. subagent-tax is the other half: the repeat cost — how many tokens get re-sent on every single subagent call because the preamble is fixed. They compose: mcp-tax measures it once, subagent-tax multiplies it by your call count.

Honest limitations

No exact token counts without real API payloads — text is estimated at ~4 chars/token, and dollar amounts use a pricing input you should verify against current published rates. It only sees local JSONL transcripts, and "waste" is a simplification: some preamble is working context, not pure waste. The README documents all of this.

Links: GitHub · pip install subagent-tax · MIT licensed, stdlib-only, zero dependencies.

If you run it, I'd genuinely like to know: what's your most expensive preamble component?

Top comments (0)