Two readers argued with my billing model in the comments, and they were right. So subagent-tax 0.3 rebuilds it — borrowing the corrected framework from mcp-tax 0.3.
The correction: the preamble re-enters the window on every turn
The old subagent-tax billed the fixed preamble once per Task() call: 51K tokens in, done. That's wrong twice over.
First, the per-turn re-entry correction from mcp-tax 0.3 (thanks to MCPulse for pressing on this): tool schemas and the rest of the preamble re-enter the model context on every API call, not once per session. A 51K preamble in a 3-turn subagent isn't a 51K tax — it's a 153K tax that has to fit inside the window three times over.
Second, the subagent multiplier: every Task() call spawns a brand-new context, so the full preamble is re-paid on spawn. It's a true amplifier. Measured in the wild by dev.to reader aidiveyt: 2,631 subagent runs ate 48.1% of all tokens.
The per-turn re-entry ledger
The report now ends with a ledger — one billing line per subagent_type, with an explicit subagent_multiplier column showing exactly how many fresh contexts each type spawned:
Per-turn re-entry ledger (corrected billing model, same framework as mcp-tax 0.3)
turns per Task call: 3 (heuristic; --turns-per-task) | spawn overhead: 2,000 tok/call | cache hit rate: 0% (spawn turn always cold)
subagent_type calls multiplier per-call tok window tok billable tok
Explore 90 90 155,000 13,950,000 13,950,000
Plan 31 31 155,000 4,805,000 4,805,000
Old model (v0.2, spawn-only) counted 6,732,000 tokens; corrected ledger counts
20,460,000 — the old model undercounted by a factor of 3.0x.
Three knobs tune the model: --turns-per-task (heuristic turns inside each subagent — parent transcripts don't show internal turns), --spawn-overhead (the Task payload + result handling per call), and --cache-hit-rate. On caching: a fresh subagent starts with a cold cache, so the spawn turn is always billed at full price — a cache TTL miss on every spawn. And as in mcp-tax 0.3, caching only lowers the price, never the window occupancy: billable_tokens and cumulative_window_tokens stay separate columns.
Same framework as mcp-tax, different side of the bill
mcp-tax audits how much your MCP server schemas cost and bills them per turn with a per-server ledger. subagent-tax shares that exact billing framework but focuses on the subagent side: the multiplier (each Task() re-pays the full preamble on spawn), the per-call spawn overhead, and the cache TTL miss at spawn — none of which exist in the MCP-server world. The old spawn-only number stays in the JSON output as old_model_tokens so you can see exactly how much it undercounted.
pip install subagent-tax
subagent-tax
61 tests green (48 existing + 13 new), stdlib only, MIT.
Top comments (0)