DEV Community

hao li
hao li

Posted on

subagent-tax 0.3: the preamble re-enters every turn — and every subagent spawn re-pays it

Two readers argued with my billing model in the comments, and they were right. So subagent-tax 0.3 rebuilds it — borrowing the corrected framework from mcp-tax 0.3.

The correction: the preamble re-enters the window on every turn

The old subagent-tax billed the fixed preamble once per Task() call: 51K tokens in, done. That's wrong twice over.

First, the per-turn re-entry correction from mcp-tax 0.3 (thanks to MCPulse for pressing on this): tool schemas and the rest of the preamble re-enter the model context on every API call, not once per session. A 51K preamble in a 3-turn subagent isn't a 51K tax — it's a 153K tax that has to fit inside the window three times over.

Second, the subagent multiplier: every Task() call spawns a brand-new context, so the full preamble is re-paid on spawn. It's a true amplifier. Measured in the wild by dev.to reader aidiveyt: 2,631 subagent runs ate 48.1% of all tokens.

The per-turn re-entry ledger

The report now ends with a ledger — one billing line per subagent_type, with an explicit subagent_multiplier column showing exactly how many fresh contexts each type spawned:

Per-turn re-entry ledger (corrected billing model, same framework as mcp-tax 0.3)
  turns per Task call: 3 (heuristic; --turns-per-task) | spawn overhead: 2,000 tok/call | cache hit rate: 0% (spawn turn always cold)

  subagent_type     calls multiplier per-call tok   window tok billable tok
  Explore              90         90      155,000   13,950,000   13,950,000
  Plan                 31         31      155,000    4,805,000    4,805,000

  Old model (v0.2, spawn-only) counted 6,732,000 tokens; corrected ledger counts
  20,460,000 — the old model undercounted by a factor of 3.0x.
Enter fullscreen mode Exit fullscreen mode

Three knobs tune the model: --turns-per-task (heuristic turns inside each subagent — parent transcripts don't show internal turns), --spawn-overhead (the Task payload + result handling per call), and --cache-hit-rate. On caching: a fresh subagent starts with a cold cache, so the spawn turn is always billed at full price — a cache TTL miss on every spawn. And as in mcp-tax 0.3, caching only lowers the price, never the window occupancy: billable_tokens and cumulative_window_tokens stay separate columns.

Same framework as mcp-tax, different side of the bill

mcp-tax audits how much your MCP server schemas cost and bills them per turn with a per-server ledger. subagent-tax shares that exact billing framework but focuses on the subagent side: the multiplier (each Task() re-pays the full preamble on spawn), the per-call spawn overhead, and the cache TTL miss at spawn — none of which exist in the MCP-server world. The old spawn-only number stays in the JSON output as old_model_tokens so you can see exactly how much it undercounted.

pip install subagent-tax
subagent-tax
Enter fullscreen mode Exit fullscreen mode

61 tests green (48 existing + 13 new), stdlib only, MIT.

Top comments (0)