DEV Community

hao li
hao li

Posted on Originally published at github.com

Your MCP servers are eating your context window. I built a CLI to audit the tax

Every MCP server you connect injects all of its tool schemas into every request. Claude Code loads them at session start, and there is no UI to temporarily disable a server you don't need right now. People have measured 41k tokens of pure schema; one estimate says 6 mid-size servers eat 10–15% of the window before the conversation even starts.

Nobody measures this per server. So I built a tiny CLI that does.

mcp-tax

mcp-tax audits the context tax of your MCP servers, then lets you launch Claude Code with the expensive ones switched off — per session, without touching your real config.

Zero dependencies. Python standard library only.

pip install mcp-tax
Enter fullscreen mode Exit fullscreen mode

See what you have configured (reads ~/.claude.json plus ./.mcp.json when present):

$ mcp-tax list
github
    npx -y @modelcontextprotocol/server-github
postgres  [off]
    uvx mcp-server-postgres --db-url ...
2 server(s), 1 disabled
Enter fullscreen mode Exit fullscreen mode

Measure the tax — it handshakes each server over stdio (initialize, then tools/list), counts tools, and estimates tokens:

$ mcp-tax audit
server    tools  schema chars  est. tokens
------  -------  ------------  -----------
github       51       118,203        29,551
postgres [off]    9        12,440         3,110
------  -------  ------------  -----------
TOTAL        60       130,643        32,661
~16.3% of a 200k context window (est. tokens = schema chars / 4)
Enter fullscreen mode Exit fullscreen mode

Toggle servers off/on (persisted in ~/.config/mcp-tax/disabled.json):

$ mcp-tax off postgres
postgres disabled (affects `mcp-tax run`)
$ mcp-tax on postgres
postgres enabled (affects `mcp-tax run`)
Enter fullscreen mode Exit fullscreen mode

Launch Claude Code without the disabled servers:

$ mcp-tax run -- -p "summarize this repo"
Enter fullscreen mode Exit fullscreen mode

This writes a filtered {"mcpServers": ...} config (disabled servers removed) and execs claude --mcp-config <file> with your args forwarded. Your real ~/.claude.json is never modified.

How the token estimate works

For each server, mcp-tax JSON-encodes the full tools/list result (compact, no whitespace) and counts characters. Estimated tokens = round(chars / 4) — the Anthropic tokenizer's ~4-chars-per-token rule of thumb on English/JSON text.

Deliberately crude. Real counts vary with the tokenizer and how Claude Code wraps schemas, so treat it as an order-of-magnitude gauge: good enough to answer "which server is eating my window?", not a billing meter.

Honest limitations

  • stdio servers only. SSE / streamable HTTP servers aren't audited.
  • Audit uses select(2) on the server's stdout pipe — fine on Linux/macOS, not on Windows.
  • The estimate ignores runtime behavior: a 2-tool server can still be expensive if its tool results are huge. This measures schema cost only.
  • The --mcp-config flag for run comes from Claude Code's documented CLI options; it couldn't be verified on the machine where this was built (no Claude Code CLI there). If the flag name ever changes, run prints the filtered config path so you can pass it manually.

Links

7 smoke tests pass, including a fake stdio MCP server that speaks the same newline-delimited JSON-RPC 2.0 framing as real servers — so the audit math is tested against a realistic handshake, including stdout noise and timeout cases.

If you run it against your own setup, I'd genuinely like to know: which server turned out to be your biggest tax?

Top comments (0)