DEV Community

Rudratosh Shastri
Rudratosh Shastri

Posted on

Your GitHub MCP server costs 55,000 tokens before your agent reads a single word.

Here's a cost most people connecting MCP servers never see on a dashboard, because it's spent before their agent says anything.

Connect the GitHub MCP server and it exposes ~93 tools. The client serializes every one of them — name, description, full JSON parameter schema, enum values, example text — into the model's context. That's about 55,000 tokens. Before you type a word.

Now connect Slack and Sentry too. Community measurements put a three-server setup around 143,000 tokens of a 200k window — spent on tool definitions the agent may never call. You paid for a 200k context and got ~57k to actually work in.

Why MCP is expensive by design

It's not a bug in one server. It's the shape of the protocol:

  • Every tool's full schema loads up front. MCP hands the model the complete menu so it can reason about what's available. Great for capability, brutal for tokens.
  • It re-enters context every single turn. This isn't a one-time load. Those 55k tokens ride along on each message in the conversation. A 20-turn session pays the tax 20 times.
  • One tool ≈ ~1,000 tokens. Field descriptions (~400), type definitions (~300), nested parameter structure (~300) — for a tool a human would describe in two sentences. Multiply by 93.

So the "hidden context tax" isn't hidden because it's sneaky. It's hidden because it's charged in the one place nobody's looking: the system prompt, on every turn.

The CLI comparison that stings

The reason "I deleted my MCP servers and my agent got faster" keeps trending is that a CLI + on-demand tools flips the model. Instead of pre-loading 93 schemas, the agent runs gh (or a 40-line script) and learns the one command it needs when it needs it. Same result, a tiny fraction of the tokens, because the tool definition never sits in context bloating every turn.

Fewer tokens per turn also means: cheaper, faster (less to process), and often smarter — a model drowning in 55k tokens of irrelevant schemas has less attention for your actual task.

But MCP isn't dead — three cases where it wins

The "MCP sucks" takes overcorrect. The protocol earns its overhead when you need what it uniquely provides:

  1. Auth and credential isolation. MCP servers hold the token so your agent (and your prompt) never see it. Rolling your own CLI wrapper means the secret lives closer to the model. For anything sensitive, that isolation is worth tokens.
  2. Multi-user / multi-tenant. One governed server serving many agents with per-user scoping beats every agent shelling out with its own creds. Governance centralizes; scripts scatter.
  3. Discovery across a big, changing tool surface. When you genuinely don't know which of 200 tools you'll need, structured discovery beats hard-coding — if you pair it with lazy loading.

And that last word matters: MCP Tool Search (protocol-level, since January) defers loading schemas until they exceed ~10% of your context, then discovers tools on demand. If your client supports it, turn it on — it's the difference between paying the 55k tax and paying for what you use.

The rule of thumb

Count the tokens your tools cost before your agent runs. If it's a big slice of your window and you only use a handful of tools, you're paying MCP's convenience tax for tools you never call. Use a CLI, or turn on lazy tool loading. Keep MCP for auth, multi-user, and genuine large-surface discovery.

MCP is a real standard solving real problems. It's just not free, and the price is charged in the currency your agent needs most — context.


Have you actually measured what your connected MCP servers cost per turn? Run the count and drop the number — I'm curious how bloated the median setup really is. 👇

I write about building with AI and the honest costs of it. Follow me here if that's your lane. 👋

Top comments (0)