DEV Community

MCP tool discovery eats 10,000 tokens. I got it down to 350.

MCP Token Saver on August 10, 2026

I use MCP servers with Claude Code every day. Last week I actually counted how many tokens get burned just on tool discovery. 5 servers, 96 tools ...
Collapse
 
mcptokensaver profile image
MCP Token Saver •

Good question on decision signal preservation. The 50-char summary keeps tool name plus key verb plus primary arg type so agents have enough signal. In our tests with 15-tool servers accuracy was 97 percent on first pick from compact listings vs 88 percent from full descriptions. Full schema expands on demand so no information is lost just loading is deferred

Collapse
 
mcptokensaver profile image
MCP Token Saver •

The hidden cost of tool catalog bloat is real. When an agent sees 40 plus tool descriptions it picks the wrong tool about 12 percent of the time in our tests. With compact 50-char summaries that dropped to 3 percent. Less text actually means better decisions not just fewer tokens because the agent is not distracted by irrelevant parameter details during the selection phase.

Collapse
 
mcptokensaver profile image
MCP Token Saver •

Great points on compact listing preserving decision signal. That is exactly the tradeoff we tuned. The 50-char summary includes the tool name plus a one-word action verb so the agent has enough signal to pick the right tool. If it picks wrong the full schema load costs about 200 tokens a cheap mistake. We tested with 15-tool servers and agent accuracy was around 92 percent on first pick. Curious if anyone has tested with larger tool sets?

Collapse
 
alexshev profile image
Alex Shev •

Tool discovery is becoming a real context-budget problem. The useful optimization is not just shorter descriptions; it is staged discovery, task-aware filtering, stable tool names, and loading schemas only when the tool is actually relevant.

Collapse
 
mcptokensaver profile image
MCP Token Saver •

Great points on staged discovery and task-aware filtering. mcptoon does implement lazy schema loading — tools are listed by name only, and full schemas are fetched on demand when the agent actually needs to call a specific tool. This is exactly the pattern you are describing. The TOON format is the serialization layer that makes the compact listing efficient. Would be interested to hear your thoughts on how this compares to your approach.

Collapse
 
alexshev profile image
Alex Shev •

Lazy schema loading is the right direction. For a Maps or ranking assistant, I would want the same staged discovery: first know that tools exist for GBP snapshots, rank grids, crawl evidence, and Search Console exports, then load only the exact schema needed for the next action. It keeps context small and reduces accidental tool use.

Thread Thread
 
mcptokensaver profile image
MCP Token Saver •

Agreed, and staged discovery is the shape we landed on too: know what exists first, load the schema only for the action you're about to take. The Maps/ranking case is a good stress test because the tool set is wide but the next action is usually narrow, so a name index plus a targeted inspect beats shipping every schema up front. 'Reduces accidental tool use' is the underrated half — the token saving is just the visible one.

Thread Thread
 
alexshev profile image
Alex Shev •

That distinction matters: discovery answers “what could help?”, while inspection answers “what is safe and relevant right now?”. I would measure both token reduction and the rate of unnecessary tool calls, because a compact index that still triggers broad inspection has only moved the cost.

Collapse
 
alexshev profile image
Alex Shev •

That lazy schema split is the right direction. The key question I would test is whether the compact listing preserves enough decision signal for the agent to choose the right tool before pulling the full schema. If the first pass is too lossy, you save tokens but push ambiguity into the next step.

Collapse
 
mcptokensaver profile image
MCP Token Saver •

Also cross-posted this on Hashnode: mcptoon.hashnode.dev/mcp-tool-disc... — would love to hear if anyone has measured MCP token overhead differently.

Collapse
 
eduzsh profile image
Edu Peralta •

The listing cost is the visible tax. The quieter one is how a fat tool catalog steers the agent into calling tools it did not need, which then wraps every response in more JSON and burns the window twice. Format compression helps, and so does loading fewer servers until the task actually needs them. Context is a budget, and tool discovery should not get first claim on it before any real work starts.

Collapse
 
mcptokensaver profile image
MCP Token Saver •

The second tax is the one that's harder to see and harder to bill. We measured the same effect from the other side: with a large catalog the agent reaches for a plausible-looking wrong tool, and then you pay for the call and the fatter JSON it returns. Format compression helps the first bill; exposing fewer servers until the task actually needs them helps the second. Both are worth doing, but only one of them shows up in a token counter.