You wired up a great MCP server. It lists 42 tools, each with a detailed schema, and the agent now has full access to your internal API, your database, your deploy pipeline.
Three weeks later you notice the agent is forgetting earlier instructions, hallucinating tool names, and hitting token limits halfway through a task. You assume your model isn't smart enough. You switch models. The problem gets slightly worse.
The issue isn't the model. It's that your MCP server is using the agent's context window as a dumpster.
The eager-load tax
Most MCP servers expose every tool at session start. The model receives the full tool catalog, all descriptions, all parameter schemas, and the resource list before it has read the first line of your codebase. On a 200k-token context window, a well-built internal API with 30 endpoints can consume 8-15% of that budget before the agent has done anything.
The agent doesn't need 30 tools. It needs 2. It needs the one that reads the file it's about to edit, and the one that runs the test it just wrote. The other 28 are dead weight that crowd out the actual task.
What actually helps
Split by capability, not by resource. Don't ship one monolithic server that knows everything. Build narrow servers: one for file operations, one for the API you're testing against, one for deployment. The agent loads the one it needs for the current task. A 6-tool server costs a fraction of a 42-tool one, and the agent's success rate on that narrow task goes up because it isn't choosing between similar-sounding tools.
Return short errors, not long payloads. When a tool fails, the agent will read the error and retry. A 400-line stack trace is worse than useless: it floods the window, the model half-reads it, and retries with the same wrong assumption. Return a 2-line error code, a pointer to logs, and a hint about which parameter was wrong. The agent can ask for more detail if it needs it.
Make the schema do the filtering. Tool descriptions should tell the model when not to call this. "Use this only for production deploys, never for staging." "Prefer this over the legacy version." The model uses these disambiguators to pick the right tool, and a well-narrowed schema means fewer wrong calls and fewer retries.
Version your tools and deprecate loudly. If you have v1 and v2 of the same endpoint, mark v1 as deprecated in the description. Agents will still call it, but less often, and you can remove it once telemetry shows zero calls. Leaving both alive "just in case" is how you end up with 42 tools.
The spec is the shortcut
Here's the part that surprised me: the highest-leverage move is to generate the MCP server from an OpenAPI spec, not hand-write it. A spec-driven generator produces a tight, consistent tool list — one tool per operation, predictable parameter names, no invented helpers, no duplicate endpoints. You can prune the spec to the paths the agent actually needs, and the server follows.
This is the exact problem we kept hitting while building Powerduck: we wanted the local-first OpenAPI editor to also be the thing that generates a minimal MCP server from the paths you care about right now, without the agent having to wade through 80 endpoints it will never touch. The spec stays the source of truth, the generated server stays small, and the agent keeps its context window for the actual problem.
The boring test
Before you ship your next MCP server, do one thing: open a fresh agent session and count how many tokens the tool catalog alone consumes. If it's over 5% of the context window for a task that needs 3 tools, you're doing it wrong. Narrow the server, shorten the errors, and let the model focus on the work instead of reading a menu.
Top comments (2)
Agree on the catalog tax, and one thing is shifting under it: some clients now defer tool schemas. Claude Code, for example, can keep MCP tools as names only and pull a full schema when the model searches for it. On those clients a 42-tool server costs much less up front. Your other points still apply, because a deferred tool is found by its name and description, so vague or overlapping descriptions now hurt at search time instead.
The bigger leak I see is tool results, not the catalog. I work on Auten, a computer-use MCP server, and one screenshot returned as an image can cost more than the whole tool list. Every "look at the screen" turn pays it again. What helped us: return text by default (the element list and the visible text with coordinates) and send an image only when the agent asks for one. Same idea as your short-error rule, applied to every successful call too.
So I'd add a second boring test next to yours: run a typical 20-step task and compare tokens from the catalog with tokens from tool results. For anything that reads screens, pages or files, results usually win by a wide margin.
Counting the tokens of the tool catalogue before shipping is a good habit. One thing I'd add from the other side: the saving is in fewer tools, not thinner ones. I maintain a small three-tool server, and a directory's quality check marked it down because the parameters had no descriptions. Adding one sentence per parameter ("every word must match, case-insensitive", plus which fields are searched) was enough for the model to send short, specific queries and to retry with fewer words when it gets nothing. A wrong call and a retry cost far more context than that sentence.
On generating from OpenAPI: one tool per operation gets you back to 80 tools quickly on a real spec, so the allow-list by tag or path at generation time matters as much as the generator.