Everyone's started auditing the first bill. Almost nobody has found the second one.
Here's what I mean. Open your agent's context and count what's already loaded before you type a single character. There are two line items, and they come from two different places:
- MCP tool schemas. Every connected server ships the full JSON description of every tool it exposes.
-
Agent skills. Every
SKILL.mdin your skills directory, loaded whole.
I measured both with tiktoken, OpenAI's official tokenizer, using the cl100k_base encoding. The tool side is bad. The skill side is worse. And the same idea fixes both.
Bill #1: the tool schemas
Connect enough MCP servers and the tool descriptions alone eat a large slice of your context window. My reference measurement is a catalog of 255 tools. Rendered as full JSON schemas, that's 71,929 tokens.
And it isn't a one-time cost. The manifest gets re-sent every turn. Ask a dozen questions in one session and you've paid that bill a dozen times.
The reason is structural. Choosing a tool only needs its name. You don't need the full schema to decide whether to call something. But the default plumbing ships the schema anyway, up front, for every tool, on every request. So you're paying for documentation you'll mostly never open.
Bill #2: the skills, the one nobody audits
The skill side is bigger, and it's bigger because of what a skill is. A tool is a name plus a schema. A skill is an entire document: instructions, examples, a workflow, a checklist. When the catalog gets exposed as "read the SKILL.md you need," the naive implementation reads the catalog.
On this machine the skills directory holds 418 skills. Concatenated whole, every SKILL.md is 1,109,242 tokens.
That can't fit in a context window, so most setups do the thing that looks careful and isn't. They load part of it, whatever the agent guesses it needs, and pay a few thousand tokens per skill, every turn, forever. You don't notice, because the bill is spread across many small, reasonable-looking decisions.
The same fix lands on both bills
The insight is almost boring once you see it. Choosing something only needs its name.
- You don't need a tool's full schema to decide whether to call it.
- You don't need a skill's full text to decide whether to read it.
So expose the names first, and fetch the bodies on demand.
Tools. 255 tools as a names-only index come to 581 tokens. That's 99.2% off the original 71,929. The full schema arrives only when you actually call a tool.
Skills. One resident pointer, 39 tokens, replaces the whole catalog. A single lookup returns the handful of matching skills, and that lookup costs 501 tokens. You read the full SKILL.md files only for the winners. Same idea, other bill.
The pointer's entire job is one sentence: don't load the catalog, ask for a match, then read only those files.
I run this as mcptoon, a zero-dependency CLI. It's pure Python standard library, 227KB, and it keeps every server and every skill configured and addressable while keeping their full text out of your context window. mcptoon manifest gives you the names-only view of your tools. mcptoon skills resolve "<what you're trying to do>" returns the best-matching skills as JSON. Nothing's pre-installed, and nothing loads until you choose it.
Measure your own number first
Don't take my catalog as yours. The tax scales with how much you've wired up, and so does the fix.
Count what you actually carry. Load your real tool list and your real skills directory, run them through tiktoken, and compare the whole thing against a names-only view. If the gap is large, you've got a bill worth cutting. If the gap is small, you don't.
That measurement is the whole point. It's also the step almost everyone skips, which is why the second bill goes unnoticed for so long.
The honest boundary
This pays off when you have many tools and skills. If you run a couple of servers and a dozen skills, none of this matters much. The tax is real but small, and a gateway is one more thing to keep running. Don't optimize for the sake of optimizing.
My rule of thumb: audit first, and only reach for a name-index layer when the measurement says your context is actually crowded.
The part worth arguing about
The tool-schema side is now a known problem. Anthropic's own engineering write-up makes the case, and a Firecrawl benchmark measured the same task at 1,365 tokens through a CLI against 44,026 tokens through MCP, a 32× gap. That conversation is happening.
The skill side isn't. And it's the larger of the two, because skills are documents, not schemas.
So here's the question I'd like answered in the comments: when your agent loads a skill, how much of it is the pointer, and how much is the whole file? If you've measured that number, I'd like to see it, because I suspect most people are paying the second bill without ever seeing the invoice.
mcptoon is open source. Search for it by name. The numbers above are tiktoken cl100k_base measurements on the tool and skill catalogs described. The skill figure reflects a **418-skill directory on one machine, so treat it as a shape, not a constant. Your own mcptoon status prints your number.
Top comments (0)