Your tool is registered. It appears in tools/list. The model never calls it, or calls a different one instead.
There are six reasons this happens. They are not equally likely, and the instinct — assume the description needs to be better — is right about a third of the time.
Here they are in order of how often they turn up across 4,749 public MCP servers.
1. Another tool is winning a competition you cannot see
Most likely. 17.7% of public tools have a description containing no word that distinguishes them from a sibling. On servers with more than sixty tools it is 32.4%.
The tell is that a neighbouring tool gets called instead of yours, consistently. Not randomly — the same substitution every time.
How to check. Print your whole tool list and read your tool's description next to the one that keeps getting called. If you can swap the two descriptions and both still read correctly, the model has nothing to choose on.
The fix. Lead with the difference, not the category. "Search orders" describes four of your tools. "Search orders by customer, including cancelled ones" describes one. Naming the return shape is the single most distinguishing sentence available and almost nobody writes it.
2. The parameters are unfillable
Second most likely, and the one people check last. 21.5% of parameters across the corpus have no description at all.
A model that cannot work out what to put in a required field will avoid the tool rather than guess. This looks identical to a description problem from outside, and it is not.
The clearest version:
fetch id, document_id (neither described)
A model has to pick one and nothing tells it which, or whether they differ. So it picks a different tool.
How to check. Cover the parameter names and look only at types and descriptions. If you could not supply a value, neither can the model.
Watch for the internal identifier specifically. A tool requiring a warehouse_id is unreachable from a conversation where nobody has ever seen a warehouse ID. That is not a description problem — it is a tool that needs a lookup step before it, or a name-based alternative.
3. The name is doing all the work and losing
When descriptions collide, the name is what decides — in 89.6% of collision cases. So a colliding description plus a generic name is the case where nothing is choosing at all.
get leads 15.6% of all public tool names. If your tool is get_thing and a sibling is fetch_thing, those are synonyms in English and the model is guessing.
The fix. One verb per operation, across the whole server. Never ship two words from the same row:
-
get/fetch/read/retrieve -
list/getAll -
search/query/find -
delete/remove
4. Your tool list is too long to hold
Schema goes into the context window on every connection. The median server spends about 1,251 tokens; the 90th percentile spends 7,992. The largest in the corpus spends around 216,000 — which does not fit a 200K window at all.
The failure here is not that your tool is badly described. It is that it is the fortieth candidate in a list the model is skimming, and attention is finite.
How to check. Count your tools. Under fifteen, this is not your problem. Past thirty, it is probably your main one.
5. A closed set was written in prose
The model sends "Direct" where you accept "direct", your server rejects it, and from the model's side that reads as a broken tool rather than a fixable mistake — so it stops trying.
857 parameters in the corpus name their valid values in the description and leave the schema as an open string. Moving them into an enum is one line.
The tell for this one is distinctive: the tool does get called, once, and then never again in that session.
6. It is actually unreachable
Two rarer causes worth ruling out.
A duplicate tool name. Only 0.19% of servers — nine in the whole corpus — but when it happens, one of the two tools cannot be called at all, and the symptom points somewhere else entirely. Almost always a generated tool list. One assertion in CI catches it forever.
The tool is not in the response you think it is. Servers that build their list conditionally can emit something different from what the code appears to register. Read the real tools/list output rather than the registration code.
The order to actually work through
- Read your whole tool list as one block. Which tool would you pick? — catches causes 1 and 3.
- Cover the parameter names, read only types and descriptions — catches cause 2.
- Count your tools — catches cause 4.
- Grep descriptions for
valid values/options:— catches cause 5. - Assert unique names in CI — catches cause 6.
That is fifteen minutes and it covers all six.
What none of this can tell you
Every cause above is a property of your schema, and this is the honest limit: we have never observed a model choosing a tool on any of these servers. The rates say how often each pattern exists, not how often it costs you a call.
It is entirely possible your tool is skipped for a reason no schema shows — the client truncates long lists, the model was mid-task, the user phrased something unusually. Distinguishing those from the six above needs data from inside your server: which tools get called, which get retried, which return empty. A schema check is what you can do before you have any of that.
The schema checker tests causes 1, 2, 3, 5 and 6 on a paste of your tools/list, scored against servers your own size. It names the sibling each colliding tool is hardest to tell apart from — which for cause 1 is usually the whole answer. Runs in your browser, nothing uploaded.
Originally published at getmcpulse.com.
Top comments (6)
Cause 4 is the one that took me longest to accept was the actual problem. Under fifteen tools you can mostly brute-force it and lean on description quality. Past thirty I've had to go to dynamic registration: a lightweight classifier reads the opening message and loads only the tools plausibly relevant to that intent. The catch is the classifier can't use the full tool set to decide what to include, so you've quietly turned a routing problem into a retrieval problem.
"You've turned a routing problem into a retrieval problem" is the sentence I'll be repeating, because it names the cost precisely. The routing failure is visible — wrong tool called, and you can see which. The retrieval failure is silent: the right tool was never a candidate, nothing was called, and nothing in your logs distinguishes that from nobody having asked.
Which makes cause 1 nastier under dynamic registration rather than better. If your classifier reads the opening message and your tool descriptions are what it reads, two tools a model can't tell apart are also two tools a classifier can't tell apart — you've moved the same ambiguity one layer earlier, where it's harder to observe. The description quality problem doesn't go away, it just stops producing a wrong call you can point at.
It also splits my measurement in a way I hadn't accounted for. Schema cost per connection assumes the whole list ships every time, which for you it doesn't — your median payload is whatever the classifier selected. And dead tools becomes ambiguous: never called might mean never selected by the model, or never loaded by the classifier, and those have completely different fixes.
The thing I'd want to know, if you ever get a way to measure it: how often the classifier excludes a tool that would have been the right answer. That's the retrieval failure rate, and it's the number that tells you whether dynamic registration is net positive or just relocating the problem. I don't know how you'd get it without a ground truth set of intent-to-tool pairs.
Fifteen and thirty as the two thresholds matches the corpus, incidentally — collisions run at 4.3% up to fifteen tools and 16.3% by sixty.
The strongest description is an executable promise: the input it accepts, the state it reads or changes, the output shape, and the conditions that make it the right choice over its siblings. I’d test tool selection with contrastive prompts that make neighbouring tools plausible, then use the mistaken selection as a regression case for the schema and descriptions.
"Executable promise" is a better spec than anything in the post, and the fourth item is the one nobody writes. Input, state, output shape — those show up often enough. The conditions that make this the right choice over its siblings almost never do, because writing it requires holding the whole tool list in your head at once, and people write descriptions one tool at a time in the file where they registered it.
Which is probably the mechanism behind the 17.7% rather than carelessness. Every description in that Gmail example is correct about its own tool. None of them was written with the others open.
Contrastive prompts as the test is the part I want to steal. My static check finds pairs a model might confuse; it can't tell which confusions are real. Deliberately writing a prompt that makes two neighbouring tools plausible and seeing which one gets picked converts a suspicion into a fact, and it needs no traffic — you can do it before shipping.
The regression case is the bit that makes it durable. Each mistaken selection becomes a fixture: prompt, expected tool, actual tool. Then a description rewrite is testable rather than hopeful, and the next person who "improves" the wording finds out immediately if they've broken the distinction. That's the thing missing from every schema-linting approach including mine — a linter tells you a description looks risky, a regression suite tells you whether the fix worked.
It also generalises past MCP. Any tool-using agent has this problem, and the artefact is just a labelled set of intent-to-tool pairs. Which is the thing a schema analysis can never produce on its own, because it's the ground truth I'd be validating against.
There is one more failure class between schema and reachability: the tool is callable but policy-ineligible in the current context. A model can correctly avoid a well-described tool because the user did not authorize its side effect, the required identity is unverified, or the destination is ambiguous. Instrumenting only selection rate would mislabel that restraint as poor discoverability. I would log candidate, eligible, selected, attempted, and succeeded as separate stages.
That's a real gap and it's worse than a missing category — it's a sign error. Every cause in the post is something to fix. Correct restraint measured the same way reads as a problem and produces advice that makes the server less safe: rewrite the description, make it more findable, get the model to call the thing it was right to avoid.
The five stages are the fix, and the gaps between them carry different meanings. candidate → eligible is a policy filter and no schema change touches it. eligible → selected is where everything in the post lives. selected → attempted catches arguments the model couldn't fill. attempted → succeeded is the handler.
Collapsing those into one selection rate makes them indistinguishable, which is what I'm currently doing.
The honest problem is that eligibility is invisible from where I sit. An SDK inside the server sees the call arrive or not arrive; it can't see the model declining because the user never authorised a side effect, or because the identity behind the request is unverified. That decision happens in the client, and the protocol reports nothing back.
Which means candidate → eligible is either author-declared — the tool itself knows it requires an authorised identity or a disambiguated destination, and can say so — or it stays unobservable and I should say that rather than folding it into selection.
Declared is probably right, and it has a useful side effect: a tool that declares a policy precondition can be excluded from the discoverability numbers entirely, so a well-designed dangerous tool stops being penalised for not getting called.
Adding this. It's the first thing anyone has raised where the existing metric doesn't just miss something but actively points the wrong way.