DEV Community

Hiroki Nomura
Hiroki Nomura

Posted on Fully Autonomous

We checked 46 popular MCP servers on Smithery: 7 in 10 tools never say when to use them

We took 46 of the most-used MCP servers on Smithery and checked the descriptions of all 1,118 of their tools. 769 tools (69%) never say when to use them. 507 (45%) never say in their description what they return. And 15 servers have tools whose descriptions are so alike that an agent can mix them up: 90 such pairs in all.

The 69% is exactly what we found last time, when we checked 8 public servers with 126 tools and 87 of them lacked it. Nine times the sample, the same share.

The checks were done with Forecall, a static linter for MCP tool descriptions that I build.

How we checked: the top 50 servers on Smithery, judged by their text only

On October 8, 2026, we took the first 500 servers in Smithery's public registry (in the order its API returns), removed duplicates, picked the 50 with the most uses, and collected their tool definitions. We scored them with Forecall's CLI (forecall 0.3.2). Four of the 50 have more than 200 tools, the CLI's limit, so we left them out: 46 servers and 1,118 tools remain.

One caveat. The tool definitions Smithery's API returns leave out the input schemas' required lists and any output schemas, so Forecall's overall scores come out lower than they should. This article does not use the overall scores. It counts only the findings that depend on the description text alone: whether a tool says when to use it, whether its description says what it returns, whether the description is too short, and pairs of tools whose descriptions are alike.

7 in 10 tools never say when to use them

Finding Tools Share
No "when to use" 769 69%
Description doesn't say what it returns 507 45%
Too short, or restates the name 101 9%

By server it is starker. In 16 of the 46 servers, not one tool says when to use it. Count the servers where 80% or more lack it, and you have half of them: 23.

"When to use" is what decides between lookalike tools. Without it, a model picks by the tools' names and the feel of their descriptions.

The same notice on every tool makes the tools indistinguishable

PubMed has the most lookalike pairs. It has only 7 tools, yet all 21 pairs are flagged as easy to mix up. Among the 12 most similar pairs, the descriptions share 73% to 94% of their words.

The cause is a long notice pasted into all 7 descriptions:

IMPORTANT - PubMed Database Scope: This server provides access to PubMed, which ONLY indexes biomedical and life sciences literature including: …

What sets each tool apart, its first sentence ("search articles", "get metadata", "find related articles" and so on), is buried in the shared paragraph.

A note that applies to the whole server belongs once in the server's instructions, which the client receives at initialization, not in every tool's description. A tool's description should say what only that tool does.

"Send a reply" and "draft a reply" are written almost the same way

Gmail has 7 lookalike pairs. The closest two share 92% of their words:

Tool Start of its description
Gmail_ReplyToEmail Send a reply to an email message, optionally with one or more file attachments.
Gmail_WriteDraftReplyEmail Compose a draft reply to an email message, optionally with one or more file attachments.

After that, both go on with the same paragraph about attaching files. The difference is a few words: "Send" against "Compose a draft".

Mix these two up, and an agent meant to write a draft sends the email. Tools that do something you can't undo, above all, should say in one sentence how to choose: "To prepare a reply without sending it, use Gmail_WriteDraftReplyEmail."

A good example: Brave Search says how its tools connect

Brave Search's descriptions, by contrast, spell out how the tools relate. Here is brave_web_search:

Performs web searches using the Brave Search API and returns comprehensive search results with rich metadata. To chain into local-POI enrichment, pass result_filter=locations and feed the resulting locations.results[].id values into brave_local_search.

It says what comes back and which value to pass to which tool next, so an agent can use several tools in sequence. All 8 of Brave Search's tools say in their descriptions what they return, and it has no lookalike pairs.

How to fix it: one sentence each for three things

From these results, this is the order to fix descriptions in:

  1. Say when to use the tool. "Use this when …", and if there is a lookalike, "for …, use X instead."
  2. Move shared notes to the server's instructions. The same paragraph on every tool buries what tells them apart.
  3. Put the difference between lookalikes into words. Send or draft? Delete or remove? Say it plainly, especially for actions you can't undo.

To check your own server, run these two commands. The first gets the tools/list from your server, the second scores it. Both run on your machine and send the results nowhere.

npx forecall dump -o tools.json -- <command that starts your server>
npx forecall lint tools.json
Enter fullscreen mode Exit fullscreen mode

For a Streamable HTTP server, pass its URL instead: npx forecall dump https://example.com/mcp > tools.json. The CLI is open source: https://github.com/forecall/forecall-cli

The weakness doesn't shrink with scale

With 8 servers or 46, 7 in 10 tools never said when to use them, popular servers included. The fix is small: one sentence on when to use the tool and one on how it differs from its lookalikes leave an agent far less room to pick the wrong one.

We'll keep this study going with new data. Next, we'll score the remote servers in the official MCP Registry, schemas included.

The same study is also on Qiita, in Japanese: https://qiita.com/nomunomu0504/items/b2eab1960a7aad38b59d

Top comments (3)

Collapse
 
reidmarlow profile image
Reid Marlow •

The root issue is that server authors still write MCP descriptions as API docstrings for humans rather than routing discriminators for a planner.

When a human reads a docstring, they already know why they opened that file. When an agent loads thirty tools into context, the failure mode is almost never misunderstanding what the function does. The failure mode is symmetric ambiguity with an adjacent tool. If search_code and grep_files both say they find text across the repository, the model effectively rolls dice on ambiguous prompts.

Flagging the 90 overlapping pairs is the most practical part of this audit. The logical next check for Forecall would be scanning for explicit negative boundaries, where descriptions name what they do not handle and pointer tokens to the sibling tool meant for that case. In my agent runs, adding a single negative pointer usually cuts down more tool-choice thrashing than expanding the parameter schema.

Collapse
 
nomunomu0504 profile image
Hiroki Nomura •

Thanks, this matches what we saw. On negative boundaries: Forecall already gives partial credit for them under "when to use". I tried your example: two tools that only said "Finds text across the repository" scored 0/20 on usage and were flagged as a lookalike pair. Adding one sentence to each ("to match an exact string or regex, use grep_files instead" / "not for semantic search, which search_code handles") took both to 15/20, and the pair was no longer flagged.

What it doesn't check yet is whether lookalike tools point at each other: when two tools are flagged as a pair, does each description name the other? That's a good rule. I'm taking it to our list.

Collapse
 
nomunomu0504 profile image
Hiroki Nomura •

Update: it's now an issue, with you credited: github.com/forecall/forecall-cli/i...