DEV Community

Vishal Habib
Vishal Habib

Posted on AI-assisted

I scanned 3,923 MCP servers. 1 in 4 tools leaves the model guessing.

TL;DR

  • I ran one static scan (mcp-doctor v1.15.6) over every public MCP server GitHub search could find: 3,923 servers, 147,646 tools, October 8–9, 2026.
  • The common problems are boring ones. About 1 in 4 tools has a parameter with no description, and 41% of servers (Go excluded, where the check doesn't apply) hand the model raw exception text. Popular servers are no better.
  • The scary one is nearly absent: 13 servers had tool-poisoning language, I read every one, and none was a real attack. The bug that actually breaks calls, Python None defaults, is in 8.6% of Python servers. 7 of the 14 fixes I've sent for it are already confirmed upstream.

Why I did this

An agent never reads your MCP server's code. It reads three things: the tool's name, its description, and the JSON schema generated from its parameters. Whatever the code means but those three leave out, the model can't see.

There are thousands of public MCP servers now and no shared bar for those three things. I build mcp-doctor, a static checker for MCP servers. So I pointed it at every public server I could find, to get a baseline for the whole ecosystem instead of one repo at a time.

How I counted

  • Which repos. GitHub search over the MCP topics (mcp-server, model-context-protocol, mcp-servers, modelcontextprotocol, mcp) and the phrases "mcp server", "fastmcp" and "model context protocol". Python, TypeScript, JavaScript and Go only. 20+ stars, no forks, no archived repos. Each query was split by star range so no slice hit GitHub's 1,000-result cap. That found 6,054 repos as of October 8, 2026.
  • How. A shallow clone of each default branch (October 8–9, 2026), one mcp-doctor --json run, then the clone was deleted. Nothing was executed and no hosted endpoint was called.
  • What counts as a server. A repo where mcp-doctor found at least one tool: 3,923 servers, 147,646 tools. 1,992 repos had no tools (clients, lists, frameworks, plus servers mcp-doctor can't read, more on that below), and 139 couldn't be scanned: 110 over the size limit, 21 failed to clone, 8 crashed mcp-doctor.
  • One version. Running the census turned up 14 bugs in mcp-doctor itself (#78–#91). Most were tool shapes it didn't recognize yet. One was a false positive and one a slowdown, both introduced by my own fixes earlier that week. I fixed them all, then rescanned so every row comes from the same release (v1.15.6), not an older, blinder scanner.

This isn't a runtime test, a security audit, or a ranking of anyone's project. I'm not publishing a per-repo grade list.

The big picture

Repos found 6,054
Servers (≥1 tool) 3,923
Tools 147,646 (median 12 per server)
Quality grade A / B / C / D / F 72.4% / 21.5% / 3.8% / 1.8% / 0.5%
Security grade A / B / C / D / F 78.9% / 11.9% / 4.1% / 2.1% / 3.0%

Most servers get an A. That's real: the basics (a description on every tool, typed parameters, a README) are mostly there. But an A is a floor, not a clean bill of health. The problems that matter to an agent sit one level down, in specific checks, and they cluster. 10% of servers hold 65% of all per-tool problems, and a quarter of servers (971) have none at all.

A note on counting. Tool counts are lopsided: one registry repo alone has 10,170 tools, and the 10 largest repos hold 19% of all tools. So for every per-tool number below I also checked it with the largest 1% of repos removed and as a per-repo average. Where those disagree, I say so.

Four things that go wrong for agents

1. Parameters the model can't interpret

48% of servers. About 1 in 4 tools (24.4% of all tools; 26.5% without the giant repos; 25.9% averaged per server) has at least one parameter with no description. This is the most common problem in the data, and it holds however you count it. The model sees a parameter called query or id and a type, but nothing about what belongs in it. It guesses, and a wrong guess looks like a tool failure.

2. Errors that don't help the model recover

41% of servers (Go excluded); 16–23% of tools depending on how you count (16.1% of all tools, 20.0% without the giant repos, 22.9% per-server average). FastMCP and the SDKs still return an error when a tool throws. What's missing is the message: the model gets raw exception text instead of "the date must be YYYY-MM-DD", so it retries blindly or gives up. (Not checked for Go, where errors are return values, or for the low-level setRequestHandler style, which has no per-tool handler.)

3. Optional parameters that reject their own default

151 of the 1,765 servers with Python tools (8.6%), 1,387 tools. This is the one that breaks calls outright. In Python, account: str = None produces a schema that says the type is string and the default is null. A client that sends the advertised default explicitly, which some agent frameworks do for every optional field, gets a validation error before the tool even runs. The fix is one annotation: Optional[str] = None.

Real examples, all already fixed by their maintainers:

It's a Python type-hint habit, and it's more common with FastMCP (12.0% of FastMCP servers) than with the official Python SDK alone (8.2%). The check only runs on Python, so there's no TypeScript number to compare.

4. Tools an agent can't tell apart

109 servers (2.8%), 985 tools. Two differently named tools with the same description. The agent has nothing to choose between them, and usually one of the descriptions was copied and is wrong.

Also common: tools with no description at all (513 servers, 13.1%), and URL parameters typed as a bare string with no format: "uri" (352 servers, 9.0%).

Does popularity buy quality?

stars servers median tools servers with an undocumented param share of each server's tools affected (avg)
20–99 2,482 12 45.9% 25.7%
100–499 918 13 48.6% 25.7%
500–999 202 12 51.0% 24.6%
1,000–4,999 213 17 60.6% 29.6%
5,000+ 108 16.5 62.0% 28.1%

No. Popular servers are more likely to have an undocumented parameter somewhere (62% of 5,000+ star servers vs 46% under 100), but most of that is size: they have more tools. Per server, the share of tools affected barely moves (about 26% vs 28–30%). Stars measure usefulness, not schema hygiene. Median quality score is 94–95 in every bucket.

By SDK

SDK (declared dependency) servers tools servers with an undocumented param servers with generic error text
TypeScript SDK 1,691 66,949 42.7% 38.4%
Official Python SDK 1,285 63,520 60.0% 43.7%
FastMCP (Python) 515 32,600 55.3% 50.9%
mcp-go 128 6,036 8.6% not checked
Go SDK 104 4,072 21.2% not checked

A repo can declare more than one SDK. Go servers document parameters far more often. A likely reason: in mcp-go the description sits right next to the parameter (mcp.WithString("city", mcp.Description("..."))), and the Go SDK reads it from a struct tag on the field. In Python, the description lives in a separate Args: docstring section that's easy to skip. I spot-checked popular Go servers and saw exactly that: one 11K-star mcp-go server puts Description(...) on every parameter line, and a 16K-star Go SDK server tags every field. Still, that's a likely cause, not something the data proves.

Security: what I'm confident about and what I'm not

mcp-doctor scores security separately from quality, so a well-documented server can't hide a security gap behind a good grade.

Tool poisoning: close to absent. A tool description that talks to the model ("ignore previous instructions", "do not tell the user") is the attack everyone worries about with MCP. mcp-doctor flags that language in 13 of 3,923 servers (17 tools). I read every hit by hand:

  • 7 are deliberate examples: security vendors' test servers (bad_mcps/, example-malicious-servers/, test-fixtures/evil-mcp-server.mjs) and a course lesson on tool poisoning (counted twice: the original and a translated copy).
  • 6 are real servers giving the model ordinary instructions: "before calling this endpoint, you must call listTables", "you MUST ALWAYS ask the user to confirm creation", a design-tool vendor's official server saying "ALWAYS CALL THIS TOOL FIRST BEFORE CALLING ANY OTHER TOOL". The closest call: "do NOT tell the user to run any CLI command — call auth_login_start immediately". That reads like poisoning out of context and is a login design choice in context. A phrase match isn't intent.
  • 0 looked like a real attempt.

None of the 13 phrase hits was a real attack, and a random sample of long descriptions (below) was clean too. That doesn't mean tool poisoning can't happen. This only covers public repos with 20+ stars, and a server can change its descriptions at runtime, which a static scan can't see. It does mean the ecosystem's real problems are the boring ones above.

Other precise checks (the pattern is unsafe whenever it appears):

  • unsafe deserialization (pickle.loads, yaml.load without a safe loader): 139 servers (3.5%)
  • hardcoded secrets: 175 (4.5%)
  • dependencies pinned to */latest or not at all: 155 (4.0%)

Heuristic checks (worth a look, not findings):

  • outbound HTTP with a variable URL (possible SSRF): 2,219 (56.6%)
  • shell/eval primitives present (not traced to tool input): 1,559 (39.7%)
  • tool descriptions over 500 characters: 816 servers (20.8%), 5,874 tools. Long descriptions are a place hidden text could sit. I read a random sample of 20 (19 readable): all were ordinary docs, like filter tables, warnings that a change can't be undone, lists of returned fields. So 816 of the 823 servers in mcp-doctor's raw "prompt injection" count are long-description hits: long docs, not attacks. (The other 7 are phrase-only; 6 servers have both.)

I'm not naming repos in this section.

How much to trust these numbers

When I checked a hit by hand, it was real. For None-default hits I reproduced the failure at runtime on each repo's own pinned SDK before opening a fix. 14 repos so far: 6 fixes landed, 1 maintainer fixed it in their own PR, 0 were rejected, and 7 are still open. Three more went in as issues. None of these maintainers work with me. That's a hand-picked sample (chosen for fix value), not a random audit, so it isn't a precision rate for every hit.

What mcp-doctor can't see. 715 repos with 0 tools found still declare an MCP SDK. That's the upper bound on missed servers; many are clients. On the independently counted coverage set (four frameworks, before this census), mcp-doctor finds 11,228 of 11,411 tools (98.4%). The biggest group here, the TypeScript SDK, hasn't had its own recall sweep yet (coverage). Update, Oct 9: that sweep is done. Across 703 repos that depend on the TypeScript SDK, recall was 80.3% (almost all of it one repo's own SDK) and is 99.15% after v1.15.8. Across 5 frameworks it's now 15,755 of 15,977 tools (98.6%). It's still blind to some dynamic registries, class-based low-level registries and in-house Go registries (github-mcp-server is the named example in the README).

Limitations

  • Static only: what the code says, not what a client sees at runtime.
  • Default branches as of October 8–9, 2026.
  • 8 repos crash mcp-doctor (a recursion limit on very deeply nested TypeScript/Go). They're counted as errors. That's my bug to fix next. Update, Oct 9: fixed in v1.15.7; all 8 now scan. The numbers above are from the census run and still count them as errors.
  • 20+ stars, four languages.
  • The checks are opinions about quality. Parameter docs and error messages are judgment calls. None defaults are objective.

If you maintain an MCP server: 10 minutes

  1. Write optional parameters as Optional[X] = None (or X | None = None on Python 3.10+), not X = None.
  2. Give every parameter a description (an Args: docstring section in Python, description in the schema elsewhere).
  3. Write an error message the model can act on ("date must be YYYY-MM-DD"), not a bare exception.

To check your own repo: pip install mcp-server-lint, then mcp-doctor path/to/server. Or add the GitHub Action (uses: vishalhabib99/mcp-doctor@v1), which suggests the fix on the lines a PR changes. No install at all: open a scan request with your repo URL and a bot replies with the report.

Method and scripts: mcp-doctor/census/2026-10 (aggregates only, no per-repo results).


Scanner: mcp-doctor. Everything else I build: github.com/vishalhabib99.

Top comments (0)

Some comments may only be visible to logged-in visitors. Sign in to view all comments.