DEV Community

Vishal Habib
Vishal Habib

Posted on Edited on

I built three tools to audit MCP servers. Each one found a bug in itself first.

Over the last couple weeks I built three small, independent CLIs that each check a different way an MCP (Model Context Protocol) server can be broken — not by reading marketing copy, but by pointing them at real, popular servers and reading what came back.

The three tools ask three different questions:

  • mcp-doctor — is it documented? Static analysis of the source: missing tool descriptions, undocumented parameters, hardcoded secrets, no error handling, no tests.
  • mcp-fuzz — does it fail safely? Actually launches the server over stdio and calls every read-only tool with schema-derived bad input (missing required fields, wrong types) to see if it returns a structured error or just crashes.
  • mcp-reality-check — does it actually work? Calls tools with realistic input and checks whether a "successful" response is actually honest: not a disguised refusal, not empty content, not ignoring its own declared output schema.

None of them use an LLM to judge anything. All three are fully static/heuristic by design — deterministic, no API key, no per-call cost.

The pattern I didn't expect: each tool found a real bug in itself

I dogfooded all three against real servers as I built them, and the same thing happened three times: the first serious dogfood pass against a real-world server found a genuine bug in my own tool's logic, not the target.

  • mcp-doctor, against homeassistant-ai/ha-mcp (4.7k stars): the secret scanner was flagging test fixtures and identifier-style constants as hardcoded credentials, and tool detection was counting mock functions inside test files as real tools. Both wrong — the server's actual score corrected from a false 55%/F to an accurate 89%/B. That fix led to a PR merged upstream: https://github.com/homeassistant-ai/ha-mcp/pull/2327
  • mcp-fuzz, against mendableai/firecrawl-mcp-server: a server that validates input with zod and correctly returns a well-formed JSON-RPC -32602 INVALID_PARAMS error was being classified identically to an actual crash, because of a blanket except Exception. Root-caused the real distinction (-32602 = server validated and rejected properly; -32603 and everything else = still counts as broken) and fixed it — but not before nearly flipping an already-published, already-cited finding on a different repo to a false pass. Caught that by re-testing the cited repo before shipping, not after.
  • mcp-reality-check, against modelcontextprotocol/server-time: no timezone hint existed in the realistic-input generator, so every timezone-shaped property got a bogus placeholder string and every call failed. Fixed by adding a real IANA timezone name as the hint.

I don't think this is a coincidence. It's what happens when you build a tool whose entire job is judging correctness, then finally point it at something popular enough to have edge cases you didn't think of. If your own tool never finds a bug in itself the first time it meets the real world, you probably haven't tested it against anything hard enough yet.

The strongest finding so far

Running mcp-fuzz against antvis/mcp-server-chart (4.3k stars, official Ant Design MCP server, 27 chart-generation tools) found that all 27 tools crash — raw JSON-RPC -32603 internal errors, not structured tool-level errors — on missing-required or wrong-type input. 133 of 214 test calls failed this way. Filed as issue #323: https://github.com/antvis/mcp-server-chart/issues/323

Traced it further: the fix already exists on main (an unreleased commit, 9fd0bb4), just never published to npm. Built main locally, replayed all four repro cases from the issue directly over stdio, confirmed the fix actually resolves what mcp-fuzz flags. Posted that as a comment instead of a "please fix this" ask, since there was nothing left to fix — just an unpublished release. Update: the fix has since landed and I rescanned the live repo to confirm it — crash resilience went from 37.85%/F to 100%/A.

Where things stand

  • mcp-doctor: 40+ real-world dogfood passes, 37 genuine bugs found and fixed, one PR merged upstream, a live public leaderboard scoring 21 real MCP servers on quality and security — each with a real per-repo badge maintainers can embed in their own README: https://vishalhabib99.github.io/mcp-doctor/
  • mcp-fuzz: 23 real-world passes, including the antvis finding above (now confirmed fixed).
  • mcp-reality-check: 20 real-world passes, on par with mcp-fuzz on track record.

None of these have real traction yet — this is still early, built solo, with basically zero stars or followers behind it. I'm writing this up because the process itself — build a tool, immediately distrust it, verify it against something real before believing its output — is the part I think is actually worth sharing, independent of whether the tools themselves ever get popular.

Update: a fourth piece, and the loop closing on an external collaboration

Since first publishing this, two things happened worth adding rather than quietly editing away.

The three became four. mcp-trust-check wraps mcp-doctor, mcp-fuzz, and mcp-reality-check into one GitHub Action that runs all three against a target server and posts a single combined score instead of three separate installs — self-verified live against the official @modelcontextprotocol/server-memory reference server (doctor 86%/B, fuzz 100%/A, reality-check 100%/A, combined 95.33%/A) before shipping. It's also a Python package (GuardedSession) applying the same combined idea to a live agent session instead of a CI run.

An external maintainer shipped a fix citing this work. Raju Dandigam, maintainer of agent-inspect (a trajectory-debugging tool for TypeScript agents), asked to feed a real mcp-fuzz session into their evidence model. The finding that surfaced — a healthy, 100%/A crash-resilience session reads as almost entirely failed once isError: true rejections collapse into a generic error status — was substantial enough that he shipped --preset behavioral-session in response, verified it against the same sanitized repro, and merged a privacy-reviewed copy of the trace into his own repo's fixtures. First time in this project an external maintainer built and merged something directly attributed to this work, not just fixed their own bug in response to a report.

Most recently, dogfooding mcp-fuzz against a fresh real target (jon-the-dev/linkding-mcp-server) surfaced a new, generic gap: a url-shaped string parameter with no format: "uri" JSON-schema hint gives a schema-only client no signal it needs to be a real URL. Confirmed it wasn't a one-off by checking five other independently-authored servers already in the dogfooding history — found the same real gap in all five — then shipped it as a new mcp-doctor check (v1.11.0 / PyPI 0.11.0).

Repos: mcp-doctor · mcp-fuzz · mcp-reality-check · mcp-trust-check

Top comments (2)

Collapse
 
alikhatersaibreakroom profile image
Ali Khater •

A tool finding its own bug is a better advertisement for characterization tests than any benchmark. For agent infrastructure, I would add one adversarial pass where the auditor receives malformed tool metadata, oversized payloads, and conflicting schemas. The interesting failures often live in the contract between components, not in the component being tested.

Collapse
 
vishalhabib99 profile image
Vishal Habib •

Good call, and it turned out to be a real gap, not a hypothetical one. Fed the schema generator some hand-crafted tools/list metadata (non-object properties, required as a string instead of a list, a schema nested a few thousand levels deep) and found three actual crashes in mcp-fuzz's own code — it had only ever been fed metadata from well-behaved real servers, so nothing had exercised the malformed-metadata path before. Fixed on main, with 4 new regression tests: github.com/vishalhabib99/mcp-fuzz — write-up in the README under "Hardened against malformed tool metadata." Appreciate the nudge — exactly the kind of thing that's easy to miss when every test target so far has been a real, spec-compliant server.