MCP has outputSchema so agents can validate tool results. But does the schema actually reject a wrong answer? I built mcp-drill -- a fault-injection harness that speaks MCP -- and scanned 31 popular servers (265 tools) including Microsoft Learn, Hugging Face, Cloudflare, DeepWiki. Result: only 3% declare a contract that would reject a corrupted response. 56% declare nothing, 42% declare a schema that happily validates garbage. Your agent cannot tell a bad result from a good one.
- Live scorecard: timurrakhmatullin86.github.io/mcp-drill
- Install:
pip install mcp-drill[scan] - How this differs from security scanning: vs mcp-scan
Why I built this
MCP is JSON-RPC over stdio / Streamable HTTP with bidirectional notifications. Normal HTTP chaos tools don't speak it. And even when you test an MCP server, you usually test your agent, not whether the server's contract protects you.
Real failures I kept hitting:
- A tool returns a well-formed but truncated payload mid-stream -- agent acts on half a JSON.
- A
fetchtool returns{"result": "ok"}with status 200 even when the tool name is wrong -- how does the agent know it failed? - A server declares
outputSchema: {type: "object"}-- great, it validates any object, including a corrupted one. Zero protection.
I wanted one command to answer: if I corrupt the response but keep its type, does your schema catch it? And separately: if I send you garbage input, do you tell me with a proper error?
So I built mcp-drill:
-
Fault injection proxy --
mcp-drill wrap --faults timeout,corrupt,truncate,malformed -- npx ...-- sits between client and server, perturbs responses deterministically (seed-able). -
Model-free scorecard --
mcp-drill scan -- npx ...ormcp-drill scan --url https://...-- no LLM, deterministic, reproducible. Every number is a property of the server.
What I measured (no LLM involved)
For each server, mcp-drill does the MCP handshake, lists tools, then runs fixed probes:
-
Output-contract coverage -- % of tools that declare an
outputSchemaat all. -
Output-contract enforceability -- of those that declare one, % whose schema rejects a corrupted-but-well-typed payload. Corruption: keep structure + type, replace every leaf with
mcp-drill-corruption/-999999999/ out-of-range. If it still validates -> vacuous. If it rejects -> enforceable. This is an outcome test, not style policing. -
Error conformance -- 3 probes: unknown method, unknown tool, missing required args. Classified as
jsonrpc_error/tool_error(good) vsaccepted/timeout/crash(bad).
Repro: pip install -e ".[scan]" && python studies/pilot/run_pilot.py studies/pilot/servers.json -- commits server list + raw results.json.
Full methodology: METHODOLOGY.md
The headline number
31 servers, 265 tools -- 3% enforceable.
| Tier | Share | Count | Meaning |
|---|---|---|---|
| No schema | 56% | 148/265 | Nothing to validate against |
| Vacuous schema | 42% | 110/265 | Corrupted payload still validates -- zero protection |
| Enforceable | 3% | 7/265 | Schema rejects the corrupted payload |
Notable:
- SDK-auto-wrapped (
x-fastmcp-wrap-result): 10 tools -- the default from the dominant MCP Python SDK (FastMCP) wraps a return as{"result": string}and calls it a contract. It's vacuous by construction. - Error handling: 30/31 servers handle bad input correctly -- the error path is healthy. The success path is not.
- The number is stable: 3% at 18 servers -> 3% at 22 -> 2% at 26 -> 3% at 31 (including remote marquee servers). Not a small-sample artefact.
Some highlights:
| Server | Tools | Enforceable | Notes |
|---|---|---|---|
| git-mcp-server | 28 | 18% | Best of the bunch |
| huggingface (remote) | 8 | 12% | Only marquee with enforceable schemas |
| filesystem | 14 | 7% | Best reference server |
| everything | 13 | 0% | Reference server, 100% vacuous |
| microsoft-learn (remote) | 3 | 0% | Name-brand doesn't help |
| deepwiki (remote) | 3 | 0% | 100% vacuous |
| playwright | 23 | 0% | No schemas at all |
| desktop-commander | 26 | 0% | No schemas at all |
Why this matters for agents
Agents increasingly act on a tool result without a human in the loop: tool A's output becomes tool B's input. The only automatic guard is: did the transport succeed + did the payload match outputSchema? If the schema is vacuous, nothing guards a well-typed but wrong result, and the agent proceeds on bad data.
This is not what security scanners (like mcp-scan) catch. Those ask "can this server be abused to do something evil?" We ask "can this server be trusted when it returns a result?" See the VS page.
And coverage is a vanity metric here. Auto-generated schemas (FastMCP infers from return type hints) raise coverage toward 100% while enforceability stays near 0% -- a vacuous default inherited by every server that doesn't override it. The gap widens as tooling improves, unless schemas add value-level constraints (enum, pattern, format, bounds).
Try it on your own server
# install
pip install "mcp-drill[scan]"
# or without install
uvx mcp-drill scan -- --help
# local stdio server
mcp-drill scan -- npx -y @modelcontextprotocol/server-filesystem /tmp
# remote Streamable HTTP
mcp-drill scan --url https://mcp.deepwiki.com/mcp
# JSON output for CI
mcp-drill scan --json -- npx -y @modelcontextprotocol/server-filesystem /tmp > mcp-drill.json
# badge (shields.io endpoint)
mcp-drill scan --badge --url https://mcp.deepwiki.com/mcp > badge.json
# fault injection proxy
mcp-drill wrap --faults timeout,truncate -- npx -y @modelcontextprotocol/server-everything
Gate in CI -- GitHub Action (no LLM, no API key):
- uses: TimurRakhmatullin86/mcp-drill@v0
with:
server: 'npx -y @modelcontextprotocol/server-filesystem /tmp'
min-error-handling: '0.9'
What to do about it
-
If you build an MCP server: declare
outputSchemas that constrain values, not just shapes. Addenum/pattern/format/ numeric bounds where semantics allow.additionalProperties: falsehelps, but alone it's not enough -- a corrupted string is still a string. Test withmcp-drill scan --jsonin CI and gate on it. -
If you consume MCP tools: don't rely on
outputSchemapresence as safety. Validate semantically downstream or usemcp-drill wrapto exercise your agent's failure paths before prod. - If you review MCP proposals: coverage will trend to 100% as generators spread. Ask for enforceability.
GitHub repo -- Apache-2.0, telemetry off, Python 3.10+. Live scorecard. PRs and issues welcome -- especially if your server scores differently and you think the harness is wrong.
Method is model-free and deterministic -- every number is a property of the server, not of whatever agent called it. The tool is the methodology, and it is released.
Top comments (8)
The vacuous-schema tier is the finding that should scare people, more than the 56% who declare nothing. "No schema" at least fails honestly — the agent has no false sense of safety. But
outputSchema: {type: "object"}is worse: it looks like a contract, it passes a linter, it makes a reviewer nod, and it protects against exactly zero real corruptions. It's safety theater that survives code review precisely because it's present.The failure mode I keep hitting in agent loops maps straight to your truncation probe: a well-formed but half-streamed JSON payload where the top-level shape is intact but a nested array got cut. Type-valid, semantically garbage, and the agent confidently acts on it. Your "keep the type, corrupt the leaves" perturbation is the correct test because it targets exactly that gap.
Curious whether you saw any correlation between SDK-auto-wrapped servers and the vacuous tier — my hunch is the frameworks that auto-generate a schema from a return annotation are quietly manufacturing most of that 42%. Great work making it deterministic and LLM-free; that's what makes the number citable.
You nailed the worst part: vacuous schemas survive code review because they look like contracts. A reviewer sees outputSchema with properties and moves on. The 42% tier is exactly this -- auto-generated schemas from FastMCP's return type hints that raise coverage metrics toward 100% while enforceability stays at 0%.
On the SDK correlation: yes, the x-fastmcp-wrap-result pattern accounts for a chunk of the vacuous tier. Any server built with FastMCP that returns a plain type gets wrapped as {"result": string}, which validates anything. The framework manufactures the schema, not the developer, so they don't even know the contract is vacuous.
The truncation probe you described (top-level shape intact, nested array cut) is the exact gap that output-contract enforceability is meant to catch. If the schema only constrains the envelope, a half-streamed payload passes validation and the agent acts on incomplete data. Adding value-level constraints (enum, pattern, bounds on array length) is what turns a schema from decoration into a guard.
The {type: "object"} schema is the perfect mascot for this problem: a contract that validates any object is security theater with a JSON syntax. Fault injection at the protocol level is the right test - almost everyone evaluates their agent against the server, never the server's contract against corruption. The fetch-returns-ok-on-a-wrong-tool-name case is the one I have hit in the wild: the agent has no channel to learn that it failed, so the failure just becomes context. 56% declaring nothing tracks with every MCP integration I have wired up; the schema slot exists and nobody fills it because nothing rejects an empty one.
The fetch-returns-ok-on-wrong-tool-name case is why error conformance needed to be a separate axis in the scorecard. The agent has no signal that anything went wrong, so the failure propagates as context. The vacuous tier is arguably worse than declaring nothing, because someone actively wrote a schema that validates everything. Considering adding a schema quality badge to CI output so server authors get a visible signal before shipping.
Would love to see this extended to stateful/session-bound tools. I'd expect the numbers to get even worse.
Agreed -- stateful tools are the next frontier and I'd bet the numbers drop below 1%. The current scorecard probes are stateless by design (single request/response), so they miss a whole class of contracts that only matter across calls: session tokens that should expire, cursor-based pagination where skipping pages silently drops results, multi-step workflows where step 2 should reject if step 1 wasn't called. Those are all observable properties of a server, so they're testable without an LLM -- but the probe design needs a session graph, not just individual perturbations. It's on the roadmap.
Some comments may only be visible to logged-in visitors. Sign in to view all comments.