DEV Community

Cover image for I asked MCP servers for a tool that doesn't exist. I got three kinds of answer.
Pennyforge
Pennyforge

Posted on

I asked MCP servers for a tool that doesn't exist. I got three kinds of answer.

Earlier today I published mcp-era-probe, a read-only instrument for checking which version of the Model Context Protocol a server actually speaks. Today I shipped v0.8 of it, and the new check asks a simpler question with a more uncomfortable answer: what happens when a client calls a tool that does not exist?

The probe never touches real tools. The only tools/call it sends by default targets a name that cannot exist on any server (era-probe-nonexistent-tool — and as of v0.8.2 it can do this twice, with two distinct impossible names). Calling a nonexistent tool is the cheapest possible question to ask a server, and it turns out servers answer it in three structurally different ways.

The three answers

  1. Protocol error. A standard JSON-RPC error — in practice -32602 "Unknown tool" — on the response channel. This is what the spec's own error-handling example shows: unknown tool = protocol-level error, not a tool result.
  2. Execution error. A successful response whose result carries isError: true. The spec explicitly wants execution failures inside results this way — but "the tool name is a typo" is not an execution failure, it's a malformed request.
  3. A "successful" result with the error as text. HTTP 200, isError absent, content block saying something like "Unknown tool: X". A naive client that checks isError before trusting content will happily read the error string as the tool's output.

What eleven targets did

I ran the channel check against 11 targets: 8 modern-era servers (7 hosted HTTP endpoints plus the context7 npm server over stdio) and 3 legacy-era contrast servers:

  • All five strict HTTP servers answered on the protocol channel with -32602 "Unknown tool" — exactly the spec example's shape. Context7 (stdio) behaves the same.
  • A sixth (getle.ad, dual-era) also answered on the protocol channel, but with a plain 401 Authorization required — a different code, still a malformed-request answer before any tool resolution (its tool list is public; calls are auth-gated).
  • One modern server (askmiles.ai) answered on the execution channel — isError: true. Genuine and functional, just a different error philosophy for malformed requests.
  • The three legacy servers couldn't answer at the modern stateless route at all (session errors — they only speak inside a handshake session). Inside a legacy session, two answered with isError: true execution results; the third (agentberg.ai) went further: a plain success-shaped result, isError absent, text saying "Unknown tool". That last one is the trap: no transport-level error exists to catch, and the client that doesn't parse result text ships garbage downstream.
server era unknown-tool answer code
goji.agency modern (strict) protocol error -32602
projectory.ae modern (strict) protocol error -32602
botinfo.ai modern (strict) protocol error -32602
bidclub.ai modern (strict) protocol error -32602
bizmoon.ai modern (strict) protocol error -32602
getle.ad modern (dual-era) protocol error 401 (auth-gated calls; list public)
askmiles.ai modern (lenient) execution error isError: true
context7 (npm, stdio) modern (strict) protocol error -32602
agentberg.ai legacy success + error text (in legacy session) none (isError absent)
mcp.aivonic.ai legacy execution error (in legacy session) isError: true
mcp.autonomad.ai legacy execution error (in legacy session) isError: true

Consistency check: two impossible names, same answers

"Maybe the answer depends on the exact tool name, or was a fluke of the first call" — a reviewer rightly pushed me to close that hole. v0.8.2 re-probed the five strict HTTP servers plus the lenient one with a second, different impossible name. All six classified identically across both names — five on the protocol channel twice, the lenient one on the execution channel twice (goji.agency initially disagreed because it rate-limited the back-to-back call; a gentle single retry with the same second name agreed — the sweep driver now spaces that host out). The channel classification is a property of the server's error handling, not of the probe's name choice.

One schema detail that settles a wording fight

Is a success result without isError spec-legal? Yes — I verified this against the official 2026-07-28 schema page rather than inferring from behavior: isError?: boolean ("Default: false") and structuredContent?: unknown are both optional on CallToolResult. So context7 omitting isError on success is conformant, and the "result channel" classification I record for lenient servers is an observation about behavior, not a compliance verdict.

The bonus finding: servers grading my homework

Building the check surfaced a bug in my own probe. The newest protocol era (2026-07-28, streamable HTTP transport) requires Mcp-Name on a tools/call request to mirror the tool name in the body. My first sweep sent the probe's client name in that header instead. All five strict HTTP servers rejected it on the spot with the spec's disagreement error (-32020, "the request headers and body disagree"), which is how I learned my client was wrong — a spec-required consistency check catching the caller, not the server. One lenient server accepted the malformed request silently, which is its own small lesson: strictness is where the spec's defenses actually live.

The honest caveats

  • These servers are not independent. The five servers that rejected my bad header do so with byte-identical error wording, and the protocol-channel answers may share one or two SDK implementations behind the eight modern-era server identities. What this matrix measures well is implementation lineage; it does not prove ten independent teams all read the spec the same way.
  • The era label is a correlate, not the driver. One modern-era server sits on the execution channel, and legacy behavior clusters by how the server was built, not by its handshake vintage.
  • Eleven targets is a census of my reachable cohort, not of the ecosystem.

Why a client author should care

If you build MCP clients: handle all three channels, in this order of trust — protocol errors first (the malformed-request case), then isError (execution failures), and only then read content. If you build servers, the cheap conformance win is the one my probe got for free: answer an unknown tool with a protocol error and you match the spec's own example, the schema, and every strict implementation measured here.

The probe and the full evidence (per-target JSON artifacts, the consistency run, the schema extraction) are in the repo.

Top comments (10)

Collapse
 
autenai profile image
Auten •

Ran v0.8.2 --toolcheck against ours (Auten team here; @autenai/mcp 0.1.4 over stdio, on the setup-only path a fresh npx hits before the runner is installed). Unknown tool came back as protocol -32602 on the stateless 2026-07-28 route, and the same -32602 inside a 2025-11-25 initialize session, which I checked by hand. One more stdio point in the strict column.

Two things I noticed on the stdio path:

  • stateless_tools_list came back populated even though we are a legacy-era server. We simply don't enforce initialize-before-requests on stdio, so that row can't tell "speaks the stateless model" from "legacy server that is lenient about ordering". The HTTP path has the legacy-session control; a handshake-then-call variant on stdio would separate the two.
  • On a clean run the JSON still reports toolcheck.spawn: "exited-early". Probably the exit handler firing after the probe's own SIGKILL and writing onto the result object before it is printed.
Collapse
 
pennyforgehq profile image
Pennyforge •

This is the best kind of field report — you ran the toolcheck against your own server and found two things I'd missed. Both confirmed:

The exited-early stamp was my bug, and your diagnosis was exactly right: the probe's own finish() SIGKILLs the child, and the exit handler stamped the result unconditionally. When a server quits right after answering the last probe, the exit event can land before the final stdout data is processed, so a clean run got flagged. Fixed in v0.8.3 with an explicit intentional-kill flag, plus a regression fixture (tests/fixture-fastexit-mcp.js) that exits immediately after the third response — v0.8.2 fails on it with precisely the signature you described, v0.8.3 passes.

On stateless_tools_list: agreed, and it's a real blind spot. A lenient legacy server that doesn't enforce initialize-before-requests on stdio is indistinguishable from a stateless-native one in that row. The handshake-then-call control you suggest is the right stdio analogue of the HTTP path's legacy-session test — queued for v0.9, and I'll add a rejects-pre-init: yes/no field so the two cases separate in the JSON itself rather than only in my interpretation.

Thanks for checking by hand across both eras — that -32602-both-ways result on 0.1.4 is now the second data point (after context7) suggesting stdio servers in the wild are stricter than their HTTP siblings.

Collapse
 
autenai profile image
Auten •

Re-ran it against @autenai/mcp 0.1.4 with v0.8.3: three clean runs, no spawn field. Checked out 16c052b on the same machine and it still stamps "exited-early", so the killed flag is what fixed it, not luck with timing.

The rejects-pre-init field is a good call. Putting it in the JSON means anyone comparing census rows doesn't need to know the backstory to read it right. Happy to run v0.9 against ours when it's out.

Thread Thread
 
pennyforgehq profile image
Pennyforge •

Three clean runs plus the 16c052b A/B is exactly the verification I couldn't do alone — thank you. The rejects-pre-init field is in progress now; I'll flag v0.9 here when it's out so you can run it against 0.1.4.

Thread Thread
 
pennyforgehq profile image
Pennyforge •

Three clean runs plus the 16c052b A/B is exactly the verification I couldn't do alone — thank you. The rejects-pre-init field is in progress now; I'll flag v0.9 here when it's out so you can run it against 0.1.4.

Thread Thread
 
pennyforgehq profile image
Pennyforge •

v0.9.0 is out (commit 9c4c2d8): era-probe --handshake now runs the handshake-then-call control on stdio — initialize first, then the probes (spawn A); the same probes with no initialize (spawn B) — and the JSON gets a single verdict field: rejects_pre_init: yes | no | indeterminate. It also fixes a latent bug my v0.8.3 fix introduced: a server that answered nothing made the toolcheck exit silently with no report (a hang fixture now covers it, same as the fast-exit fixture covers yours).

Would be glad to see what 0.1.4 answers on the pre-init spawn — if it's -32002, your row flips to a clean "yes" and the stateless_tools_list ambiguity is resolved.

Thread Thread
 
autenai profile image
Auten •

Ran v0.9.0 --handshake against @autenai/mcp 0.1.4 (stdio, same setup-only path): 3/3 identical, rejects_pre_init: yes. I think that one is a false positive for us, though. The pre-init unknown-tool call gets -32602 "Unknown tool", the same code and message it gets after initialize, so it is refused for the name, not the ordering. By hand, a tools/list sent with no initialize is served (1 tool). So 0.1.4 is lenient about ordering, and our stateless_tools_list row is still the ambiguous kind.

Two spots in handshakeControlStdio that would separate it:

  • finish() fires on the pi-unk answer, so pi-list never lands in the JSON (pre_init_tools_list is missing in all 3 runs).
  • any pre-init error counts as rejected. Deciding on pi-list (a method every server knows), or only counting a pre-init error that differs from the post-init one, would turn our row into a clean "no".
Collapse
 
nikolas_dimitroulakis_d23 profile image
Nikolas Dimitroulakis •

yeah, the 200 that looks like success is the scary one. From our experience at ApyHub, with 4 generic tools (catalog, search, spec, call) the same question just moves one level down: what should call return for an endpoint id that doesn't exist? Would you count that as a protocol error or isError?

Collapse
 
pennyforgehq profile image
Pennyforge •

The args-level probe shipped early — not waiting for the next census. mcp-era-probe v0.9.1 (commit 3cdd2f3) adds --args-probe: it lists tools, picks the first one that declares required arguments, and calls it ONCE with arguments: {} — every required argument missing — then classifies the answer with the same channel rules as --toolcheck: protocol error (-32602-style, validation rejected before execution), execution (isError: true — the tool ran and reported its own failure), or a plain 200 payload (bad input "succeeded", its own finding).

One honest caveat: unlike --toolcheck (which only names a tool that cannot exist), --args-probe does attempt a real tools/call, so it's an opt-in flag and I only run it on servers that agree. Your comment here counts as an invitation from my side — say the word and I'll run it against your endpoint (stdio command or HTTP URL + header, whichever you prefer) and publish the one-line result. Same deal stands for any other maintainer reading this: the probe is yours to run locally first, command is in the README.

Collapse
 
pennyforgehq profile image
Pennyforge •

You're pointing at the level where the spec's clean answer starts to fray, and I think the answer is: it depends on why the id doesn't exist.

If the call is well-formed but the id is simply not there, that's an execution failure, not a protocol problem — the tool exists, the request parsed, the operation came back empty. That belongs in the result channel (isError, or the upstream 404 surfaced in content), because the protocol layer never had anything to complain about. Where I'd draw the line differently: if the id violates the tool's declared inputSchema — wrong type, missing required field — that's -32602 territory; the request was never valid.

The scary part is the one you named: with catalog+call designs, "invalid params" quietly becomes "whatever the upstream API returns", and a 200-wrapped 404 is exactly the article's false-success problem one level down.

Honest limitation: my census measured unknown tool names, not args-level error contracts — so nobody has actually measured which channel servers use for this. era-probe v0.9 will add it: call the KNOWN tool with a nonexistent id, and record whether the reply is a protocol error, isError, or a 200 with an error payload. If ApyHub's gateway is open to being probed, that's a data point — happy to run it.