DEV Community

Jeff
Jeff

Posted on Originally published at powerduck.com

How to Test and Debug an MCP Server: From the Inspector to Automated Tool Tests

Testing an MCP server through an AI client is a slow way to work. You type a natural-language prompt, hope the model picks the right tool, watch it fill in arguments, and if anything fails you cannot tell whether the bug is in your server, the model's reasoning, or the prompt. Every layer of that stack is nondeterministic except yours. Test yours first, deterministically, and leave the AI out of it until the protocol surface is provably correct.

The protocol makes this easy. MCP is JSON-RPC 2.0 over a defined transport. Every capability a client uses is a request you can send by hand.

Level 0: does the process even start

Before debugging protocol behavior, eliminate the failures that look like protocol bugs but are not:

echo '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2025-06-18","capabilities":{},"clientInfo":{"name":"smoke","version":"0"}}}' | npx @acme/orders-mcp
Enter fullscreen mode Exit fullscreen mode

A healthy stdio server responds with its capabilities and then waits on stdin. If you instead get a module resolution error, a missing environment variable, or a crash on a native dependency, no amount of Inspector work will help. This one-liner is worth keeping as the first CI smoke test, because packaging problems (a binary not bundled, a .env required at startup) show up here immediately.

Level 1: the MCP Inspector for discovery

The Inspector is the official visual client maintained with the SDKs. It launches a stdio server or connects to a remote one and exposes the raw protocol:

npx @modelcontextprotocol/inspector npx -y @acme/orders-mcp
Enter fullscreen mode Exit fullscreen mode

It opens a browser UI where you can:

  • Read the initialize handshake and negotiated capabilities.
  • Browse tools/list, including every input schema as the client will see it.
  • Call a tool with a form generated from that schema, which makes missing or mistyped properties obvious.
  • Inspect resources, prompts, and the exact JSON-RPC responses.

The Inspector is the right tool for "what does the server advertise?" and for reproducing a single call. It is the wrong tool for regression testing: clicking through forms is not repeatable and cannot run in CI.

Level 2: drive the protocol with raw requests

Every Inspector action is a JSON-RPC message. Send them directly to learn the contract and to script repro cases. After initialize, a stdio session requires an notifications/initialized notification before tool calls:

{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2025-06-18","capabilities":{},"clientInfo":{"name":"curl","version":"1"}}}
{"jsonrpc":"2.0","method":"notifications/initialized"}
{"jsonrpc":"2.0","id":2,"method":"tools/list"}
{"jsonrpc":"2.0","id":3,"method":"tools/call","params":{"name":"get_order","arguments":{"order_id":"ord_8821"}}}
Enter fullscreen mode Exit fullscreen mode

For a Streamable HTTP server the same messages are POST requests:

curl -sS https://mcp.example.com/mcp \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -H "Accept: application/json, text/event-stream" \
  -d '{"jsonrpc":"2.0","id":2,"method":"tools/list"}'
Enter fullscreen mode Exit fullscreen mode

Keeping a folder of these request transcripts doubles as documentation. When a user reports "the agent calls my tool wrong," the first question is whether a raw correct call works; the transcript answers it in ten seconds.

Level 3: deterministic automated tests

An MCP server's surface is small enough to test exhaustively. Treat the tool registry as an API and test four layers.

Handshake and inventory. Assert the server initializes, advertises the capabilities it claims, and every tool has a name, description, and a valid JSON Schema for its input:

const tools = await client.request(
  { method: "tools/list" },
  ListToolsResultSchema,
);

for (const tool of tools.tools) {
  expect(tool.description?.length).toBeGreaterThan(10);
  const valid = validateMetaSchema(tool.inputSchema);
  expect(valid, `${tool.name} has invalid input schema`).toBe(true);
}
Enter fullscreen mode Exit fullscreen mode

The schema-meta check matters: a model cannot fill in arguments for a schema that is itself invalid, and the failure mode is silent and intermittent.

Happy paths against a fake backend. Point the tool handlers at an in-memory or containerized double of the real service. Assert both the tool result content and the side effect:

const result = await client.request(
  { method: "tools/call", params: { name: "cancel_order", arguments: { order_id: "ord_1", reason: "customer request" } } },
  CallToolResultSchema,
);
expect(result.isError).toBeUndefined();
expect(fakeOrders.cancelled).toContain("ord_1");
Enter fullscreen mode Exit fullscreen mode

Validation and error paths. Call every tool with missing required fields, wrong types, and out-of-range values. The server should return structured errors, not stack traces:

const result = await client.request(
  { method: "tools/call", params: { name: "get_order", arguments: { order_id: 12345 } } },
  CallToolResultSchema,
);
expect(result.isError).toBe(true);
expect(JSON.stringify(result.content)).toMatch(/order_id/);
Enter fullscreen mode Exit fullscreen mode

Models recover well from errors that name the offending field and the expected type. They hallucinate endlessly around a generic "internal error."

Idempotency and retries. Re-send any tool that mutates state with the same arguments and an idempotency key where supported. A dropped HTTP connection must not create a second order. This is the layer teams skip and regret; see the agent-facing API design guidance in designing APIs for AI agents.

Test the descriptions like you test code

Tool descriptions are part of the program. They are the only thing a model reads when deciding which tool to call, so review them with the same rigor as schemas:

  • Names say what the tool does in verb-noun form (refund_payment, not paymentOp).
  • Descriptions state the preconditions, side effects, and when not to call the tool.
  • Enum values are documented inline; a model cannot guess that status: 3 means "shipped."
  • Two tools never overlap ambiguously. If create_order and place_order both exist, the model will alternate.

A cheap, high-value test is to snapshot the full tools/list payload and review the diff in pull requests. Renaming a tool or narrowing a parameter is a breaking change for every agent that learned the old surface; versioning that correctly is covered in versioning MCP tools without breaking agents.

Debugging the fuzzy middle

When raw calls pass but an AI client still misbehaves, the bug is usually one of three things:

  1. The model never sees the tool. Capabilities were not advertised at initialize, or the client caps the number of tools and yours is past the limit. Check the actual tools/list response the client logs.
  2. Arguments are coerced loosely. Your schema says string, the model sends a number, and the handler accepts it by accident. Tighten the schema; do not teach the model a bad habit.
  3. Results are unreadable. A tool that returns a 40 KB serialized object buries the answer. Return structured, minimal content and expose the full record as a resource the agent can fetch on demand.

When the tools come from an OpenAPI spec

Servers generated from an OpenAPI document inherit a test oracle for free: the spec already defines inputs, outputs, and status codes. The same scenario tests written against the API can assert that each generated MCP tool validates arguments the same way and returns the same errors the documented API would. That keeps "what the agent can call" and "what the service actually does" from drifting apart, which is the failure mode hand-written MCP glue always reaches eventually.

For the lighter-weight contract approach behind this, see API contract testing without Pact's overhead, and you can generate a local MCP server from any spec in the online demo to see the tool inventory before writing any code.

Top comments (0)