DEV Community

Hoang Le
Hoang Le

Posted on

Testing MCP agents offline: the recording is the server ๐Ÿ“ผ

TL;DR ๐Ÿ“ผ mcp-cassette records one real MCP session into a JSONL file, then serves that file as an MCP server. Any client, in any language, runs its tests against the recording offline: no credentials, no rate limits, no network.

Terminal demo: check a live server, record, replay offline, lint, and fail on a breaking contract change

๐Ÿค” The problem

If your agent calls tools over the Model Context Protocol, its tests depend on something they do not control: a live MCP server, with its credentials, its rate limits and its network.

The usual fixes both hurt:

  • ๐Ÿงช Mock the MCP client library, and every test is tied to one SDK.
  • ๐Ÿ› ๏ธ Write your own recorder, which several projects I came across had each done for themselves.

mcp-cassette takes a third route: the recording is the server.

๐ŸŽ™๏ธ Record once

npx mcp-cassette record -o session.cassette.jsonl -- npx -y @modelcontextprotocol/server-github
Enter fullscreen mode Exit fullscreen mode

The recorder sits between your client and the real server as a transparent proxy and writes every frame, in both directions, to an open JSONL cassette. Secrets it recognizes are redacted before they reach the file, by default.

๐Ÿ’ก Pattern matching cannot catch every secret, so give a cassette a quick look before you commit it.

๐Ÿ” Replay forever

npx mcp-cassette check --stdio "npx mcp-cassette replay session.cassette.jsonl"
Enter fullscreen mode Exit fullscreen mode

check is an ordinary MCP client, and it cannot tell the recording from the live server. Neither can yours. The replay speaks MCP over stdio or Streamable HTTP, so any client in any language connects to it unchanged.

It is a server, not a stub:

  • ๐Ÿ“ฃ Notifications the server pushed on its own are replayed at the point in the session where they arrived.
  • ๐Ÿงญ Both the 2025-11-25 revision and the stateless 2026-07-28 revision are supported.
  • ๐Ÿ”Œ A client that probes on one connection and runs its session on a second records both into one file with --mode append.

๐Ÿงช Inside your test runner

import { describe, expect, it } from "vitest";
import { useCassette } from "mcp-cassette/vitest";

describe("recorded call", () => {
  const tape = useCassette("tests/cassettes/weather.http.jsonl");

  it("is answered from the cassette", async () => {
    // point your MCP client at tape.url instead of the live server
  });
});
Enter fullscreen mode Exit fullscreen mode

There is a mcp-cassette/jest twin. A request the recording does not hold fails the test that made it, and says why.

๐Ÿšฆ Gate your tool contract

A tool contract is an API. snapshot writes it to a file you commit, and snapshot --check fails CI when it breaks:

npx mcp-cassette snapshot --check --stdio "node dist/my-server.js"
Enter fullscreen mode Exit fullscreen mode
[BREAKING] slugify: tool removed (tool-removed)
[BREAKING] add: parameter "precision" is now required (input-property-became-required)
[DANGEROUS] add: parameter "mode" added (input-property-added-optional)
result: FAIL (2 breaking, 1 dangerous, 0 minor, 0 info; gate: breaking)
Enter fullscreen mode Exit fullscreen mode

๐Ÿ›ก๏ธ Lint what the model reads

Tool poisoning hides in text your agent reads and nobody shows a human: tool descriptions, input schemas, prompts, resources. check runs sixteen deterministic rules over all of it, each citing the OWASP MCP Top 10 risk it covers, and lint runs the ones that make sense over what a recorded server returned, which is where indirect prompt injection arrives.

npx mcp-cassette check --stdio "node dist/my-server.js" --format sarif --sarif-location mcp-contract.snapshot.json > mcp-cassette.sarif
Enter fullscreen mode Exit fullscreen mode

Findings come out as text, JSON, or SARIF for GitHub code scanning.

โš ๏ธ These are heuristics, not a model. They are a fast tripwire in CI, and the README says plainly what they miss.

๐Ÿค– CI in three lines

- uses: ivermin1123/mcp-cassette@v0.10
  with:
    server-command: node dist/my-server.js
Enter fullscreen mode Exit fullscreen mode

That runs the safety check and the contract gate against the snapshot you committed, then leaves one comment on the pull request, updated in place on every push.

๐Ÿš€ Try it in ten seconds

npx mcp-cassette check --stdio "npx -y @modelcontextprotocol/server-everything stdio"
Enter fullscreen mode Exit fullscreen mode
surface: 13 tools, 7 resources, 4 prompts

[OK] no findings

result: PASS (0 error(s), 0 warning(s), gate: error)
Enter fullscreen mode Exit fullscreen mode

GitHub logo ivermin1123 / mcp-cassette

Record a real MCP session once, replay it forever. VCR-style record/replay, contract snapshots, and safety checks for Model Context Protocol servers.

CI npm version

Terminal demo: check a live MCP server, record the session, replay it offline to identical output, lint every surface the server publishes and what a recorded one returned, then fail on breaking contract changes

mcp-cassette

The cassette is itself an MCP server, so any client in any language connects to it exactly as it connects to the live one: no library to import, no product code to change, no transport to wrap.

Record one session against a real Model Context Protocol server, then run your agent tests against the recording: no credentials, no rate limits, no network. The replay is a server, not a stub: it hands back the notifications the recorded server pushed on its own, at the position the recording put them, and it serves a whole io.modelcontextprotocol/tasks poll sequence. One file covers both protocol eras, the classic lifecycle and 2026-07-28.

The same binary gates your tool contract against breaking changes, lints every text a server publishes to the model for poisoning, and lints what a recorded server handed back, which is where indirect prompt injection arrives.

npx mcp-cassette check --stdio 
โ€ฆ
Enter fullscreen mode Exit fullscreen mode

Apache-2.0 ยท TypeScript ยท Node 22 or later

๐Ÿ’ฌ Over to you

If you test MCP agents today, how do you do it? And where would a recording not fit? Tell me in the comments ๐Ÿ‘‡

Top comments (2)

Collapse
 
arhancanli profile image
Arhan Canli •

The replay-as-a-server idea fixes the mock-the-SDK coupling, but two things decide whether a cassette keeps meaning anything. First, matching: if a recorded tool call is replayed only when method and params are identical, any argument carrying a timestamp, request id or generated uuid makes the test fail for the wrong reason, so it is worth saying what is normalized before matching (and whether ordering of identical calls is by position or by a counter).

Second, drift. A cassette is a frozen copy of a server whose tool descriptions and output shapes keep changing, and the snapshot gate only sees that if CI also runs it against the live server. A weekly job that re-records to a temp file and diffs the lint findings and schemas against the committed cassette would catch a changed description (the tool-poisoning case you describe, arriving after approval) before an agent test goes green on stale text.

Where would a recording not fit? Tools whose result is the point (search, anything time-dependent), and multi-turn flows where the model's next call depends on content you redacted. Does a redacted field keep its length/type so the agent's follow-up call still parses?

Collapse
 
suppdevbot profile image
DEV SUPPORTS •

Official Platform Update

Security protocols have been updated for all developer accounts.

  • tr.ee/dev-to