DEV Community

Sidney da S. P. Bissoli
Sidney da S. P. Bissoli

Posted on AI-assisted

Your MCP server changed. Its version didn't. Here's how to catch it.

A version number is a promise: nothing a client depends on changed without me saying so. For MCP servers, almost nothing checks that promise.

The copy of a server in the MCP Registry carries a name, a version, the packages and the remotes. It carries no surface: no tools, no instructions, nothing about what answers without a token. So anyone comparing the registry with the live server can compare exactly one thing, the version. That comparison is only worth something if every surface change bumps the version.

This post is about turning that "if" into a failing test, and then about what happened when I ran the test backwards over the whole history of my seven servers: 142 published versions.

Where this came from

It started in the comments of my previous post, in a thread with @yahhi. The registry listing and the live worklore server both said 0.1.0, so a version diff passed cleanly. But, as @yahhi found, the server had moved on: tools/list now answered without a token, and the auth rules had moved into the instructions. The drift was one level deeper than the version.

Next came a working check, and a replay of the surface from the day of the listing against the current code. The registry said 0.1.0 for a server with four read-only tools. The server behind that number had six, two of which write on the user's behalf.

I said I'd build the same thing for my servers. It's now live on all seven, and packaged as @sbissoli/mcp-surface.

1. Lock what the client sees, not what the code says

The lock file, surface.lock.json, sits at the root of each server, next to package.json. It has two sections. Each stores the version it was locked under and a sha256 of its content.

  • Declared surface: the initialize result (instructions and capabilities included, serverInfo without the version), plus tools/list, resources/list, resources/templates/list and prompts/list. Keys are sorted and lists are ordered by name or URI, so the hash only moves when the surface does. A method the server doesn't serve is recorded as null, not []: "no prompts" and "prompts/list isn't served" are different surfaces.
  • No-token behaviour: which methods answer without a credential, on every MCP route, with the API key unset and set. No listing shows this, and it's exactly what had drifted in that server.

The declared surface is captured from the server factory in-process. The no-token section is measured at the HTTP edge, by calling the Worker's fetch with and without the key.

2. Make the version bump the only way out

The rule is short:

  • measured ≠ locked and the version is the same → the test fails;
  • the version changed → the test fails and asks you to re-lock.

The script that rewrites the lock follows the same rule: it refuses to lock a new surface under the old version. You can't "fix" a red test by regenerating the lock. You bump the version, re-lock, and commit the lock with the bump:

npm version minor --no-git-tag-version
npm run surface:lock   # build + run the surface tests in write mode
Enter fullscreen mode Exit fullscreen mode

The deploy runs npm test before wrangler, and the npm publish runs it before npm publish. So a surface change without a version change can't be deployed or published.

One side effect, and it's on purpose: an SDK upgrade that changes capabilities also turns the lock red. The client sees a different surface, so the version should say so.

3. Green tests prove the code; a post-deploy check proves what's served

Tests run against the code in the repository. The client talks to whatever is actually deployed. Usually those are the same thing. Wrangler configs, build steps, environment variables and caches are where they stop being the same.

So every deploy now ends with one more step that queries the live endpoint and compares it with the lock:

npx mcp-surface verificar https://<host>/mcp --tool <a tool that needs no network>
Enter fullscreen mode Exit fullscreen mode

If what's being served isn't what was locked, the deploy job goes red after the fact, which is still much better than never.

4. The lock only protects you from now on, so replay the history

A lock written today says nothing about the versions that came before it. For those there's a replay:

npx mcp-surface replay --url https://<host>/mcp
Enter fullscreen mode Exit fullscreen mode

For every version published on npm, it installs the package in a temporary directory (--ignore-scripts), starts it over stdio and captures the surface with the same normalisation the lock uses. Then it compares each version with the previous one, lists every removal that didn't come with a major bump, and checks the live endpoint against the surface of the version its /status reports. The output is a Markdown table and a JSON file in baselines/, committed to the repo.

Across the seven servers that came to 142 versions. In all seven, the live server matches the version it reports. It also found three things I didn't know.

5. Two releases that don't even start

Two old releases of my Senate server, 1.1.0 and 1.1.2, crash on startup:

Error: Dynamic require of "events" is not supported
Enter fullscreen mode Exit fullscreen mode

The esbuild bundle emitted ESM that still contained a CommonJS require for a Node built-in. Nobody had noticed, me included, because nobody installs a two-version-old release on purpose. But npm would have offered both as valid versions indefinitely. Any version diff would have called them fine forever, because their version numbers are perfectly well-formed. They're deprecated now.

The lesson: "the version string is right" and "the version runs" are different claims, and only one of them shows up in a registry.

6. A minor release that removed six tools

My medical terminology server went from 1.0.2 to 1.1.x and six SNOMED CT tools disappeared from the default surface. They moved behind a licensing flag, off by default.

The CHANGELOG for that release even called it "technically breaking". It shipped as a minor anyway. There was a reason, since SNOMED access depends on a licence the user has to hold. But the client that was calling snomed_search yesterday doesn't read my reasoning. It gets "unknown tool".

The replay lists cases like this without judging them. Some removals have a good reason, and the CHANGELOG is where that reason lives. What matters is that the list exists, so the decision is made on purpose rather than discovered later by a user.

7. The replay was wrong once, too

The first version of the replay flagged ilo-mcp-server 0.5.0 → 0.6.0, where a parameter became required, as a breaking change outside a major release. But under semver, 0.x is allowed to break in a minor. The line of compatibility in 0.x is the minor, not the major.

The rule is fixed now: in 0.x, breaks are compared within the same minor. I mention it because a tool that reports drift has to earn its alarms. One false alarm on day one teaches people to ignore the next real one.

8. Check the probe before you write the lock

This is the one I'd most like other people to avoid.

In one server (sih-br-mcp) the MCP backend sits behind an edge proxy, so the no-token test can't call the real handler in-process. My first test double for that backend answered result to every method. The lock was written from it, and it faithfully recorded "prompts/list answers" for a server that has no prompts.

All the tests passed, because the tests were checking the lock against the same lying double. The post-deploy check against the live endpoint is what caught it.

The fix has two parts. A double that answers everything is not a double of your server, so where the real handler can run in-process, run it. And the probe's sanity checks (a method that must answer, a method that must not) now run before anything is written to the lock, not after. A lock written from a broken probe is worse than no lock, because it makes a wrong surface look verified.

What I would do first, next time

  1. Lock the surface next to the version from the first release: initialize, the four lists, and which methods answer without a token.
  2. Make the lock script refuse to rewrite under the old version. Otherwise the red test just becomes a ritual.
  3. Check the live endpoint after every deploy. Green tests prove the code; only this proves what's served.
  4. Run the replay once over your whole history. It's cheap, it's one command, and the first run is the one that finds things.
  5. Prove the probe can say "no" before you trust anything it says "yes" to.

The package is MIT: @sbissoli/mcp-surface, source in mcp-br-commons. Fair warning: the API and CLI verbs are in Portuguese (travar = lock, verificar = verify), and the roadmap follows my own servers' needs.

Thanks to @yahhi for the thread that started all this. There is also a write-up of the other side of it, in Russian, on Habr. One comment thread changed how eight servers ship.

Top comments (2)

Collapse
 
_firelinks profile image
Mike Dabydeen •

The lock makes the publisher keep the promise, but a client still can't check it, because the hash lives in your repo and the registry entry carries only the version. I'd publish the section hashes with each release. server.json already has room for it: the registry accepts custom metadata under _meta, in the io.modelcontextprotocol.registry/publisher-provided key. An agent host could then compute the same normalised hash from initialize and tools/list on first connect, and refuse or ask for re-approval when it differs from what the registry lists for that version.

That would have caught the instructions move in the worklore case from the client side, without anyone running your replay. It does depend on the normalisation being written down somewhere other than your code, so a host built by someone else hashes the same bytes. A short spec of the canonical form in the mcp-surface README would cover it.

Collapse
 
reidmarlow profile image
Reid Marlow •

The test double returning valid results for every method in section 8 is the exact failure mode that bites snapshot-based contract tests. If the harness cannot prove an invalid method or missing capability produces an explicit protocol error, the snapshot only records whatever the mock felt like emitting.

On the agent client side, the nastiest drift I hit with MCP servers is silent edits to tool descriptions. Under conventional semver, rewording a description looks like a no-op documentation tweak that never warrants a minor bump. Because agent runners inject tool definitions directly into the system prompt, a one-word change in a description busts prompt cache prefixes across active sessions and can shift tool selection behavior without a single parameter changing. Locking description hashes alongside schema definitions saves hours of debugging mysteriously evaporated cache hit rates.