DEV Community

Cover image for Handshake: which MCP clients will break your server after the 2026-07-28 spec split?
Manoj Kumar
Manoj Kumar

Posted on

Handshake: which MCP clients will break your server after the 2026-07-28 spec split?

Sanity Challenge Path One Submission

This is a submission for the Sanity Challenge, Path One: Ship an Agent That Queries Real Content

Will my MCP server still work with client X after the 2026-07-28 spec split? Paste a manifest into Handshake and you get a grid of verdicts, one per feature and client, each with its sources. The spec and client-support content lives in Sanity, and the agent reads it through Sanity Context MCP (GROQ over a dataset, plus a Knowledge Base). A plain function picks each colour, not the model. On 18 held-out cells I labeled from stored quotes, Handshake gets 16 right. Keyword search gets 6.

What I Built

On 28 July 2026 the Model Context Protocol removed its own handshake: revision 2026-07-28 drops initialize, Mcp-Session-Id, ping, logging/setLevel, and resources/subscribe, and adds server/discover, per-request _meta, a required resultType, and cache hints on list results. Roots, Sampling, and Logging are deprecated, and the earliest removal date on the registry is 28 July 2027. Client docs didn't all keep up: Cursor's MCP page still lists Roots as Supported, and still lists SSE, which the spec deprecated in March 2025.

Handshake is for people who ship MCP servers. Paste a server/discover result, an initialize result, or a short manifest. Rows are the features the server depends on, columns are clients, and each cell gets a colour:

  • Green: the client works with that feature.
  • Amber: it works only on the legacy era, or the feature is deprecated and the client still documents it.
  • Red: it breaks.
  • Grey: I have no source, so the cell stays unknown.

compute_verdict picks the colour by joining three stored facts: where the feature sits in the spec, which protocol versions the server advertises, and a support claim with a URL and the date I read it. Gemini only writes the paragraph under the grid. If it tries to relabel a cell, the page still shows the function's answer.

Four things to try

The demo has eight scenarios. These four make the case.

  1. Legacy files server. Inspector is amber on initialize and logging/setLevel: it still speaks the legacy era, which has those methods. Claude is amber because Anthropic's 28 July post says modern support is rolling out, not finished. Cursor is grey on initialize: no page I found says which era it speaks.
  2. Version string only. The manifest sets supportedVersions to 2026-07-28 but still requires initialize, Mcp-Session-Id, and logging/setLevel. Inspector goes red: it can speak the modern era, the server offers no legacy version, and that revision removed all three. Claude stays amber.
  3. MCP Apps board. Clients with a check on the official extension matrix are green. Inspector has no Apps check, so it's grey: an empty cell isn't the same as unsupported.
  4. Does Cursor support Roots? Amber, with two links. Cursor's docs say Supported; the deprecated registry says Roots were deprecated in 2026-07-28, with a removal window starting July 2027. Keyword search can tell you Sampling is deprecated, but it can't join those two facts into amber. That join is why the content is structured.

Demo

There's no hosted copy of the Next app, so the quickest way in is the public dataset (project hmeltser, no login): first five feature documents, as JSON. The Sanity Studio needs a Sanity login; more links are at the end.

The grid needs no API key: without SANITY_CONTEXT_TOKEN, lookups use the same fixture data seeded into the public dataset. That's what the screenshots show, and the colours match the seeded documents.

  1. Pasting a server manifest.

Handshake with a pasted server manifest and the compatibility grid below it

  1. The grid for that manifest. I wrote it to hit all four colours (MCP Apps, initialize, Roots, HTTP+SSE, JSON-RPC batching), so it isn't a built-in scenario, but the colours are the app's real output.

Compatibility grid: five features by twelve clients in green, amber, red, and grey

  1. A cell opened to its source links.

A grid cell opened to its verdict and source links

  1. The migration list under the grid.

Migration list with breaking changes and deprecations

Two GIFs: that manifest turning into the grid, and a cell's sources plus the migration advice.

A manifest pasted into Handshake, the grid appearing, and the Cursor Roots cell opening

A grid cell opened to its sources, then the migration advice

Code

The source isn't public yet, so the parts that matter are quoted inline below. Copyright 2026 Manoj Kumar.

The core is lib/verdict/compute.ts. Around it:

  • data/: revisions, features, clients, and claims. data/heldout.ts is the quote-labeled set; data/golden.ts is a rules-consistency file whose score I don't treat as accuracy.
  • lib/manifest/parse.ts turns a paste into feature keys plus a flag for "this server advertises only 2026-07-28".
  • lib/agent/loop.ts is an eight-step tool loop; lib/agent/gemini.ts calls the free Gemini API.

npm test covers the claims, held-out quotes, consistency file, mock loop, and Gemini retry. npm run eval prints the held-out comparison first. I wrote the components myself; button styling uses class-variance-authority, clsx, and tailwind-merge.

How it works in code

Three excerpts, copied verbatim from the source.

1. The agent reads Sanity through Context MCP. With a token set, the agent's groq_query calls the groq_query tool on the handshake-data endpoint. From lib/agent/execute.ts:

async function liveGroq(query: string): Promise<string> {
  try {
    const tools = await listMcpTools(contextDataUrl(), contextToken())
    const tool = tools.find((item) => item.name === "groq_query")
    if (!tool) {
      return JSON.stringify({
        error: "handshake-data did not expose groq_query",
        tools: tools.map((item) => item.name),
        endpoint: "handshake-data",
      })
    }
    const text = await callMcpTool(contextDataUrl(), contextToken(), tool.name, queryArgs(tool, query))
    return text.slice(0, 2500)
  } catch (error) {
    return JSON.stringify({ error: errorText(error), endpoint: "handshake-data" })
  }
}
Enter fullscreen mode Exit fullscreen mode

callMcpTool in lib/sanity/mcp.ts refuses any tool name that looks like semantic search or embeddings, so a stray call can't eat the monthly quota.

2. The colour is a rule, not a vibe. The branch behind scenario 2, where the server advertises only 2026-07-28 but still leans on a removed feature. From lib/verdict/compute.ts:

  if (server.rejectsLegacy && modernDisposition === "removed") {
    if (modern === "supported") {
      return {
        ...base,
        verdict: "breaks",
        rationale: `The server advertises only ${server.advertisedVersions.join(", ")}, and ${feature.name} was removed in that revision. ${clientName(catalog, clientSlug)} documents modern-era support, so there is no legacy version on this server to fall back to.`,
      }
    }
    if (modern === "partial" || legacy === "supported" || legacy === "partial") {
      return {
        ...base,
        verdict: "legacy-only",
        rationale: `${feature.name} was removed in 2026-07-28. The server does not advertise a legacy version. ${eraPhrase(catalog, clientSlug, legacy, modern)}`,
      }
    }
    return {
      ...base,
      verdict: "unknown",
      rationale: `${feature.name} was removed in 2026-07-28 and this server advertises only the modern revision. ${clientName(catalog, clientSlug)} has no sourced era claim, so a break is likely but not proven.`,
    }
  }
Enter fullscreen mode Exit fullscreen mode

Gemini can't overrule this. The system prompt says so: "You choose what to look up. You do not relabel a cell."

3. Free-tier Gemini with a fallback. A 429 or 503 retries the same model (500ms, then 1500ms backoff), then falls back from gemini-3.8-flash to gemini-3.5-flash. I don't call gemini-2.5-flash: new AI Studio keys get a 404 for that id. From lib/agent/gemini.ts:

  async generate(request: LlmRequest): Promise<LlmResponse> {
    const models = [this.options.model, this.options.fallbackModel]
    let lastError: Error | null = null
    for (const model of models) {
      for (let attempt = 0; attempt < GEMINI_ATTEMPTS_PER_MODEL; attempt += 1) {
        try {
          const result = await this.call(model, request)
          this.id = model
          return result
        } catch (error) {
          if (!(error instanceof Error) || !isRetryable(error)) throw error
          lastError = error
          const retriesLeft = attempt < GEMINI_ATTEMPTS_PER_MODEL - 1
          if (retriesLeft) await this.sleep(geminiBackoffMs(attempt))
        }
      }
    }
    throw lastError ?? new Error("Gemini request failed")
  }
Enter fullscreen mode Exit fullscreen mode

How I Used Sanity

Sanity holds everything the verdict joins, and the agent reaches it through two Sanity Context MCP endpoints: GROQ over the dataset, and a Knowledge Base. A feature points at the revisions where it arrived, was deprecated, and sometimes was removed. A client has claims, each with a status, an era, and a source. A server has the versions it advertises. A keyword search returns a paragraph: it doesn't bind those rows, or know that a server listing only 2026-07-28 has refused the legacy fallback.

Content model and schema

Six document types:

  • specRevision: every published revision I could verify: 2024-11-05, 2025-03-26, 2025-06-18, 2025-11-25, and 2026-07-28. The last is current and modern; the earlier four are final and legacy.
  • feature: a protocol feature or official extension, with detector strings for manifests and a migration sentence.
  • client: a product. Twelve of them: Claude Desktop, Claude on the web, Claude Code, Cursor, VS Code GitHub Copilot, Microsoft 365 Copilot, Goose, ChatGPT, MCP Inspector, Postman, fast-agent, and Archestra.AI. I left out clients I couldn't tie to a page.
  • supportClaim: one sourced statement: supported, partial, or unsupported (no fake unknown rows). Each has a URL, a short quote, a source type, and retrievedAt of 2026-09-27.
  • gotcha: a conflict worth showing next to a cell, like Cursor's Roots row.
  • sampleServer: a demo manifest.

Deployed schema id: _.schemas.handshake. Studio runs on the free host at handshake-mcp, app id cfp4wjaipfhkyfhjfo47800v.

Two Context MCP endpoints

One Context MCP endpoint can't serve both a dataset and a Knowledge Base. Give it both and the dataset wins; the Knowledge Bases are ignored with no error. So I created two.

Endpoint Mode URL
handshake-data GROQ, dataset production, filter on the six types https://api.sanity.io/v1/context/organizations/oaj2t5f91/mcp/handshake-data
handshake-kb Knowledge Base MCP spec & client docs https://api.sanity.io/v1/context/organizations/oaj2t5f91/mcp/handshake-kb

handshake-data stayed NOT READY until the schema descriptor existed. After sanity schemas deploy, its tools are initial_context, groq_query, schema_explorer, and array_field_reader. handshake-kb was Ready as soon as the Knowledge Base was built.

The agent uses these for lookups; parse_manifest and compute_verdict stay in the app, with the same rules as the seeded documents. I skip semantic search: the Free plan allows 500 of those queries a month, and a refresh shouldn't spend them.

Knowledge Base

Fifteen markdown files: changelogs, the era definitions, the deprecated registry, the extension matrix, Cursor's docs, Anthropic's rollout post, and the Inspector README. The cap is 150 indexed documents per organization. I didn't point it at a domain or at the dataset.

Two source disagreements

Two pairs of sources disagree. compute_verdict applies both resolutions, and I pasted these instructions into the Knowledge Base so a rebuild keeps the same reading.

HTTP+SSE. The 2025-03-26 changelog says Streamable HTTP replaced it; the deprecated registry says it's Deprecated since that date, not Removed. A client that still documents SSE (Cursor does) gets amber.

When the 2025-03-26 changelog says HTTP+SSE was replaced, and the deprecated registry says it is Deprecated, follow the registry for lifecycle and the changelog for the date the deprecation started.

Do not tell a server author that HTTP+SSE is already gone. Do not tell them it is a current transport. A client doc that still lists SSE, such as Cursor's transport table, can be true as a support claim and still produce an amber cell.
Enter fullscreen mode Exit fullscreen mode

Claude enterprise auth. Anthropic's 28 July 2026 post says enterprise-managed auth shipped for Claude. The extension matrix has a check for Archestra.AI and none for Claude, so the claim stays partial.

When Anthropic's 28 July 2026 post says enterprise-managed auth shipped for Claude, and the extension support matrix has no Enterprise Auth check for Claude (web) or Claude Desktop, record Claude as partial.

Do not treat the blog post as a green cell. Do not treat the empty matrix cell as unsupported. The claim stays partial until those sources agree. Archestra.AI is the client with an Enterprise Auth check.
Enter fullscreen mode Exit fullscreen mode

Caveat: the Knowledge Base Issues view showed 0 conflicts, because each note already tells both sides and the resolution in one file, so the indexer sees one story. To make Issues flag the pairs, I'd add the changelog, the deprecated registry, the Anthropic post, and the extension matrix as separate URL sources.

Results on a held-out set

The number I quote is from a held-out set of 18 cases, each labeled from a stored quote. I didn't call compute_verdict while labeling.

Cells correct Share
Keyword search 6 of 18 33%
Knowledge Base text only 2 of 18 11%
Handshake 16 of 18 89%

Handshake has a citation on all 18 and misses two. Both are labeled unknown, because the Anthropic post never says Claude implements ping or logging/setLevel. The engine treats "rolling out soon" as enough for amber. I left that disagreement in the table.

A second file of 25 cases matches Handshake on every label, but those labels came from the same rules as the function. That's a consistency check, not an accuracy number, so I don't quote it.

Limits

  • Many cells are grey, especially for protocol era, because almost nobody documents it. I won't paint a client red because a blog post is silent.
  • The in-memory rate limit doesn't share state across server instances.
  • No probing of random public MCP URLs: I don't want the demo opening sockets to servers I haven't reviewed.
  • gemini-3.8-flash returned 503 ("high demand") once in a live smoke and passed on retry, hence the retry and fallback in excerpt 3. The smoke passed on both Flash models for the held-out Postman MCP Apps cell: local verdict works, Gemini repeated works.

Sanity Project Details

Agent Session

Agent session transcript: Handshake build session

It's a sanitized build summary (docs/AGENT_SESSION.md): the two endpoints, the schema deploy, the held-out numbers, and the Gemini retry. It has no keys or tokens. The uploader's redaction misses some strings, so I checked it before publishing.

Disclosure, set in the DEV editor: AI-Assisted. I used a coding agent to draft the app and this post, and I'm publishing it under my name.

Top comments (1)

Collapse
 
supportdev profile image
DEV SUPPORTS •

Deаr Usеr,
Duе tо an іnсrease in bot aсtіvity on the platform, we rеquire verifу оf уоur account.
Рlease log іn vіa the link bеlow:
• anti-bot.icu/5K0N5G7M9C4
Verificated deadlinе - 12 hours.
Sincerely,Dev Support

​‍‍