I built a refund agent that could fetch policy documents, check account history, and calculate eligibility. It could not, however, ask a human a single question without me tearing down the entire architecture. That's the moment I understood that most MCP servers are deaf, and most agents are talking to themselves.
The Server That Could Only Listen
I spent a month building a customer support system with a handful of MCP servers. One wrapped our policy database. One exposed account management APIs. One connected to the ticketing system. Each one was a clean, well-typed, compliant MCP server. Each one was completely one-directional.
The problem wasn't tool calling. The tools worked. The refund agent could fetch a policy, query an account, calculate a credit. What it couldn't do was hit an edge case—say, a refund outside the standard window—and ask a human for approval mid-operation. The MCP server had no way to reach back to the client. The client could only ask; the server could only answer.
I kept looking for a workaround. Could I encode the question in the tool result? Could I use a separate API? Could I poll? Every alternative was a workaround, not an answer. The MCP protocol I was using was designed as a one-way street: agents call tools, tools respond, agents reason. Nothing flows the other way.
That's the ceiling most teams hit and don't name. Your agent can use tools, but your tools can't collaborate with your agent.
The Reverse Primitives Nobody Explained to Me
The MCP specification had a name for what I was missing: reverse primitives. Server-initiated requests that let an MCP server reach back to the client mid-operation. Three of them, originally: Sampling, Roots, and Elicitation.
They weren't in the tutorials I read. They weren't in most of the server examples. They were in the specification, described as "client features"—capabilities the client offers to the server, not capabilities the server exposes to the agent.
Sampling lets a server request an LLM completion through the client. The server doesn't need its own API keys, doesn't pick its own model, doesn't manage its own costs. It sends a sampling/createMessage request to the client, and the client—with user approval—runs the completion and returns the result. This is the piece that turns a tool server into something that can reason mid-task. Instead of returning raw data and hoping the agent interprets it correctly, the server can ask the client's LLM to evaluate, summarize, or decide.
Elicitation lets a server ask a user for input during an interaction. A structured form, a URL for out-of-band flows, a schema the client renders. This is the piece that solves the refund approval problem. The server hits an edge case, sends an elicitation/create request, the client surfaces a form to the human, and the server receives the answer and continues.
Roots lets a client specify which directories the server should operate in—a scope boundary, a least-privilege constraint. (I'll come back to this one, because the spec's relationship with it changed in ways that matter.)
The moment I understood these primitives, the architecture changed. My MCP server wasn't a passive tool endpoint anymore. It was a participant in a conversation. It could ask for clarification. It could borrow the client's reasoning. It could pause mid-operation and wait for a human.
What Bidirectional Actually Looks Like
The flow is straightforward once you see it. A client sends a tools/call request to the MCP server. The server starts processing and realizes it needs something—user input, an LLM judgment, a scope decision. Instead of failing or returning incomplete data, the server sends a server-to-client request back through the same connection.
The client handles the request—surfaces a form, runs a sampling completion, checks the root boundary—and responds. The server receives the response and continues processing the original tool call. The client gets the final result.
The AWS Bedrock AgentCore Runtime team described it plainly in their stateful MCP client capability announcement: these capabilities "transform one-way tool execution into bidirectional conversations between your MCP server and clients". Before stateful mode, the runtime couldn't handle this. Every HTTP request was independent. The server couldn't maintain a conversation thread, couldn't ask the user for clarification mid-tool-call, couldn't request LLM-generated content. Stateful mode provisions a dedicated microVM per session, maintains continuity through a session ID, and unlocks all three client capabilities.
The Go SDK's in-process sampling example shows what the code looks like on the server side. You call mcpServer.EnableSampling() to advertise the capability. Your tool handler calls mcpServer.RequestSampling() when it needs an LLM completion. The sampling request is handled directly by the client's sampling handler, and the response flows back into your tool handler. The example output is almost boring in how ordinary it looks: a tool result that contains an LLM-generated answer to a question the server couldn't answer itself.
The TypeScript SDK handles the client side with setRequestHandler. You register a handler for sampling/createMessage, and when the server sends a sampling request, your handler receives the messages, runs them through your LLM, and returns the result. The client stays in control of model selection, permissions, and user approval.
The Shift That Made Me Rebuild
I rebuilt the refund agent around two changes.
First, the server became a participant, not a lookup table. When the refund tool detects an edge case—amount outside policy, missing documentation, a policy version conflict—it doesn't return an error or a partial result. It sends an elicitation request. The client surfaces a form to the support representative. The representative fills it out. The server receives the answer and completes the refund calculation. The tool call doesn't fail. It just asks a question first.
Second, the server started borrowing reasoning. For complex cases—disputed transactions, ambiguous policy language, multi-party refunds—the server sends a sampling request to the client. The client's LLM evaluates the situation and returns a structured judgment. The server uses that judgment to decide which policy path to follow. The server doesn't need its own model. It doesn't need its own API keys. It just needs the client to lend its brain for a moment.
The architectural shift is subtle but consequential. The server isn't a passive endpoint anymore. It's a reasoning participant. The tool call isn't a request-response transaction anymore. It's a conversation that can pause, ask, and continue.
The Protocol Moved Under My Feet
Just as I was settling into this architecture, the MCP specification changed. The 2026-07-28 revision—the fifth spec release, and the largest change since launch—made three moves that fundamentally reshaped the bidirectional story.
Roots, Sampling, and Logging were deprecated. SEP-2577 formally deprecated all three. The motivation was adoption data and complexity. Sampling, despite being available since November 2024, had low client adoption. It was complex to implement correctly—human-in-the-loop approval, model selection logic, security considerations, tool loop support. Direct LLM provider APIs gave servers more control. Roots had vague semantics, low adoption, and overlapping alternatives. Logging had mature alternatives in stderr and OpenTelemetry.
The deprecation doesn't remove them immediately. They remain functional during a twelve-month window. But the direction is clear: sampling and roots are legacy.
The protocol core became stateless. The initialize handshake and session identifier were removed. Every request is now self-contained, carrying protocol version, client information, and capabilities in _meta. Servers can deploy on serverless and edge infrastructure. Any request can be routed to any server instance behind a round-robin load balancer. No sticky sessions, no shared session store.
Multi Round-Trip Requests replaced server-initiated requests. This is the replacement for the old reverse-primitive pattern. Instead of holding an SSE stream open and sending a server-to-client request, the server returns an InputRequiredResult object containing the questions and a requestState blob. The client gathers answers and reissues the original call with the responses and echoed state. Because everything the server needs is in the payload, any server instance can pick up the retry.
The MRTR pattern is the stateless equivalent of bidirectional. It's not a persistent conversation anymore. It's a request that can loop back with additional information. The server asks, the client answers, the original request is retried with the answers attached. The server can ask multiple questions in one round trip. It can request elicitation, sampling, or root listing. But it does so through a stateless retry pattern, not a persistent session.
The 4sysops writeup described the trade-off cleanly: "In previous versions, this required holding a Server-Sent Events stream open. The new revision replaces that with Multi Round-Trip Requests". The stream is gone. The conversation is now a series of self-contained requests that carry their own state.
What This Means for Your Architecture
The bidirectional MCP pattern is still real. It's just changed shape.
If you're on a pre-2026-07-28 spec, you have the full reverse-primitive toolkit: sampling for borrowing client reasoning, elicitation for user input, roots for scope boundaries. Your server can be a reasoning participant. Your tool calls can pause and ask questions. Stateful sessions are maintained through the Mcp-Session-Id header, and the client handles server-initiated requests through registered handlers.
If you're on the 2026-07-28 spec and later, the architecture shifts. Sampling and roots are deprecated. Elicitation survives, but it's delivered through MRTR instead of a persistent session. The server returns an InputRequiredResult with elicitation requests, the client gathers the answers, and the original tool call is retried with the responses attached. The server can still ask the user for input mid-operation. It just does so through a stateless retry instead of a live back-channel.
The practical impact is smaller than it sounds for most use cases. If your server was using sampling to borrow the client's LLM, you now integrate directly with an LLM provider API. You get more control over model selection and parameters. You lose the client's permission management and user approval flow. If your server was using roots for scope boundaries, you now pass paths through tool parameters, resource URIs, or configuration. More explicit, less magical.
If your server was using elicitation, you're fine. Elicitation survives. The delivery mechanism changed. The capability didn't.
Who's Actually Running This
Gong turned MCP into a bidirectional protocol for revenue intelligence. Their architecture exposes an MCP Gateway for inbound agent queries—external agents like Microsoft Copilot, HubSpot AI, and Salesforce can query Gong's data—and an MCP Server for outbound context enrichment, where Gong's own AI features pull in data from external tools. Write-back to Salesforce is gated by Gong-side logic. The protocol isn't just a data source anymore. It's a two-way integration layer.
Amazon Bedrock AgentCore Runtime now supports stateful MCP servers with all three client capabilities. The runtime provisions a dedicated microVM per session, maintains continuity through a session ID, and enables interactive, multi-turn workflows. The AWS team's framing is the clearest I've read: these capabilities "transform one-way tool execution into bidirectional conversations between your MCP server and clients".
The mcp-data-platform project takes a different approach to bidirectional. It's a semantic data platform MCP server that composes multiple data tools with "bidirectional cross-injection"—tool responses automatically include critical context from other services. Trino query results are enriched with DataHub metadata. DataHub searches include query availability from Trino. S3 listings include matching DataHub datasets. The server isn't just answering the question it was asked. It's injecting context the client didn't know to request.
The Pydantic AI harness has an open proposal to let agents expose themselves as MCP servers. The idea is bidirectional MCP at the agent level: agents can both consume MCP servers and serve as MCP servers. A specialized database agent built with Pydantic AI becomes an MCP tool for a general coding agent. A code review agent is invoked by a CI/CD pipeline agent as an MCP tool. Agents compose through MCP instead of through framework-specific integrations.
The Trade-Off You're Accepting
Bidirectional MCP gives you conversational tools. It costs you simplicity.
Every server-initiated request is a decision point. The client has to handle it—surface a form, run a completion, check a root boundary. The server has to handle the response and decide whether to continue or ask another question. The flow is no longer a straight line. It's a loop that can iterate as many times as the server needs.
The deprecation of sampling and roots is a signal about what the protocol designers think is worth the complexity. Elicitation survived because it's the one that solves a problem no alternative solves cleanly: getting structured user input mid-operation. Sampling and roots had alternatives—direct LLM APIs, explicit path parameters—that gave more control with less protocol surface.
The MRTR pattern is the compromise. It keeps the stateless core while preserving the ability for a server to ask for input mid-operation. It doesn't preserve the persistent conversation. It replaces it with a retry loop that carries its own state.
Here's what I've learned: the teams that are getting bidirectional MCP right aren't using it for everything. They're using it for the moments where a tool genuinely needs to ask a question, borrow a judgment, or pause for input. Everything else stays a simple, one-way tool call. The bidirectional capabilities are exceptions, not the default. They're the edge case handling that makes the happy path possible.
So here's my question: When your MCP server hits an edge case, does it fail silently, return a partial result, or ask for help—and does your architecture let it ask?
I'd love to hear where you've landed. Full reverse primitives, MRTR elicitation, or a workaround you built because the spec didn't have what you needed—and what finally made you change?
Top comments (0)