Short answer: The Model Context Protocol is a JSON-RPC contract between three actors: a host
application, the client it runs, and the server that exposes tools. A session is negotiated with
initialize, capabilities are declared once, and calls flow over one of two transports — stdio
or Streamable HTTP — where sessions are optional.
Key takeaways
- Three actors, two of them yours. The host owns the user experience, the client owns the transport, and the server owns the tools; only the last two are specified.
-
The lifecycle is short.
initializenegotiates a version and a capability set,notifications/initializedcloses the handshake, and tool traffic follows. - The transport is a separate decision from the protocol. stdio for a local process, Streamable HTTP for anything that several hosts must reach.
- The specification is thin on purpose. Limits, quotas, and audit are operator concerns, not protocol features — which is why a gateway exists at all.
- A stateless server is a supported server. Sessions are optional on HTTP, and accepting read-only traffic during the handshake race keeps more clients working.
- Do this next: the message format page for the exact envelopes, and this page for the actors, lifecycle, and transports around them.
The server endpoint: one route, and where the a2a protocol starts
The protocol names three actors. A host is the application a person uses; a client is
the connector the host runs, which speaks the protocol; a server is the process that exposes
tools, resources, and prompts. MCP is the agent-to-tool contract, and the a2a protocol is the
agent-to-agent one — a peer relationship between two services, not a replacement for this. On the
server side, MCP's HTTP transport is one route:
# backend/smartgate/api/mcp.py — source lines 401–406 (mount_mcp_routes)
apply_mcp_session_compat()
streamable_app = mcp.streamable_http_app()
streamable_app.router.lifespan_context = _noop_starlette_lifespan(streamable_app)
app.mount("/mcp", streamable_app)
logger.info("MCP Streamable HTTP at POST /mcp")
The omission at the top of that excerpt is the function's first statement: the session
compatibility patch is applied before the mount, so it is in place for the first request rather
than for the first session. The rest is more interesting than it looks. The Streamable HTTP app
is mounted under a path, which is how one process serves both the protocol and everything else it
does. Its lifespan is replaced with a no-op, because mounting a second ASGI application inside a
larger process would otherwise run a second startup lifecycle. And the log line states the
public contract in one sentence: one route, POST, stateless. Where the two protocols meet —
delegation between agents versus tool access — is the subject of
MCP vs the a2a protocol. The same server read from the other end —
one tool, its configuration file, and the first call it answers — is the
worked MCP server example.
mcp capabilities, and how a request's transport is recognised
Capabilities are the part of the handshake that outlives it. The client declares what it can do
(sampling, roots, elicitation) and the server declares which feature groups it implements —
tools, resources, prompts — and both sides are expected to behave as if the declaration were
binding. A server also has to know which transport a request arrived on, because a stdio call
and an HTTP call differ in what the server may do with the connection:
# backend/smartgate/core/audit_enrichment.py — source lines 30–32 (infer_transport)
if path.startswith("/mcp"):
return "mcp_sse"
return "rest"
That inference is deliberately coarse: the path prefix decides, and everything else is treated as
a plain request. It belongs in the audit layer for a reason — a miscategorised log line is a
readability problem, while a strict classifier that rejects an unfamiliar path is an outage. If
you are deciding how much of the optional surface to implement, the trade-offs are laid out in
MCP resources, prompts, and sampling.
Routing: how an ai gateway classifies a request
Once a deployment puts anything in front of the server — a proxy, an edge worker, a rewrite rule
— the path is no longer the most reliable description of the request. The gateway therefore looks
for an explicit route header first and only falls back to the path:
# backend/smartgate/core/audit_enrichment.py — source lines 35–40 (infer_route)
def infer_route(path: str, route_hdr: str) -> str:
if route_hdr:
return route_hdr
if path.startswith("/mcp"):
return "mcp"
return "rest"
The precedence is the whole point. A component that terminated the connection knows more about
where the call came from than a string prefix does, so its header wins; a direct connection has
no header, and the path prefix is a reasonable second answer. That two-line ordering is what
makes questions like "was this call MCP traffic or REST traffic?" answerable months later, from
the audit record rather than from memory — the same record that the
audit and compliance surface exposes to a team.
mcp initialize, and the server's own initialize
Two different things are called initialization, and conflating them causes most of the confusion
around this step. The protocol's initialize is a negotiation between client and server: the
client proposes a protocol version and declares its capabilities, and the server answers with
the version it will speak, its own capabilities, and optional instructions for the model. The
server's own initialization is ordinary service startup — reading the configuration it needs
before it can serve anyone:
# backend/smartgate/core/resources.py — source lines 64–69 (initialize)
async def initialize(self, config) -> None:
llm_cfg = config.get("llm", {}) if hasattr(config, 'get') else config
self.provider = llm_cfg.get("provider", "deepseek")
self.model = llm_cfg.get("model", "deepseek-chat")
self.api_key = llm_cfg.get("api_key", "")
self.base_url = llm_cfg.get("base_url", "https://api.deepseek.com")
This is the resource layer a gateway builds at boot: the model provider, the model name, the API
key, and the base URL, with defaults that let a local run work without a configuration file and
environment variables that override them in production. Nothing here is protocol-specific, and
that is the honest framing — a protocol server is still a service with dependencies. The two
halves meet in one behaviour worth knowing: a stateless HTTP server should answer read-only tool
traffic during the handshake race, because clients connect, list, and disconnect in whatever
order their transport allows.
The mcp protocol on the wire: JSON-RPC with a lenient front door
The message layer is JSON-RPC 2.0, and every request has the same job to do before a framework
parses it: arrive as bytes that the parser will accept. Real clients do not always oblige, so the
gateway puts a normalising middleware in front of the application — but only in front of the part
that needs it:
# backend/smartgate/api/mcp_sse_compat.py — source lines 108–116 (NormalizeJsonRpcMiddleware)
async def __call__(self, scope: Scope, receive: Receive, send: Send) -> None:
if scope["type"] != "http" or scope.get("method") != "POST":
await self.app(scope, receive, send)
return
path = scope.get("path", "")
if "messages" not in path:
await self.app(scope, receive, send)
return
# backend/smartgate/api/mcp_sse_compat.py — source lines 131–138 (NormalizeJsonRpcMiddleware)
async def replay_receive() -> dict[str, Any]:
nonlocal sent
if sent:
return {"type": "http.disconnect"}
sent = True
return {"type": "http.request", "body": body, "more_body": False}
await self.app(scope, replay_receive, send)
Read the guards as carefully as the fix. The middleware touches POST requests whose path contains
the messages route segment, and passes everything else straight through, so tool responses and
unrelated endpoints pay nothing for it. The body is buffered until the transport reports the last
chunk, normalised once, and then handed to the application through a receive callable that
returns the rewritten body on the first call and a disconnect on every call after it — the ASGI
idiom for "these bytes were already read". Getting that closure wrong is how a server appears to
hang on the second read, with the timeout blamed on the network. The exact envelopes, and one
tools/call traced from request to result, are in the
MCP JSON-RPC round trip.
mcp vs a2a protocol: two envelopes, two jobs
Both protocols talk JSON, and that is where the resemblance ends. MCP describes how an agent uses
a tool: a method, a params object, a lifecycle around a session. The a2a protocol describes how
two agents exchange a task: who asked, what state the work is in, and how the result comes back.
They are complements — an agent that owns a task may call tools over MCP while coordinating with
another agent over a2a — and the shared engineering problem is the same, because both arrive as
bytes that a strict server will reject:
# backend/smartgate/api/mcp_sse_compat.py — source lines 67–76 (normalize_jsonrpc_body)
def normalize_jsonrpc_body(body: bytes) -> bytes:
"""Coerce non-object JSON-RPC params (e.g. []) to {} for pydantic validation."""
if not body:
return body
try:
data: Any = json.loads(body)
except (json.JSONDecodeError, UnicodeDecodeError):
return body
if not isinstance(data, dict):
return body
# backend/smartgate/api/mcp_sse_compat.py — source lines 80–87 (normalize_jsonrpc_body)
params = data.get("params")
if params is None:
data["params"] = {}
changed = True
elif isinstance(params, list):
# Cursor: tools/list, notifications/initialized with "params": []
data["params"] = {}
changed = True
The excerpt shows the two decisions a tolerant parser makes. A message body that is not JSON, or
not a JSON object at all, is returned untouched: leniency is not the same as repair, and a
genuinely malformed request still fails where it should. A valid message with no params, or with
params as an empty array — which one widely used client sends for parameterless methods — gets
an object instead, so validation passes and the call means what it always meant. The full
reference for request, notification, and result shapes lives on the
; this page stays on
the actors and the lifecycle around them.
What a model gateway must limit on the tool path
Tool calls are cheap to send and expensive to serve, so every serious deployment needs a limit
somewhere. The gateway applies two of them to the tool path at once — one per API key and one per
team — over a rolling window, and it decides what to do when the counter store itself is
unavailable:
# backend/smartgate/core/rate_limiter.py — source lines 51–91 (check_rate_limit_mcp)
async def check_rate_limit_mcp(
key_id: str,
team_id: str,
*,
per_key_limit: int,
team_ceiling: int,
window_seconds: int = 60,
) -> dict:
try:
redis = await get_redis()
except Exception as exc:
logger.warning("mcp rate_limit fail-open: %s", exc)
return {"allowed": True, "fail_open": True}
try:
bucket = int(time.time() / window_seconds)
key_bucket = f"rate_limit_mcp:{key_id}:{bucket}"
team_bucket = f"rate_limit_mcp_team:{team_id}:{bucket}"
key_count = await _incr_bucket(redis, key_bucket, window_seconds)
if key_count > per_key_limit:
ttl = await redis.ttl(key_bucket)
return {
"allowed": False,
"retry_after": max(1, ttl),
"limit_scope": "mcp_key",
}
team_count = await _incr_bucket(redis, team_bucket, window_seconds)
if team_count > team_ceiling:
ttl = await redis.ttl(team_bucket)
return {
"allowed": False,
"retry_after": max(1, ttl),
"limit_scope": "mcp_team",
}
return {"allowed": True}
except Exception as exc:
logger.warning("mcp rate_limit fail-open: %s", exc)
return {"allowed": True, "fail_open": True}
Three details repay attention. The window is a time bucket, so the count is shared by every
instance of the service without a lock. A rejected call comes back with a retry_after value
derived from the bucket's remaining time and a limit_scope naming which ceiling was hit, which
is what lets a client back off intelligently instead of retrying blindly. And the failure path is
explicitly fail-open: if the store cannot be reached, the call is allowed with a warning rather
than denied. That is an availability decision, not a security one — where a hard ceiling is
required, the quota and budget controls are the right instrument, as described in
enforce a token quota per team.
context window management is a client problem, not a protocol feature
Nothing in the protocol says how much of a tool result belongs in the next prompt. That is the
host's decision, and it is where context window management actually happens. What a server can do
is bound what it returns and make retrieval an explicit choice rather than a side effect of a
large response:
# backend/smartgate/core/resources.py — source lines 141–149 (search_vectors)
async def search_vectors(self, collection: str, query_vector, top_k: int = 10, filters=None):
from qdrant_client.models import Filter as QFilter
result = await self._client.search(
collection_name=collection,
query_vector=query_vector,
limit=top_k,
query_filter=filters,
)
return result
top_k and the filter are the budget, expressed as a query instead of as a truncation: ask for
ten passages from one collection rather than every match, and the prompt cost is predictable
before the model sees it. The gathering behind the query — chunk size, overlap, when to re-index
— is a separate problem, covered in
context window management techniques.
mcp inspector: watching a live protocol session
An inspector is a client, and that single fact explains most of what it can and cannot show you.
It connects over one of the transports, sends initialize, and renders what comes back. The
three things worth reading in that output are the negotiated protocol version, the capabilities
the server declared, and the tools/list payload — names, descriptions, and the annotations that
tell a host which calls may run without asking the user. An inspector that shows a method list
but no annotations is showing you the smaller half of the contract. When nothing appears at all,
the question is not "is the inspector broken" but "did the request reach the server": a call that
never arrived leaves no audit row, while a call that arrived and failed leaves one. We compare
the available tools in
MCP inspector alternatives.
The mcp specification: the normative core, layer by layer
The specification is easier to hold in your head as four layers, only two of which are
normative for every server. The message layer is JSON-RPC 2.0 with MCP-defined methods
(MCP specification). The transport
layer is stdio or Streamable HTTP, and a server implements at least one — on HTTP, session
management is optional and a server may answer each request independently
(transports). The
feature layer is what the server publishes: tools, resources, and prompts, with tools defined
by two methods, tools/list and tools/call
(tools). The client
layer is optional on both sides: sampling, roots, and elicitation exist, and a host may
implement none of them. Authorization is optional too, and where an HTTP transport supports it
the server behaves as an OAuth 2.1 resource server
(authorization).
Nothing in those four layers bounds how often a caller may call, what a call costs, or how long a
record is kept — which is precisely the space a gateway occupies. The
specification walkthrough takes the same layers with
more examples. The dated revisions, the schema files they ship, and the changelog between them are
collected in the MCP specification and schema docs.
mcp spec dates, and what a server must answer
Every revision of the specification is a dated document, and the date travels in the handshake.
When a client sends initialize, it states the protocol version it wants to speak; the server
replies with that version when it supports it, and otherwise with the newest version it does
support. A client written against an older revision therefore keeps working, and "which version
do you implement?" has a precise answer instead of a vague one. The practical work this implies
is small and worth doing once. Pin the revision you target in your own test suite and assert the
transport you serve there, because a transport change is the one difference a client cannot
negotiate around after the fact. Log the version pair when a handshake fails; it is the first
fact that separates a version mismatch from a transport problem. The dated revisions and their
transport rules are catalogued in
MCP protocol versions and transports.
Where an mcp server keeps protocol-adjacent state
The protocol is stateless and a deployed server is not. Rate-limit buckets, budget counters, and
team-scoped memory all have to survive a request and usually have to be shared by every instance
that serves the same key, which means a store outside the process. How the server finds that
store is a configuration question with a cheap and honest check at the start:
# lib/redis/config.ts — source lines 27–28 (isTcpRedisConfigured)
const url = redisUrl();
return Boolean(url && (url.startsWith("redis://") || url.startsWith("rediss://")));
The check accepts only a direct TCP Redis URL, in either the redis:// or rediss:// form, and
it returns a boolean instead of raising. Both properties are deliberate: the scheme test is
specific enough that a REST-style store URL will not be mistaken for a TCP one, and a boolean
means a misconfigured environment degrades the features that need shared state instead of failing
every request. The tools themselves are documented separately, in the
MCP tools reference.
Where the model context protocol stops and your stack starts
| Layer | Who owns it | What the protocol says |
|---|---|---|
| Host application | You | Nothing — the host sits outside the specification |
| Client and transport | An SDK, or your own code | JSON-RPC over stdio or Streamable HTTP |
| Server features | SmartGate, or your server | Tools, resources, and prompts, declared as capabilities |
| Limits, quotas, audit | Your gateway | Not specified; left to the operator on purpose |
The table is the useful summary of this page. The protocol standardises the middle two rows and
says almost nothing about the outer ones, which is why two deployments that both "support MCP"
can behave completely differently under load. If you are evaluating a server, the questions worth
asking are all in the bottom row: what bounds a key, what happens when a limit is hit, where the
record goes, and how long it stays. Everything above that row is a compatibility question with a
published answer. The vendor-side half of that question — which Claude clients and servers exist,
and how Anthropic announces them — is worked through in
Anthropic MCP clients and servers.
How to get started
- Read the envelopes first. The exact request, notification, and result shapes are on the message format page; this page explains the lifecycle they travel through.
- Connect once, with a free key. The connect page generates the block for your host, and the endpoint takes a plain bearer key.
-
Send
initialize, thentools/list. Two requests answer most of the compatibility questions: what version the server speaks, what it can do, and what the tools are. - Make one real tool call and follow it. The audit record shows the same call the inspector saw, from the server's side (tool documentation).
- Then decide what to implement. The features you skip — resources, prompts, sampling — are optional, but the says which revision made each one legal.
The mcp server page covers the other half of the picture: what a server
exposes, and how SmartGate's seven-tool server is put together.
FAQ
What problem does the Model Context Protocol solve?
It standardises how an AI host discovers and calls tools, so one integration works across hosts
instead of one integration per host-application pair.
Which actors does the protocol define?
Three: the host application, the client it runs, and the server that exposes tools. Only the
client and the server are specified.
Are sessions required?
No. Session management is optional in the Streamable HTTP transport, and a server may answer each
request independently.
How does a client know what a server supports?
From the initialize response, which carries the negotiated protocol version and the server's
capability groups, and from tools/list, which carries the tool metadata.
Is MCP the same as the a2a protocol?
No. MCP connects an agent to tools and context; the a2a protocol connects agents to each other. A
system can use both, and most will.
Does the protocol cover cost control?
No. Rate limits, quotas, budgets, and audit retention are operator concerns; the specification
leaves them out, which is why they belong in a gateway.
Limitations and what this does not do
- The protocol is not a policy engine. It describes how to call a tool, not how often, at what cost, or under whose authority; those questions need a layer above it.
- Transport choice is not reversible for free. A stdio server is a local process by construction; moving to HTTP changes the session model, and a client written for one may need work for the other.
- Capabilities are declarations, not guarantees. Both sides can declare a feature and implement it partially; the handshake records intent, not behaviour.
- Nothing here is a security boundary. Authorization is optional in the specification and careful in a deployment; treat a tool call as you would treat any other remote call.
- Dated revisions move. The layer names are stable, but supported versions and feature groupings change between revisions, so read the revision you target rather than a summary.
Sources
- Model Context Protocol — specification: https://modelcontextprotocol.io/specification/2026-07-28
- Model Context Protocol — transports, including stateless Streamable HTTP: https://modelcontextprotocol.io/specification/2026-07-28/basic/transports
- Model Context Protocol — server tools (
tools/list,tools/call): https://modelcontextprotocol.io/specification/2025-06-18/server/tools - Model Context Protocol — authorization: https://modelcontextprotocol.io/specification/2026-07-28/basic/authorization
- Anthropic — introducing the Model Context Protocol: https://www.anthropic.com/news/model-context-protocol
- SmartGate — product site and pricing: https://smartgate.network
Method note
The code in this article is not transcribed. Each block was cut directly out of the slice body
returned by the SmartGate slice API and re-asserted byte-for-byte as a substring of that body
before publication, and the first line inside every fence records the file and the exact source
lines. Symbols were pinned with whole-name containment (rule A level 2) and confirmed by the
service's slot-proof endpoint. Two sections quote trimmed windows of files that the sibling
message-format page quotes in full, so the two pages share no excerpt; the excerpt of
mount_mcp_routes and isTcpRedisConfigured likewise starts after the declaration line the
mcp-server page shows. Sections 8:1 (mcp inspector), 9:1 (mcp specification), and 10:1
(mcp spec) pinned no slice: they are written from the public specification, quote no code, and
their abstention is recorded rather than hidden. No third-party client or server implementation is
quoted anywhere on this page.
Slice provenance
| # | SERP keyword | Symbol | File | Source lines | How it was pinned | sha256(12) |
|---|---|---|---|---|---|---|
| 1 | a2a protocol | mount_mcp_routes |
backend/smartgate/api/mcp.py |
401–406 | rule A L2 → slot-proof | e66f6b69a172 |
| 2 | mcp capabilities | infer_transport |
backend/smartgate/core/audit_enrichment.py |
30–32 | rule A L2 → slot-proof | fc93297c9232 |
| 3 | ai gateway | infer_route |
backend/smartgate/core/audit_enrichment.py |
35–40 | rule A L2 → slot-proof | 8cb200979cc6 |
| 4 | mcp initialize | initialize |
backend/smartgate/core/resources.py |
64–69 | rule A L2 → slot-proof | ef0d1ff04523 |
| 5 | mcp protocol | NormalizeJsonRpcMiddleware |
backend/smartgate/api/mcp_sse_compat.py |
108–116, 131–138 | rule A L2 → slot-proof | 472493b3162e |
| 6 | mcp vs a2a protocol | normalize_jsonrpc_body |
backend/smartgate/api/mcp_sse_compat.py |
67–76, 80–87 | rule A L2 → slot-proof | d8359617452f |
| 7 | model gateway | check_rate_limit_mcp |
backend/smartgate/core/rate_limiter.py |
51–91 | rule A L2 → slot-proof | b540f8154beb |
| 8 | context window management | search_vectors |
backend/smartgate/core/resources.py |
141–149 | rule A L2 → slot-proof | 0e5cc4aeeb5b |
| 9 | mcp server | isTcpRedisConfigured |
lib/redis/config.ts |
27–28 | rule A L2 → slot-proof | 7ca086a475d5 |
Every fenced block above was cut from the slice body and re-asserted against it byte-for-byte before
publication. 9 of 12 sections pinned, 0 abstentions, 3 misses.
Top comments (0)