20 of 20
claude -pruns on Claude Code 2.1.285 opened our stdio MCP server withserver/discover, a probe the docs say stdio servers only get whenMCP_PROTOCOL_NEGOTIATION=autois set, and in the 2 runs where the server answered it,initializewas never sent, so the instructions we had put in theinitializereply never reached Claude. The rest matched the docs: the instructions arrived as a<system-reminder>in the first user turn (never in the system prompt), cut at 2,048 characters, and the capped block added 721 tokens to the first request in English and 1,727 in Japanese.
An MCP server can hand the client a paragraph of plain text when it connects: the instructions field. It is the one place where a server author gets to tell the model, in prose, what the server is for and when to reach for it. Claude Code's MCP docs say the field "becomes more useful with tool search enabled", since only tool names and server instructions load at session start, and our tool search measurement ended on the same advice. This time we measured the field itself: where Claude Code puts the text, how much of it survives, when it arrives, and which of the server's replies it is taken from.
The lab is one small stdio server whose instructions carry codewords at both ends and position markers in between, twenty claude -p runs against it, and three kinds of evidence per run: the model's answer, the session transcript under ~/.claude/projects/, and a log the server writes of every JSON-RPC message it receives. Everything below was run on 2026-09-30 with Claude Code 2.1.285 (claude --version), --model opus (which resolved to claude-opus-5-5), and Node v22.22.2 on macOS.
What the spec and the docs say
In the 2025-11-25 revision of the MCP specification, instructions is an optional field of the reply to initialize, and the lifecycle page is strict about ordering: "The initialization phase MUST be the first interaction between client and server." Its example reply ends with "instructions": "Optional instructions for the client", and the schema describes the field like this:
Instructions describing how to use the server and its features. This can be used by clients to improve the LLM's understanding of available tools, resources, etc. It can be thought of like a "hint" to the model. For example, this information MAY be added to the system prompt.
The latest revision, 2026-07-28, has no initialize handshake: version, identity and capabilities travel as metadata on each request. The page at the same lifecycle path is titled "Versioning and Compatibility" there, and it names the old way Legacy: "protocol versions that establish a session with an initialize handshake (2025-11-25 and earlier)." In that revision, instructions is a field of the result of a new method, server/discover, described as "Natural-language guidance describing the server and its features. This can be used by clients to improve an LLM's understanding of available tools (e.g., by including it in a system prompt)." For stdio, a client that speaks both eras "SHOULD probe with server/discover before sending any other request", and if the server "returns any other error, or does not respond within a reasonable timeout: the server is legacy. Fall back to the initialize handshake."
Claude Code's MCP page, fetched the same day, makes three statements about the field. With tool search, "Only tool names and server instructions load at session start". Then: "Claude Code truncates each tool description and each server's instructions at 2,048 characters by default. Keep them concise, and put critical details near the start." And the limit can be changed with CLAUDE_CODE_MAX_MCP_DESCRIPTION_LENGTH, an environment variable whose reference entry says it "Accepts a positive whole number in plain digits. Anything else is ignored and the default applies."
The same page has a section on the two MCP client runtimes. The v2 runtime adds the 2026-07-28 revision, and on v2 Claude Code "Asks HTTP servers whether they support the newer revision, and uses it with those that do. It also asks claude.ai connector servers in sessions where it fetches feature flags. To have it ask stdio servers, or connector servers in every session, set MCP_PROTOCOL_NEGOTIATION to auto. It connects to every other server as v1 does." The environment variable's own entry repeats the default: "Without the variable, Claude Code probes HTTP servers, and also probes claude.ai connector servers in sessions where it fetches feature flags."
So by the docs, a stdio server should see initialize first unless the user has opted in. The rest of this article hangs on that sentence.
The lab
The server is a Node script with no dependencies that speaks newline-delimited JSON-RPC over stdio. It has three tools (lookup_part, list_bins, lab_ping), so that tool search has something to defer, and it appends every message it receives to rpc-log.jsonl. An environment variable in the MCP config chooses the instructions: none (the field is left out of the reply), English text of a given length, or Japanese text of a given length. Two more switches make it answer server/discover the way a 2026-07-28 server would, or hold back its initialize reply for eight seconds.
The text is built so that Claude can only report what it actually received. It starts with BEGIN codeword: <word>-<four digits>. and ends with END codeword: <word>-<four digits>., and in between it repeats a few sentences about a parts inventory with a marker about every 250 characters, written as [at N], where N is the marker's own character offset. Each text has its own pair of codewords, derived from its length and language, and in the dual-era setup the two replies carry different pairs, so an answer shows which text Claude was given.
Each server setup had its own config file, passed with --strict-mcp-config so that no other MCP server, and none of the account's claude.ai connectors, joined the session. The configurations that change a Claude Code environment variable reuse one of these files and set the variable on the claude process. The file for the 20,000-character text:
{
"mcpServers": {
"lab": {
"type": "stdio",
"command": "node",
"args": ["/path/to/server.js"],
"env": { "LAB_INSTR": "ascii:20000" }
}
}
}
Eighteen runs used this command and this prompt, from an empty directory (the two late-server runs change the prompt and two flags, as their own section below describes):
claude -p "$PROMPT" --output-format stream-json --verbose --max-turns 1 --model opus \
--settings '{"disableAllHooks": true}' --strict-mcp-config --mcp-config cfg/a20k.json \
--debug-file runs/R05.debug.log
Do not call any tools; answer only from what is already in your context. An MCP server named lab may have given you instructions. Reply with exactly one line in this format: BEGIN=<the codeword after "BEGIN codeword:"> END=<the codeword after "END codeword:"> LAST=<the largest number N in any marker written as [at N]>. Write NONE for any value you cannot find.
We started the runs from inside another Claude Code session, so a wrapper removed the environment variables that session exports (CLAUDECODE and its neighbors) before calling claude. The --settings override turned hooks off, so the user-level notification hook stayed quiet.
For the numbers, we read the transcript. The first assistant record's usage gives the size of the first request: input_tokens plus cache_read_input_tokens plus cache_creation_input_tokens. On 2.1.285 the transcript also stores each block of context that Claude Code adds to the first user turn as an attachment record with its rendered text, next to a prompt_snapshot record that holds the system prompt and the tool definitions. Those records let us say where the instructions went, and how many characters of them, without relying on what the model said.
We ran ten configurations twice each, twenty runs in all. In every pair, the first request matched to the token.
The results in one table
| Configuration | First request (tokens) | vs. no instructions | What reached Claude | Claude's answer |
|---|---|---|---|---|
No instructions field |
20,472 | nothing | NONE for all three | |
| 1,000 characters, English | 20,849 | +377 | all 1,000 | both codewords, LAST=763 |
| 5,000 characters, English | 21,193 | +721 | first 2,048 + … [truncated]
|
BEGIN only, LAST=2009 |
| 20,000 characters, English | 21,193 | +721 | first 2,048 + … [truncated]
|
BEGIN only, LAST=2009 |
20,000, limit variable set to 25000
|
27,050 | +6,578 | all 20,000 | both codewords, LAST=19808 |
20,000, limit variable set to 25,000
|
21,193 | +721 | first 2,048 + … [truncated]
|
BEGIN only, LAST=2009 |
| 3,000 characters, Japanese (7,956 bytes) | 22,199 | +1,727 | first 2,048 + … [truncated]
|
BEGIN only, LAST=2026 |
1,000 characters, ENABLE_TOOL_SEARCH=false
|
36,871 | (all tools upfront) | all 1,000 | both codewords, LAST=763 |
1,000 characters in each reply, server also answers server/discover
|
20,849 | +377 | the server/discover text only |
the discover codewords |
1,000 characters, initialize answered after 8 s, CLAUDE_CODE_MCP_STARTUP_WAIT_MS=0
|
20,590, then 21,288 / 21,208 | nothing on request 1, all 1,000 on request 2 | both codewords, LAST=763 |
The limit variable is CLAUDE_CODE_MAX_MCP_DESCRIPTION_LENGTH. In all twenty runs, Claude's answer matched what the transcript says it was given: the right codewords where the text was delivered, NONE where it was not, and the last marker before the cut where it was truncated. It did not invent a codeword once. The twenty runs together were reported at $1.02 in total_cost_usd.
The rest of this article takes the table apart one column at a time.
The handshake: server/discover came first in 20 of 20
The server's log told the first story. In every run, the first message Claude Code wrote to the server's stdin was not initialize but server/discover, always with the id server-discover-probe-1. From the second run on, the server also logged the _meta of each request, and the probe looked like this in all 19 of those runs (shortened; the _meta also carries the client's capabilities):
{"method": "server/discover", "id": "server-discover-probe-1",
"params": {"_meta": {"io.modelcontextprotocol/protocolVersion": "2026-07-28",
"io.modelcontextprotocol/clientInfo": {"name": "claude-code", "version": "2.1.285"}}}}
In 18 runs, the server did not know the method and answered with the JSON-RPC error -32601. Claude Code then sent initialize with protocolVersion: "2025-11-25", followed by notifications/initialized and tools/list, and the debug log recorded "protocolEra":"legacy","negotiatedProtocolVersion":"2025-11-25". That is the fallback the 2026-07-28 spec prescribes. In the server's log, initialize arrived 5 milliseconds to 1.1 seconds after the probe, and in those 18 runs the instructions from the initialize reply were the ones Claude received.
In the two dual-era runs, the server answered the probe with a DiscoverResult whose instructions carried one pair of codewords, and it was ready to answer initialize with a different pair. Claude Code never asked. The next message was tools/list, with protocol version 2026-07-28 in its _meta, and the debug log recorded "protocolEra":"modern","negotiatedProtocolVersion":"2026-07-28". Claude answered with the server/discover codewords both times, and the initialize codewords appear nowhere in either transcript.
Both behaviors follow the 2026-07-28 spec. What does not line up is Claude Code's own account of when it probes. MCP_PROTOCOL_NEGOTIATION was not set: it was not in the shell environment or in the env block of ~/.claude/settings.json, and the machine has no managed settings file. By the docs, a stdio server would then be connected "as v1 does", with initialize first. All twenty debug logs contain the line mcp runtime arm: v2 (source: growthbook), which we read as a fetched feature flag putting the session on the v2 runtime. The docs describe that part too: in sessions where it fetches feature flags, Claude Code "uses the v2 runtime on Claude Code v2.1.232 or later." Nothing in the logs says why the stdio server was probed. It may be a flag, it may be documentation that has fallen behind the code, and from the outside we cannot tell which. The observable fact is that the documented default for stdio servers did not happen in 20 of 20 runs.
For a legacy server that answers an unknown method with an error, as ours did, the probe changed nothing that Claude saw. The case to watch is a server that implements both eras. Once server/discover succeeds, the initialize reply is never requested, and any instructions that live only there are lost. That can happen quietly, for example when a server builds its legacy reply and its new reply from different fields.
Where the text lands: a reminder in the first user turn
In every run that had instructions, the text appeared in exactly one place, an attachment of type mcp_instructions_delta. Its rendered text is a <system-reminder> block that begins like this:
<system-reminder>
# MCP Server Instructions
The following MCP servers have provided instructions for how to use their tools and resources:
## lab
BEGIN codeword: WAXWING-4465. The lab server tracks spare parts for a small hardware workshop. ...
For our 1,000-character instructions, the rendered block was 1,167 characters, so the heading and the wrapper add 167 characters. The attachment stores the server name under addedNames and the text under addedBlocks, next to an empty removedNames, and the late-server runs below show it behaving as a delta: it is added when a server's instructions first become available. In the transcript it sits in the chain of attachments that follows the prompt, after the deferred tool list and the agent listing and before the skill listing. When the server sent no instructions field, there was no such attachment at all, not even an empty heading.
We also checked the other places the BEGIN codeword could have gone. The system prompt in the prompt_snapshot record was 6,662 characters in all twenty runs and never contained the codeword or the "MCP Server Instructions" heading. The deferred tool list carried the lab server's three tool names and nothing from its instructions, and the tool definitions did not contain it either. So the placement the MCP spec gives as its example, the system prompt, is not where Claude Code puts this text. The spec only says "MAY", so this is not a violation, but it matters if you think of instructions as system-level text. In Claude Code they are a reminder attached to a user turn.
Turning tool search off did not move it. With ENABLE_TOOL_SEARCH=false, the 17 tools that are otherwise deferred (14 built-in, 3 from the lab server) went out as full definitions, the search tool itself left, and the first request grew to 36,871 tokens. The instructions were in the same mcp_instructions_delta attachment, the same 1,167 characters. The docs' sentence about tool search holds, since the instructions did load at session start with tool search on, but they load at session start without it as well.
How much: 2,048 characters, then a 13-character marker
With 5,000 and 20,000 characters of English, the text under ## lab was 2,061 characters long. The first 2,048 matched the server's text exactly, and Claude Code appended … [truncated]: an ellipsis, a space and the word in brackets, 13 characters in all. In the eight runs where the text was cut, Claude reported END=NONE and the last marker before the cut, 2009 for the English text and 2026 for the Japanese one.
The cut is silent outside the debug log. Stderr was empty in every run, and the stream-json output never mentioned it. The debug file had one line per run, for example MCP server "lab": Server instructions truncated from 20000 to 2048 chars.
The unit is characters, as the docs say, not bytes. The Japanese instructions were 3,000 characters and 7,956 bytes in UTF-8, and the block kept the first 2,048 characters, which is 5,452 bytes. In the tool search article we paraphrased the limit as 2KB. Measured, it is 2,048 characters, and in Japanese that came to more than 5KB.
Because the cap counts characters, what it costs in tokens depends on the language. The capped English block added 721 tokens to the first request, and the capped Japanese block added 1,727, 2.4 times as many for the same 2,048 characters. The 1,000- and 20,000-character English runs put the heading and wrapper at about 51 tokens and our English filler at about 3.06 characters per token, which predicts the capped English block to within 3 tokens. The Japanese block came to about 1.23 characters per token. The cap limits how much a server can say, not what it costs.
Raising the limit worked as documented. With CLAUDE_CODE_MAX_MCP_DESCRIPTION_LENGTH=25000, all 20,000 characters arrived, the debug log had no truncation line, Claude reported the END codeword and the last marker, 19808, and the first request grew by 6,578 tokens. Writing the same number as 25,000 did nothing. The comma made it an invalid value, the default applied, and the text was cut at 2,048 characters, with no warning about the value in stderr or in the debug log. That is also what the docs say. Keep in mind that this is a Claude Code environment variable, set on the user's side, so a server author cannot assume it.
When: only if the server is connected when the request is built
In -p mode with --mcp-config, Claude Code waits for servers that are still connecting before the first turn, "up to the MCP_TIMEOUT startup timeout, 30 seconds by default", according to the CLI reference, which adds that a server with a cached tool list skips the wait. Our server connected in 0.2 to 3.1 seconds, always before the first request went out, so in eighteen runs the question never came up. Interactive sessions are different; the environment variable reference says "MCP startup is non-blocking by default: servers connect in the background and their tools become available as they finish."
To see what a late server looks like from the inside, we made the server hold its initialize reply for 8 seconds and set CLAUDE_CODE_MCP_STARTUP_WAIT_MS=0, which the docs give as the way to skip the first-turn wait. The prompt asked Claude to run sleep 10 with Bash first and answer afterwards, with --max-turns 3 and --allowedTools "Bash(sleep:*)" in place of --max-turns 1, so that the session would make a second request.
In both runs, the init event listed the server as pending, and the first request went out within a second of Claude Code starting to connect to it. That request, 20,590 tokens, had no instructions attachment. Its deferred tool list carried a note instead: "The following MCP servers are still connecting — their tools (typically named mcp__*) are not yet available but will appear shortly:", followed by lab. The server finished connecting 8.3 and 8.6 seconds after Claude Code had started connecting to it. When the Bash result came back, Claude Code attached two deltas to it, the three lab tool names and the mcp_instructions_delta block, and the second request carried both. It was 21,288 tokens in the run whose first reply had opened with a thinking block, and 21,208 in the other. Claude then answered with both codewords and LAST=763.
So "at session start" means "in the first request built after the server connected". A server that is slow to answer initialize misses the first request of a session that does not wait for it, and its instructions reach Claude with whatever request comes next. By the same logic, a single-request -p run with the wait turned off would never show Claude the instructions of a server that connects late. We did not run that case.
What we would do as server authors
These are our readings of the twenty runs, not statements from the docs.
Put the instructions in both replies if the server speaks both eras. Claude Code 2.1.285 sent server/discover to our stdio server in every run, and once that probe succeeded, initialize was never sent. Build both replies from one string, and look at the server's own log to see which method actually arrived.
Assume only the first 2,048 characters exist. Everything after them was invisible to the model in every run, and nothing outside the debug log said so. If the text is longer, the part that tells Claude when to search for your tools belongs at the top, as the docs advise.
Budget in tokens, not characters, if the instructions are not in English. The same cap cost 2.4 times as many tokens in Japanese.
Do not count on system-prompt treatment. The text arrives as a reminder next to the user's first message. If a sentence only makes sense as a standing system rule, test that Claude follows it in that position.
Start fast. A server that connects after the first request is built is not in that request.
What we did not measure
Everything ran on one machine, one account, one model (Opus 5.5) and one version (2.1.285), all headless. We did not open an interactive session, so the late-server behavior there is inferred from the docs and from the -p runs with the wait turned off.
We did not run MCP_PROTOCOL_NEGOTIATION=legacy or MCP_SDK_GENERATION=v1. By the docs, either should keep a stdio server on initialize. We stopped at twenty runs, and we cannot say whether the probe we saw comes from a feature flag specific to this account.
The dual-era server was the 136-line lab script. We did not test any MCP SDK, so we do not know which SDKs answer server/discover or what they put in its instructions. We also did not test a legacy server that crashes or hangs on an unknown first method instead of returning an error.
Only one server was connected per run. We did not test how several servers' instructions are ordered, or whether each server gets its own 2,048 characters. We did not measure tool descriptions, which the same limit also covers.
The +377, +721, +1,727 and +6,578 figures are differences from the no-instructions runs and include the heading and wrapper. The ENABLE_TOOL_SEARCH=false runs had no matching baseline, so we report their total only. The attachment and prompt_snapshot records are not documented output and could change in any release.
We only checked whether Claude could read the codewords back. Whether it follows instructions delivered in this position is a different question, and we did not test it.
Reproduce it
This is a shorter version of the server: one tool, fixed codewords, and no Japanese or slow modes. We checked its replies by piping JSON-RPC lines into it; the numbers above come from the longer script.
// lab.js: a stdio MCP server that logs every message and returns long instructions.
// LAB_CHARS sets the length. LAB_DUAL=1 also answers server/discover, with other codewords.
const fs = require('fs');
const log = (m) => fs.appendFileSync(__dirname + '/rpc-log.jsonl', JSON.stringify(m) + '\n');
function instructions(begin, end, n) {
const tail = ` END codeword: ${end}.`;
let s = `BEGIN codeword: ${begin}. `;
let next = 250;
while (s.length < n - tail.length) {
if (s.length >= next) { s += `[at ${s.length}] `; next += 250; }
s += 'Use lookup_part when the user names a part number or asks about stock. ';
}
return s.slice(0, n - tail.length) + tail;
}
const n = Number(process.env.LAB_CHARS || 1000);
const fromInit = instructions('HERON-4821', 'OSPREY-3094', n);
const fromDiscover = instructions('TANAGER-6881', 'WAXWING-2170', n);
const dual = process.env.LAB_DUAL === '1';
const tools = [{ name: 'lookup_part', description: 'Look up one part number.',
inputSchema: { type: 'object', properties: { part: { type: 'string' } } } }];
function handle({ id, method, params }) {
log({ t: new Date().toISOString(), method, id });
if (id === undefined) return null;
const ok = (result) => ({ jsonrpc: '2.0', id, result: dual ? { resultType: 'complete', ...result } : result });
if (method === 'server/discover' && dual) return ok({
supportedVersions: ['2026-07-28'], capabilities: { tools: {} },
_meta: { 'io.modelcontextprotocol/serverInfo': { name: 'lab', version: '0.0.1' } },
instructions: fromDiscover, ttlMs: 0, cacheScope: 'private' });
if (method === 'initialize') return ok({
protocolVersion: params.protocolVersion, capabilities: { tools: {} },
serverInfo: { name: 'lab', version: '0.0.1' }, instructions: fromInit });
if (method === 'tools/list') return ok(dual ? { tools, ttlMs: 0, cacheScope: 'private' } : { tools });
return { jsonrpc: '2.0', id, error: { code: -32601, message: 'Method not found' } };
}
let buf = '';
process.stdin.setEncoding('utf8').on('data', (chunk) => {
buf += chunk;
let i;
while ((i = buf.indexOf('\n')) >= 0) {
const line = buf.slice(0, i).trim();
buf = buf.slice(i + 1);
if (!line) continue;
const reply = handle(JSON.parse(line));
if (reply) process.stdout.write(JSON.stringify(reply) + '\n');
}
});
LAB=$(mktemp -d) && cd "$LAB" # save lab.js here
cfg() { printf '{"mcpServers":{"lab":{"type":"stdio","command":"node","args":["%s/lab.js"],"env":{"LAB_CHARS":"%s","LAB_DUAL":"%s"}}}}' "$LAB" "$1" "$2"; }
cfg 20000 0 > long.json; cfg 1000 1 > dual.json
P='Do not call any tools. Reply with the codeword after "BEGIN codeword:" and the one after "END codeword:" in your context, or NONE for any you cannot see.'
mkdir empty && cd empty
for c in long dual; do
claude -p "$P" --max-turns 1 --output-format json --strict-mcp-config \
--mcp-config "../$c.json" --debug-file "../$c.log" | jq -r '.result, .session_id'
done
jq -r .method ../rpc-log.jsonl # which handshake your version sends
grep -h "instructions truncated" ../long.log
To see what reached Claude, open the transcript for a session id under ~/.claude/projects/ and print the block with jq -r 'select(.attachment.type == "mcp_instructions_delta") | .attachment.addedBlocks[]'. If your method log starts with initialize, your session is connecting to stdio servers the way the docs describe, and the dual run will show the initialize codewords instead.
The lab here is one Node file and a handful of claude -p flags, which makes the handshake log cheap to check again after each Claude Code upgrade.
If your version sends initialize first, paste your server's method log in the comments below.


Top comments (0)