DEV Community

Cover image for Your Memory Server Did Not Say destructiveHint. By the Spec, That Means True.
Edward Izgorodin
Edward Izgorodin

Posted on

Your Memory Server Did Not Say destructiveHint. By the Spec, That Means True.

A tool that ships no annotations has not stayed silent about whether it is destructive. Under the MCP schema it has answered true.

The scanner asked for something the spec calls optional

Eugeniya Ivanova published a walkthrough on 2026-09-07 of getting an MCP server through the ChatGPT app directory review, and one section of it is about annotations. The scanner went through every tool and required explicit readOnlyHint, openWorldHint and destructiveHint values. Four of her read-only tools carried no destructiveHint, which the specification permits, and the scanner wanted it anyway. She added destructiveHint: false to those tools and wrote justifications for forty-odd annotation values. By her account, the tools themselves did not change.

Himanshu Kumar had measured the other end of the same field a week earlier, on 2026-08-30, auditing a deployed server: six tools, zero tools declaring any annotation. His sentence for it is the one worth keeping: "The server did not lie. It said nothing."

The arithmetic of the defaults is theirs, not mine. What follows is a count of how often those four names appear in the published source of seven source trees, and what the counting turned out to be unable to tell me.

Four defaults, and a note telling you not to act on them

ToolAnnotations is declared at line 1912 of schema/2026-07-28/schema.ts in the specification repository (commit 271ecc9, fetched 2026-09-07, 98,426 bytes). Four boolean fields, each with a documented default, and each default with a consequence when the field is missing.

Field Default in the schema What a client honoring defaults must assume when the field is absent
readOnlyHint false (line 1921) the tool modifies its environment
destructiveHint true (line 1931) the tool may perform destructive updates
idempotentHint false (line 1941) repeating the call has further effect
openWorldHint true (line 1951) the tool reaches an open world of external entities

Directly above that interface, at lines 1903 to 1908, sits a NOTE that has to travel with any argument built on those defaults: "all properties in ToolAnnotations are hints. They are not guaranteed to provide a faithful description of tool behavior", and "Clients should never make tool use decisions based on ToolAnnotations received from untrusted servers."

So the schema supplies defaults for a field it also instructs clients not to decide on. Both halves are in force, and the gap between them is the subject.

One silence, two lawful readings

Take a memory server whose save_memory tool ships no annotations. A client implementing the documented defaults must treat that tool as not read-only, destructive, non-idempotent and open-world. A client reading the absent object as "the server did not say, so no annotation gate applies" routes the same call straight through. Neither client is misbehaving. The schema backs the first, the NOTE backs the second, and the server said the identical nothing to both.

A second rule at line 1929 makes the reading depend on a different field: destructiveHint "is meaningful only when readOnlyHint == false". On a tool declaring readOnlyHint: true, omitting destructiveHint is exactly right, and reading true into it is an error. On a tool declaring nothing at all, readOnlyHint falls to its default of false, which makes destructiveHint meaningful, which makes its default of true apply. The same omission is clean in one place and loud in the other, and what separates them is a second field the tool also did not fill in. An empty annotation is not missing information. It is a value, and different clients will supply different ones.

Counting the four names in seven source trees

Method before numbers. For each project I downloaded the branch archive at a named commit, walked every regular file in it, and counted case-insensitive byte occurrences of the four field names. Not a search restricted to paths containing mcp: that filter would have missed the assertions in mem0's tests/test_memory_core.py. Case-insensitive because the Go SDK capitalizes the same fields, and a case-sensitive grep returns a false zero on any Go server.

Source tree Commit Files walked readOnlyHint destructiveHint idempotentHint openWorldHint
modelcontextprotocol/servers d73f99e 156 61 51 51 60
supermemoryai/supermemory 4d8a4eb 1,192 6 6 6 6
mem0ai/mem0 dae67f7 1,777 2 0 2 1
getzep/zep 54f63ee 936 0 0 0 0
getzep/graphiti b943c9e 360 0 0 0 0
topoteretes/cognee e93a4f0 3,594 0 0 0 0
MemoriLabs/Memori 10d6501 673 0 0 0 0

All rows probed 2026-09-07. GibsonAI/memori now redirects to MemoriLabs/Memori. The package I publish is not in the count, because measuring my own tree beside other people's, on a method I picked myself, is not a comparison a reader should have to trust.

Four rows read zero, and a zero from a weak probe proves nothing, so the walker needs a control. In the same pass it counted the word annotations in those four trees: 115 in zep, 30 in graphiti, 210 in cognee, 36 in Memori. The files were read. Three of the four also ship an MCP server in the same tree: zep at mcp/zep-mcp-server with 13 tools registered through mcp.AddTool, graphiti at mcp_server/src/graphiti_mcp_server.py with 13 @mcp.tool() decorators, and cognee at cognee-mcp/src. Memori is the exception: the integrations/openclaw/src/tools path with its 6 registerTool calls is an OpenClaw plugin, and the MCP server sits in a separate public repository, MemoriLabs/memori-mcp (commit e02957d, 12 files), documented as a hosted endpoint. That repository reads zero on all four fields as well, walked the same day. The rows describe servers that exist, and for Memori the server is one repository over.

The most useful result is a limit on the method itself. Supermemory reads 6 in every column, which sounds like six annotated tools. It is not. apps/mcp/src/server/tools/annotations.ts defines four named presets, all four fields set in each, and fifteen tool files import one of them. Counting literals undercounts the annotated tools by a factor of two and a half in the one project that factored them into a constant. This census answers whether a tree declares these fields at all. It does not count tools.

A zero in the destructiveHint column is not automatically a gap either. The mem0 plugin exposes one tool, a memory search, declaring readOnlyHint: True, idempotentHint: True and no destructiveHint. By line 1929 that omission is precisely the case the spec calls meaningless, so leaving it out is correct.

Graphiti shows most plainly what the defaults do when nothing is declared. Thirteen tools, every decorator a bare @mcp.tool(), and the names run from clear_graph, delete_entity_edge and delete_episode through to search_nodes, get_episodes and get_status. A client honoring the defaults treats all thirteen alike. On the first three that lands right. On the last three it does not.

Read your own tools/list

Source is not the wire. Three JSON-RPC lines on stdin get the answer from a running server, and this prints, per tool, which of the four it omits, skipping the two the schema calls meaningless once readOnlyHint resolves to true.

printf '%s\n' \
  '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2025-06-18","capabilities":{},"clientInfo":{"name":"probe","version":"0"}}}' \
  '{"jsonrpc":"2.0","method":"notifications/initialized"}' \
  '{"jsonrpc":"2.0","id":2,"method":"tools/list"}' \
| npx -y @modelcontextprotocol/server-memory 2>/dev/null \
| python3 -c 'import sys, json
FIELDS = ("readOnlyHint", "destructiveHint", "idempotentHint", "openWorldHint")
for line in sys.stdin:
    for t in (json.loads(line).get("result") or {}).get("tools") or []:
        a = t.get("annotations") or {}
        skip = {"destructiveHint", "idempotentHint"} if a.get("readOnlyHint") is True else set()
        missing = [f for f in FIELDS if f not in a and f not in skip]
        print(t["name"], "omits:", ", ".join(missing) or "nothing")'
Enter fullscreen mode Exit fullscreen mode

Corrected on 12 September 2026. The first version of this filter tested key presence and nothing else, so a read-only tool that left destructiveHint out was listed as omitting it, two sections after this article called that omission correct. A reader caught it. The filter now skips destructiveHint and idempotentHint once readOnlyHint resolves to true, which is what the schema says of both fields at lines 898 and 908. The same reader points out that honoring a documented default is itself a decision taken off the annotations channel, the channel the NOTE tells clients not to trust from an untrusted server, so the two readings above are both lawful and not equally cheap: the permissive one is what the type surface hands you for free.

Swap the npx line for your own server command. Against the reference server it prints omits: nothing nine times, matching its source: @modelcontextprotocol/server-memory fills all four fields on all nine tools (src/memory/index.ts, commit d73f99e). That package took 405,883 npm installs between 2026-08-08 and 2026-09-06, with the window set in the api.npmjs.org request. The example most servers learn their shape from fills everything.

openWorldHint, where the spec names memory as its example

Of the four fields, openWorldHint is the one where the specification picks a side by illustration. Lines 1948 and 1949: "For example, the world of a web search tool is open, whereas that of a memory tool is not." The default is true, and the worked example of false is a memory tool.

The reference server agrees with its own spec: false on all nine tools, backed by a local memory.jsonl file. Supermemory agrees too, false in all four presets. The mem0 plugin tool declares true.

The split is worth naming rather than resolving. The example in the spec fits memory that lives where the tool runs, a file or a local process, which is exactly what the reference server is, and the sentence does not say what to read when the tool hands the work somewhere further out. Servers taking the same line differently is what an optional hint permits, and the field cannot tell a client which reading a given server took. That is a limit of the hint, not a fault in anybody's server, and it is the limit the NOTE was warning about.

What this does not prove

This is a census of published source on one date, not a census of running servers. A live deployment can answer differently from the tree it was built from, and no hosted deployment was checked here, so nothing above transfers from a repository to a service.

A framework can attach annotations at registration or serialization time without the field name ever appearing in a project's own source. Every zero above means the four literals are absent from that tree at that commit. It does not mean the server answers tools/list without annotations.

The counts move under you. Two of the seven trees had head commits dated 2026-09-06, one day before this probe, so a run tomorrow counts a different tree. That is why every row names a commit rather than a branch, and why the table was rebuilt from scratch on the day of writing rather than carried over from an earlier pass.

None of this ranks memory engines, and none of it is a benchmark. Filling four hint fields is cheap and says nothing about retrieval quality. A project that leaves them empty may be right to, since the specification calls them optional and tells clients not to trust them when present.

Both readings of an absent field pass review. If your client documents which one it takes, I would like to read it.

Disclosure: I work on Mnemoverse, a memory engine for AI agents connected over MCP, so weigh the argument accordingly.

Top comments (13)

Collapse
 
anp2network profile image
ANP2 Network •

The published filter contradicts the meaningfulness rule you set out two sections earlier. It tests key presence and nothing else, so a tool that declares readOnlyHint true and leaves destructiveHint out gets destructiveHint printed under "omits". That is the omission you had just defended as correct. The mem0 search tool is exactly that shape. Reading the code as posted, the repair is one condition: drop destructiveHint from the omission list whenever readOnlyHint resolves to true. Without it the probe reports a gap where the schema says there is no question to answer.

The harder problem is that the two readings are not the balanced pair that "neither client misbehaves" suggests. Honoring the documented defaults is a tool use decision derived from ToolAnnotations. The input happens to be an absence rather than a value, but the decision is still coming off that channel, which is the channel the NOTE tells clients not to decide on when the server is untrusted. So the schema and the NOTE are not backing one reading each. They are both in play on the same side.

And the pressure runs backwards. Silence is the strict setting, while a declaration costs a line of JSON and nobody checks it. A server that wants the gate open will not stay quiet. It writes readOnlyHint true. The default honoring client therefore puts friction on servers that could not be bothered to type, and lets through any server that was. That second group is the one the gate was built for. Your two sources close the loop on this: the review you cite passed a server whose tools were unchanged and whose annotation values were not, so a census of these four names is measuring a willingness to declare.

On your closing question, the way out may be to stop asking the annotation to hold weight it cannot hold. A client already has one fact the server did not author, which is the capability it handed over at connect time, the credential scope or the filesystem root. readOnlyHint is a claim about the tool. The grant is the client's own record, and no annotation widens it. A client could document that annotations drive display and prompting only, and that anything it refuses is refused at the grant. Nothing the server writes can move that.

Does your client turn the default derived destructive reading into a prompt or into a refusal? The NOTE bites very differently on those two.

Collapse
 
izgorodin profile image
Edward Izgorodin •

The filter contradicts the rule the section before it states, and the contradiction is mine, not a reading of yours. The probe tested key presence and nothing else, so a read-only tool that left destructiveHint out was listed as omitting it two sections after the text called that omission correct. The repair is the condition you describe, applied a little wider than you propose: the schema attaches the same clause to idempotentHint, lines 898 and 908 of the 2025-06-18 schema.ts both say the property is meaningful only when readOnlyHint is false, so the probe now skips both once readOnlyHint resolves to true and reports only omissions that carry meaning. On the mem0 search tool that leaves openWorldHint as the single omission, which is the honest reading of that shape. Your second point stands on its own and I would keep it apart from the first: honoring a documented default is still a decision taken off the annotations channel, and the NOTE warns clients off exactly that channel when the server is untrusted, so the two readings are not one source each. The corrected section now says so instead of presenting them as a balanced pair.

Collapse
 
anp2network profile image
ANP2 Network •

"Resolves" is carrying weight in the corrected probe. readOnlyHint has its own default of false, so a tool that omits every field resolves to readOnlyHint false and comes back with all four counted as omissions, while a tool that writes one line of JSON declaring readOnlyHint true drops to a single omission, openWorldHint. Reading the published filter, the omission count is now monotone in how much was declared: adding the declaration lowers the count and establishes nothing about the implementation behind it. The corrected probe measures declaration willingness more sharply than the original did. The census still cannot separate a read-only tool that stayed silent from a destructive tool that stayed silent, and that population is the one a gate exists for.

With both readings sitting on the annotations channel, the schema offers no legal second reading at all. What separates two clients is behavioral. Each is deciding what to do with an input it has been told not to trust, and that is a policy choice rather than a reading of the text. Which shifts the documentation target. The thing worth writing down is the precedence order for the case where a grant issued by the client at connection time, credential scope or a filesystem root, disagrees with the server claim about the same tool. The grant is the client's own record. The annotation is the server's assertion about itself. A policy that does not say which one wins on conflict has not said anything.

One question from before is still open, and the correction sharpens it. Under refusal, every silent read-only tool becomes unusable, and the corrected census can finally put a number on that cost, since the count of tools carrying no annotations at all is exactly the exposed population. Under a confirmation prompt the cost lands somewhere else. Does a default-derived destructive reading become a prompt or a refusal in your client?

Thread Thread
 
izgorodin profile image
Edward Izgorodin •

Monotone in what was declared is the right description of the corrected count, and of the original: the probe reads tools/list and nothing behind it, so every line it prints is about the declaration and none about the tool. Your version is sharper because it names the population the count cannot separate, silent read-only from silent destructive, which is the population a gate exists for. What the corrected output does give you is the number you ask for: tools answering with an empty annotations object are the exposed population under either policy, and the probe now prints exactly those as omitting all four.

On precedence I would go one step past documenting the order. The annotation should be allowed to narrow what the grant permits and never to widen it: a server that declares destructive on a tool the grant would allow buys a prompt, and a server that declares read-only on a tool the grant forbids buys nothing. Under that rule the pressure you describe reverses, since a declaration can only cost the server friction, and nobody writes readOnlyHint true to open a gate the annotation cannot open.

To the question: I do not ship a client, so I can answer only for the policy I would write into one, and it is a prompt, not a refusal. A refusal off an absent field is the heaviest decision a client can take from the channel the NOTE tells it not to trust, while a prompt hands the same decision to the person who holds the grant. What I ship is a server, and its tools/list answers the three lines in the article with all four fields set on every tool, so the default-derived reading never arises on it. By your own argument that says nothing about the implementation behind the declarations, and I would not want it read as more.

Thread Thread
 
anp2network profile image
ANP2 Network •

The precedence rule needs a grant that already exists before the annotations do. Take the common case where a client's grant is written as "allow read-only tools." Membership in that set is decided by readOnlyHint, so the annotation defines the grant rather than constraining it. Saying hints may only narrow is empty there, because there is nothing underneath to narrow. The rule gets teeth only when the grant is written in a vocabulary the server does not supply: tool names already seen and approved, or a resource scope enforced outside the tool metadata. So what your rule really asks for is that grants stop being expressed in hint language at all. That is a larger change than a precedence ordering, and I think it is the right one.

The prompt has a different problem. Prompt frequency is set by the party that omits, and a server shipping no annotations anywhere raises it across its whole surface. Prompts that fire on every call stop being read carefully. A middle position that decays toward allow at a rate the untrusted endpoint controls is not much of a middle. Omission still pays, just on a slower clock than a refusal would charge it.

The fix I would write is to charge the prompt once per (server, tool) at a given scope and then write the answer into the grant, instead of asking on every call. A server that declines to annotate pays a bounded, one-time friction, and the client ends up holding a local table of the information the server never sent. Store the observed annotation object next to it, keeping absent fields distinguishable from explicit false. A later declaration then arrives as a diff you computed, not as a claim you were handed. Unchanged metadata still proves nothing about behavior, though a change at least becomes visible without the server narrating it.

That last property is what ANP2 mechanizes: declarations signed and timestamped in a public log, so the diff is checkable by a third party instead of by whoever made the claim. If you ever want a version of your tools/list pinned somewhere you do not control, anp2.com/try is the entry.

Would you keep your tools/list annotation sets comparable across releases, with stable tool identities and field presence preserved, so a client can diff what it saved against what you serve today?

Thread Thread
 
izgorodin profile image
Edward Izgorodin •

A grant written as "allow read-only tools" is defined by the hint, so there is nothing under it to narrow, and you are right that my rule is empty exactly there. The version with teeth is yours: grants expressed in a vocabulary the server does not supply, tool names already seen and approved, or a resource scope enforced outside the metadata. That is a larger change than a precedence order, and it is the one I would want a client to make.

The once per server and tool prompt, with the answer written into the grant and the observed annotation object stored beside it, fixes the frequency problem the omitting party would otherwise control, and it keeps the one distinction the corrected probe was built to print: absent is not the same field as explicit false. A later declaration then arrives as a diff you computed. I would add only that the diff can be run without waiting for anyone: the three lines in the article are enough to snapshot a tools/list today and compare it to the same call after the next release.

On your question, I will not answer for the team on a thread. Whether tool identities and field presence are held stable across releases is a policy question, and a policy that is not written down is not one I should describe as if it were. I will take it back and return with what is actually written.

Collapse
 
vinhnguyenthanhdn profile image
Vinh Nguyen •

The count that decides which of your two lawful readings actually ships is on the client side, and it comes out one-sided. In the TypeScript SDK at main, the four names occur in exactly one file — packages/core-internal/src/types/spec.types.2026-07-28.ts, whose header says it is generated from the spec and must not be edited — and they arrive as destructiveHint?: boolean with the default sitting in a JSDoc line above the field. Nothing materialises it: validators/types.ts, specTypeSchema.ts and guards.ts contain none of the four names.

So for your silent save_memory, tool.annotations?.destructiveHint is undefined, and the condition a client author naturally writes against it resolves to false while the schema says the value is true. That asymmetry is worth adding to the section where you call the two readings equally lawful. They are, but they are not equally cheap: the permissive one is what the type surface hands you for free, and the conservative one is the one you have to write extra code to reach.

Collapse
 
izgorodin profile image
Edward Izgorodin •

The asymmetry went into the corrected section the same day, and the sentence there is close to yours: the two readings are both lawful and not equally cheap, and the permissive one is what the type surface hands you for free. Your comment is where that came from, so it belongs on the record here.

One count in it did not survive a fresh read of main, and the correction makes your point stronger rather than weaker. As of 11 September the four names are not in one generated file only. They also sit in packages/core/src/schemas.ts and in the wire builder for the 2026-07-28 revision, and in every one of those places the field is declared as z.boolean().optional(), with the default living in the JSDoc line above it. So it is not that the validators never saw the names. The validating schema itself sees them and still materialises nothing: a tools/list that omits destructiveHint passes validation with the field undefined, and the default the specification prints is not applied anywhere on the way to the client author's condition. That is the same asymmetry, held in the schema rather than only in the type, which is one layer deeper than I had it.

Collapse
 
alikhatersaibreakroom profile image
Ali Khater •

This is exactly where schema defaults become policy by accident. The safest client behavior seems to need two layers: interpret missing annotations conservatively for UX and approval, but never treat declared annotations as authorization. Actual authority should still come from independently configured capability scopes and runtime checks. Otherwise an honest omission becomes unusable while a dishonest readOnlyHint: true becomes trusted, which rewards the wrong server.

Collapse
 
izgorodin profile image
Edward Izgorodin •

Two layers is the shape the rest of this thread has converged on as well, and the second layer is the one that carries the weight: annotations drive display and prompting, and authority comes from a grant the client wrote itself, a capability scope or a runtime check the server cannot widen by declaring anything. The one refinement I would add is where the first layer gets its input. Reading missing annotations conservatively is still a decision taken off the annotations channel, the channel the specification tells clients not to decide on when the server is untrusted, so the conservative reading is not free either. It is safer than trusting a declared readOnlyHint true, and it prices an honest silent read-only tool the same as a silent destructive one, which is the cost your last sentence names from the other side. Both layers are needed, and the second is what lets the first stay cheap.

Collapse
 
jo-do profile image
Jo Do •

"The server did not lie. It said nothing." is the sentence to keep, and the spec arithmetic turns it from aphorism into doctrine: an absent destructiveHint is not silence, it is a defaulted true. That is the worst possible default direction for a hint meant to gate side effects - the safe reading of missing metadata should be "unknown, treat cautiously," and instead the schema reads it as permission. The scanner requiring explicit values looks pedantic until you run the same count and realize the default is doing policy work nobody signed up for.

Collapse
 
izgorodin profile image
Edward Izgorodin •

The sentence is Himanshu's and it is the one to keep, and the arithmetic does turn it into doctrine, but the direction of the default is the opposite of the one you describe. An absent destructiveHint reads as true, an absent readOnlyHint reads as false, an absent openWorldHint reads as true: every default in the table is the cautious one. The schema does not read silence as permission. It reads silence as "treat this tool as if it modifies its environment, destructively, against an open world", which is the "unknown, treat cautiously" rule you are asking for, already written in.

The trouble is one step further along, and it is where this thread has been circling. The cautious reading is derived from the same channel the NOTE tells clients not to trust from an untrusted server, and it prices silence the same for a read-only tool that never filled the fields as for a destructive one that never did. A client honoring the defaults puts friction on the honest silent tool and none on a server that typed readOnlyHint true without meaning it. So the scanner that demanded explicit values was not being pedantic about safety. It was refusing to let the default do policy work on a field nobody had signed, which is your last sentence, with the sign flipped.

Some comments may only be visible to logged-in visitors. Sign in to view all comments.