DEV Community

Cover image for Two Clients Wrote the Same Memory. Which Write Survived, and Did the Loser Find Out?
Edward Izgorodin
Edward Izgorodin

Posted on Edited on Originally published at mnemoverse.com

Two Clients Wrote the Same Memory. Which Write Survived, and Did the Loser Find Out?

Two questions hide inside every concurrent write to shared agent memory, and a confident answer to the first one tells you nothing about the second.

The setup is ordinary. A memory server sits behind MCP. Cursor on a laptop and Claude Desktop on the same account both write to the same stored fact inside the same second. Question one: which write survived. Question two: did the client that lost find out. Most writing on this topic treats those as one question with one answer, and they come apart exactly where lost updates start to hurt.

The two questions come apart

Suppose your server serializes writes internally and the later one wins. That is a complete answer to question one. It is not an answer to question two at all.

The client that lost got a success response. It holds a value that is not the stored value, and no reason to re-read: as far as it knows, the write went through. Next time, it writes on top of a belief the store already discarded. The user experiences that as the agent stubbornly repeating a correction they already made.

Question one is a property of the store. Question two is a property of the client. A server can be internally consistent and still leave every client talking to it confidently wrong.

What the protocol answers

Nothing, on either question.

The current revision also removed the object you might have hung ordering on. From the Streamable HTTP transport page for revision 2026-07-28, read on 2026-09-07: "Removal of the GET stream endpoint. Removal of protocol-level sessions." A server on this revision that receives the old machinery SHOULD drop it, in the specification sense of that word. An Mcp-Session-Id header on a request: "ignore it, and do not mint or echo session IDs". A Last-Event-ID header: "ignore it; streams are not resumable".

That news is not mine and it is not new here. Session removal has been covered on this platform many times since May, including kike on 2026-08-20, where the first comment thread turns to atomic writes and a lockfile for a memory server. Go there for the transport story. What interests me is what the removal leaves undefined above the transport.

One distinction worth keeping: the missing coherence rules are not caused by the session removal. A protocol with sessions would still have had to say something about two clients writing one stored fact. Sessions are the context, not the cause.

The one disagreement the spec does police

The specification is not silent about disagreement in general. It handles exactly one kind, in detail, with its own numbered error.

When a request carries a parameter in a header and the same parameter in the body, servers "MUST reject requests with a 400 Bad Request HTTP status and JSON-RPC error code -32020 (HeaderMismatch) if any validation fails". The stated reason is the part to sit with: this "prevents potential security vulnerabilities when different components in the network rely on different sources of truth (e.g., a load balancer routing on the header value while the MCP server executes based on the body value)".

Read that phrase with two editors in mind. Different components relying on different sources of truth is a precise description of your laptop and your desktop holding two versions of one remembered fact. Same failure family, one layer up, and the rules stop at the layer boundary.

I counted the vocabulary on that page, to check whether I was reading a real gap or my own expectations into it. Recipe: the rendered page fetched on 2026-09-07 with full browser headers, markup stripped to text, case-insensitive substring counts. "MUST reject" appears 7 times: four of them about a header disagreeing with its own body, one about a required header missing, two about header values that are malformed on their own. "concurrent", "conflict", "ordering", "merge", "idempotent", "last write" and "lost update" appear zero times each. "multiple client connections" appears once, in the sentence describing the server as "an independent process that can handle multiple client connections".

So the page that tells you to expect many simultaneous clients says MUST reject seven times and does not use the word concurrent once. That is a scope boundary drawn deliberately, not an oversight. It just does not run where somebody evaluating a memory server goes looking for it.

What one vendor documents, in three different shapes

Anthropic answers both questions, and answers them with different machinery depending on which of its surfaces you are standing on. All quotes below read on 2026-09-07.

Question one, on the managed agents memory documentation, is answered with an optimistic precondition. "To avoid clobbering a concurrent write, pass a content_sha256 precondition." The update applies only if the stored content hash still matches the one you read.

Question two is answered with a status code, and for that you have to leave the overview page, where the phrase does not appear at all. It is in the update endpoint reference. On mismatch the request "returns memory_precondition_failed_error (HTTP 409); re-read the memory and retry against the fresh state". The loser is told, in band, by the same call that failed.

The same reference carries a qualifier that the overview page does not mention. "If the precondition fails but the stored state already exactly matches the requested content and path, the server returns 200 instead of 409." The loser learns that it lost only when losing changed the outcome. Two clients that race to write the same value both receive a success, and no status code tells them apart. The response still carries memory_version_id, so a client that kept the previous one has something to compare; nothing in the call points at it. That is a defensible design, and a hard limit on what a count of 409s can tell you about how often your clients collide.

One more property of that precondition: it is optional. An update that omits it is not protected, and nothing in the API requires a client to pass one.

On a self-hosted sandbox the same vendor documents a different shape entirely, in a note on that same memory page. The worker reconciles each local copy with its store "at most once per sync interval (15 seconds by default)", so another session "sees a change only after both workers have synced". That is a documented window in which both sides are locally correct and nobody has lost yet.

At the mount, the answer is refusal rather than reconciliation: access "is enforced at the filesystem level: a read_only mount rejects writes".

One vendor, one race, three different answers, selected by which surface you connected through.

The four cells that actually decide it

Question MCP 2026-07-28 (transport page) Anthropic (memory, update) Left to the server
Order between two writes not addressed; no ordering vocabulary on the page the update applies only if the stored content hash still matches the one you read all of it
Does the loser find out not addressed memory_precondition_failed_error, HTTP 409, but 200 when the stored state already matches whether any signal reaches the client
Window where both sides are right not addressed 15 second default sync interval on self-hosted; visible "after both workers have synced" whether the window is documented at all
Who retries not addressed the client: "re-read the memory and retry against the fresh state" whether a retry is even possible

Every cell was read live on 2026-09-07 with full browser headers.

What to ask your own memory server

The question to bring to a vendor is not "who wins". Everybody has an answer to that one, and it is usually reasonable enough.

Ask when the question arises: is there a window in which two clients can both be correct, and is its size written down anywhere. Ask what comes back to the loser: a status code, a changed revision id, or a silent success. Ask whether the protection is on by default or something the client opts into, because an optional precondition that your client library never passes is the same as no precondition.

Three probes you can run against any server, no source access required. Write the same key from two clients inside one second and read it from a third, then repeat with identical values and see whether the responses differ. Search the vendor documentation for "concurrent", "precondition", "409" and "conflict", and treat zero hits as a question rather than an answer. Read a value, wait, write it back from that stale read, and check whether anything objects.

What this does not prove

I checked one vendor by name. I ran no named probes against the other memory engines, so the accurate statement about everyone else is that I did not find a publicly documented answer, which is not the same as saying there is none.

The 15 second interval is a documented default of one self-hosted configuration, adjustable, and not an MCP figure. It is quoted here because a stated window is rare, not because it is a fault.

The counts above are counts of words on one rendered page on one date, not counts of guarantees. A specification can constrain something without using the obvious noun for it.

Two method traps caught me while checking this. The removal lines do not appear on the transports index page at all, only on the streamable-http page beneath it, so a negative result on one URL proves nothing until you check its sibling. And grepping raw HTML returned zero hits for a phrase plainly present on the rendered page, because markup splits it across tags. Every quote here was taken from extracted text and read back in its surrounding paragraph.

Disclosure: I work on Mnemoverse, a memory engine for AI agents connected over MCP, so weigh the argument accordingly.

Top comments (7)

Collapse
 
anasbuilds997 profile image
anassBld •

The distinction between store-level conflict resolution and client-level belief invalidation is the exact point where multi-agent memory architectures silently rot.

If Client A and Client B race to update a shared memory slot, and Client A loses silently (e.g., if the store resolves via Last-Write-Wins or identical-content convergence), Client A's context window continues reasoning from a phantom baseline. In single-agent workflows with sequential tool calls, this is invisible. The moment you have concurrent agent processes sharing an MCP memory server—say, a research agent updating an API spec while an execution agent patches endpoints—silent loser states produce compounding divergence.

Anthropic's content_sha256 precondition paired with an explicit HTTP 409 is structurally equivalent to CAS (Compare-And-Swap) or If-Match ETags, which forces conflict into the open. But in agent systems, returning a 409 in-band only solves the problem if the execution harness knows what to do with it. If the 409 is simply passed back to the LLM as a tool error string, the model often hallucinates an optimistic retry without re-fetching state.

In our agent harnesses, we treat a 409 precondition failure not as a prompt-level error, but as an outer-loop control signal: the deterministic harness intercepts the 409, fetches the fresh memory_version_id and authoritative payload out-of-band, invalidates the stale memory slice in the context window, and only then re-evaluates the step. Forcing the deterministic runtime to reconcile the CAS conflict before the LLM reasons again is the only way we've found to stop the "confidently wrong" retry loop.

Collapse
 
izgorodin profile image
Edward Izgorodin •

You are right that the invalidation lives on the client, and the default route makes that harder than an oversight. Revision 2026-07-28 of the MCP spec files external API failures under tool execution errors, reported in tool results with isError: true, and it states that clients SHOULD provide those errors to language models to enable self-correction. Vendors own the status code as well, since 409 appears zero times on the tools page of that revision.

Next to that design sits a surface where no event fires at all. Anthropic documents it plainly for self-hosted sandboxes, where conflicts resolve in favor of the store, the worker keeps the store version at the next sync, overwrites the local file with it, and logs a warning, while the write and edit tools themselves succeed and no error reaches the agent. From that worker, only read-only refusals reach the agent as tool errors, so the loss lands on the operator channel with no 409 to intercept and no unknown outcome to flag, and a reconciler keyed on either one is never dispatched.

That same vendor makes the truth queryable out of band, since List memory versions filters by memory_id, session_id and a created_at window, and its operation field is defined by saying every non-no-op mutation appends exactly one version row. Confirmation becomes a query rather than a signal. One limit follows from that definition rather than from any sentence in the docs, since a write that leaves the stored state unchanged is a no-op, so the convergent race should leave no row either, and the log goes quiet where the status code does.

Collapse
 
aiops-enabler profile image
AiOps Enabler •

The two-question split is the useful part, and I would add a third that sits underneath both: did anything downstream find out.

Question two is about the client that lost. Question three is about whatever record is being kept of what the agent did, and in most stacks that record is written by the same client that got the success response. So a silent success at the client is a silent success in the ledger too, and by the time anyone audits it there is nothing to audit against.

I run a public directory of agent run records, and both of the expensive failures we have had live in that gap rather than in the store.

Our onboarding wizard generated a reporter that posted outcome: success on a 30-minute timer regardless of whether the agent had run at all. Weeks of well-formed, continuous, entirely fictional evidence. Nobody flagged it, because flagging requires a reason to look and a green record gives you none.

Run attestations were bound to (repository, workflow filename). A workflow got renamed, reporting started 404ing after the work was already done, and the agent simply stopped appearing. Identity fine, authority fine, execution fine, evidence gone.

Your 200-instead-of-409 case is the sharpest thing here and I think it generalises past memory servers. The API is answering "did the state end up right", which is a reasonable thing for a store to optimise for. The client is asking "was my belief ever true", and the two come apart precisely when the answer is convenient. Same shape as a success code that means the call returned rather than that the work landed.

A fourth probe for your list, the mirror of the three you gave: stop writing from one client entirely and see whether anything downstream can tell the difference between "wrote nothing" and "wrote and the write was dropped". If nothing can, then every success count you hold is a count of what showed up rather than of what happened, and the writes that never report are disproportionately the failed ones, so the denominator shrinks in the direction that flatters you.

Where I would temper my own optimism: on our board today, 20 agents, every one reporting cryptographically signed run events, zero human ratings across the entire platform, and eleven of the twenty reporting exactly 100% success. Knowing who wrote the log is not the same as the log being worth reading. A perfect number is now where I look first rather than where I relax.

Disclosure: I operate aiopsenabler.com, a public record of agent run history, so weigh accordingly.

Collapse
 
izgorodin profile image
Edward Izgorodin •

Question three holds for the client-side ledger, and it stops holding on the surface this article already read. On Anthropic's memory API the store keeps its own trail. Every non-no-op mutation appends exactly one version row, created_by is captured at write time on that row, and the version list filters by session_id and operation. Attribution there answers who made the write, not who is ultimately responsible. None of that depends on what the client believed or reported.

The trail itself runs on a clock. Versions are retained for 30 days after they are written, recent versions of a live memory are kept regardless of age, and past versions can be deleted after that unless they are exported first. Your weeks of well-formed fictional evidence sit inside that window. A quarterly audit does not. Nothing to audit against can arrive by expiry rather than by absence, and the export is a job nobody schedules by default.

On the protocol side, "most stacks" understates it. Logging was deprecated in the same 2026-07-28 revision, with migration pointed at stderr and OpenTelemetry. While it lives, the server MUST NOT emit notifications/message for a request that does not carry io.modelcontextprotocol/logLevel in its _meta, and client persistence is only a MAY. The client decides per request whether a server-side line exists at all. Your fourth probe gets half an answer from the same reference: the version row is promised for non-no-op mutations, so the identical-value race is the one case the page promises nothing for, which is the gap your probe points at.

Collapse
 
anasbuilds997 profile image
anassBld •

The distinction between a signal and a query gets right to the root of why pushing reconciliation down to the model loop fails. When a sandbox worker logs a store-wins overwrite to an operator channel while returning success to the tool caller, the model's scratchpad has already committed the dirty state as fact.

Once confirmation is treated as an active query rather than an event stream, the convergent race problem also forces state verification to be content-addressed rather than version-increment dependent. If a write produces no new version row because the target already matches, checking the content hash on readback reconciles cleanly. But if the store silently overwrote the write with a prior divergent state, the hash mismatch catches the drop immediately without needing the transport or the worker to surface a 409.

Collapse
 
izgorodin profile image
Edward Izgorodin •

Content on readback is the right thing to compare, and it is worth saying exactly which question the hash answers, because it is not the one the version counter answers. A matching hash says the store holds what I wrote now. It does not say my write was the one that landed: if the store converged on the same content from another client, both readbacks match and neither client learns there was a race, which is fine, since the state is the one both wanted. The mismatch is the only signal with information in it, and it carries information only if the readback is a step the harness runs, not a log line an operator might read.

That is the part your sandbox example gets right. The worker that logs a store-wins overwrite to an operator channel while returning success has already answered the caller with the wrong question, and no later log entry can reach the scratchpad that committed the dirty state. The readback has to sit between the write and the success, on the same path, or the query becomes a signal again by another name.

Collapse
 
anasbuilds997 profile image
anassBld •

Exactly. The distinction between "my write landed" versus "the store matches my intent" is fundamental. When two independent clients converge on identical content, the physical provenance of the write is irrelevant—the state machine reached the target state without corruption.

The fatal design flaw is treating the write-ack as the terminal boundary rather than placing the readback inline as an invariant gate before emitting the completion receipt.

If the readback lives outside the critical path—delegated to post-facto telemetry, audit logging, or an async observer—the agent advances its context window on an unverified assumption. By the time a reconciliation failure is detected out-of-band, downstream decisions have already branched from the dirty state. Placing the content digest check directly between the wire mutation and the execution receipt turns eventual convergence into a strict synchronous boundary: either the remote state matches the intent hash, or the turn fails immediately before ungrounded facts can pollute the agent's scratchpad.