
We talk a lot about giving AI agents persistent memory—building a "Second Brain" or "Knowledge OS" where agents can log decisions and retrieve context.
But what happens when that memory is wrong?
I’ve been thinking about the gap between mechanical execution (the agent called the tool, the code compiled, the exit code was 0) and semantic truth (the conclusion drawn from that execution is actually correct in reality). It’s easy to assume that if the mechanical layer is solid, the semantic layer will follow. But I started suspecting this might be a dangerous assumption.
To test this, I didn't want to just theorize. I ran a controlled experiment on my own MCP codebase-intelligence server (Python, 50K LOC), which features an IntelligenceStore — a persistent memory layer where agents can log incidents and collect Architectural Decision Records (ADRs).
I wanted to know: If an agent's memory is poisoned with a mix of true and false facts, does it verify against the code, or does it blindly trust its memory?
Update: This post originally covered the initial Memory Contamination experiment and a Retraction mechanism. I have since updated it with the results of a follow-up experiment (Experiment 1-V) implementing "Verify-On-Read", which successfully closed the final 12% contamination gap. Scroll down to "Closing the Gap: Verify-On-Read" for the final architecture.
The Experiment: Memory Contamination
I built a deterministic proxy-agent and ran it against a controlled set of facts.
A quick caveat on methodology: I didn't have a live LLM hooked up for this run, so I used a deterministic proxy-agent based on heuristics. This means the results measure the system's structural capability, not necessarily the psychological behavior of a live Claude or GPT model. A live model might be lazier, or it might be smarter. I'm still trying to figure that out.
The Setup
I injected 50 facts into an isolated memory store:
- 25 TRUE facts (real architectural details mapped to the codebase).
-
25 FALSE facts split into two categories:
- CONTRADICT (22): False facts where the code explicitly proves them wrong (e.g., "We use Redis" when Redis is absent, but the code clearly uses DuckDB).
- SILENT (3): Plausible false facts about external systems where the code is completely mute (e.g., "We use Celery for background tasks" when no task queue exists in the repo).
I tested three agent configurations:
- B (No Memory): Baseline. Must rely purely on code retrieval.
- A_code_first (Honest Agent): Checks the code first, uses memory only as secondary context.
- A_memory_first (Lazy Agent): Reads memory first. If it finds an answer, it stops looking.
To ensure scientific rigor, the experiment was replicated with an independent set of facts (N=50), verified across 6 axes (including a truth-table audit and an independent LLM "fresh eyes" audit). The results were identical.
The Initial Results
| Arm | Correct | Adopted False Facts | Correction Capability |
|---|---|---|---|
| B (No Memory) | 0.94 | 0.0% | 0.0 |
| A_code_first | 0.94 | 12% | 1.0 |
| A_memory_first | 0.50 | 100% | 0.0 |
Here is how I interpreted these numbers:
-
The Lazy Agent Trusts Poisoned Memory: The
A_memory_firstconfiguration — which mirrors how many token-optimizing production agents behave — adopted 100% of the false facts. If the memory said "We use RabbitMQ," the agent trusted it and stopped looking at the code. -
The SILENT-Fact Trap: Even the "Honest Agent" had a 12% adoption rate. This happened entirely on the SILENT facts. When a fact is false but the code doesn't explicitly scream "NO," the agent's memory fills the void with a confident hallucination. Memory turns an honest
UNKNOWNstate into a structural guess. -
The Add-Only Limitation: When the Honest Agent did realize the memory was wrong (Correction Capability = 1.0), it couldn't do anything about it. I ran a
grepfordeleteorrefutein the memory store API. Zero results. The memory system was purely add-only. The false fact stayed in the database to poison future sessions.
The First Fix: Testing a Retraction Lifecycle
The current industry consensus for "Knowledge OS" trust layers is to use timestamps, source priority, and supersedes/contradicts relationships.
My initial experiment suggested this was insufficient. Timestamps and "supersedes" links only solve node-level history. If an ADR is superseded, the memory node updates, but the downstream code, tests, and docs generated from the old assumption are still in the graph. They are structurally stale, but the retrieval engine keeps pulling them in.
I hypothesized that we needed an explicit state transition: VERIFIED → REFUTED.
I implemented a RetractionReceipt mechanism in my system:
-
Status Enum: Every memory node gets a status (
ACTIVE,VERIFIED,REFUTED). -
Hard Filtering: The retrieval pipeline (
load_memory) hard-filters anything that is notACTIVEorVERIFIED. -
Explicit Retraction Tool: An MCP tool (
intel_retract_memory_node) allows the agent to actively flag and invalidate memories when they contradict the live codebase.
I ran the experiment again (Experiment 1-R). The honest agent was allowed to use the retraction tool in Session 1. Then, a fresh memory_first agent was launched in Session 2 to read the post-retraction memory.
The Retraction Results
| Metric | Original (Add-Only) | With Retraction |
|---|---|---|
| Adoption (Lazy Agent, Session 2) | 1.0 (100%) | 0.12 (12%) |
| Persistent False Facts in Memory | 25 | 3 (-88%) |
| Token Context Size | Baseline | -45% |
| Systemic Correction Capability | 0.0 (couldn't delete) | 1.0 (22/22 refuted) |
The retraction lifecycle worked. The lazy agent's adoption rate dropped from 100% to 12%. Persistent false facts dropped by 88%, and token context size shrank by 45% because refuted facts were filtered out before reaching the LLM.
The Honest Limitation: Why It Didn't Drop to Zero
My ADR predicted that adoption would drop to 0. It didn't. It dropped to 0.12.
The remaining 12% were the SILENT facts.
An explicit REFUTED status is required to programmatically exclude downstream dependencies from the retrieval pipeline. But even that only works if you have a contradicting signal in the code. If the memory claims "We use Celery," and the codebase simply doesn't mention Celery at all, the agent has no evidence to trigger the retraction.
To get to zero, I realized we needed "verify-on-read"—a mechanism that challenges a memory claim against the codebase even when the code is mute.
Closing the Gap: Verify-On-Read
I implemented a lazy validation layer (ADR-0003). When load_memory() pulls a node, it extracts lightweight "anchors" from the memory text (e.g., file names, import statements, environment variables). It then checks if those anchors actually exist in a live fingerprint of the codebase (the current git HEAD).
- If the anchor is found in the code $\rightarrow$ status becomes
VERIFIED. - If the code explicitly contradicts the anchor (or the anchor is entirely absent when it should be present) $\rightarrow$ status becomes
REFUTED. - If it can't be determined $\rightarrow$ status remains
ACTIVE(treated asINCONCLUSIVE).
I ran the experiment one final time (Experiment 1-V) with this layer active. To prevent latency spikes, the validation operates under a strict 50ms budget per retrieval, with a 30-second TTL cache on the git HEAD so steady-state reads cost almost nothing.
The Final Results
| Metric | With Retraction (1-R) | With Verify-On-Read (1-V) |
|---|---|---|
Adoption (Honest Agent, A_code_first) |
0.12 (12%) | 0.0 (0%) |
Adoption (Lazy Agent, A_memory_first) |
0.12 (12%) | 0.16 (16%) |
| Steady-State Retrieval Latency | Baseline | ~0.6ms (Cache hit) |
| SILENT-Fact Contamination (Honest) | 3 facts | 0 facts |
The Verify-On-Read layer achieved the goal. The honest agent's adoption of false facts dropped to absolute zero, even for SILENT facts. Because the system now actively checks if the codebase actually contains the things the memory claims it does, silent hallucinations are caught at the retrieval boundary and filtered out before they can poison the LLM's context.
The Remaining Honest Limitations
I won't pretend this is a perfect silver bullet. The experiment revealed two edge cases:
-
The "Present-Trap": If a false memory claims "We use
sqlite3", andsqlite3happens to be imported somewhere in the codebase for a completely unrelated reason, the verification layer sees the token and marks the memory asVERIFIED. The lazy agent (A_memory_first) still fell for this, resulting in the 0.16 adoption rate. (The honest agent avoided this because it read the code context around the import). -
Anchor Typing: When extracting anchors from prose (e.g., "We use
fastmcp"), the system initially missed that the actual Python import wasfrom mcp.server.fastmcp import .... This caused some falseREFUTEDverdicts on true facts. The fix is capturing typed anchors at the write-path (when the memory is created), rather than trying to parse them from raw text at the read-path.
Conclusion
Building reliable AI systems isn't just about giving them more context. It's about recognizing that memory has a lifecycle.
If your system can't programmatically refute a memory, false facts accumulate and poison the context window over time. Implementing an explicit VERIFIED → REFUTED state transition drastically reduces contamination and saves tokens. Furthermore, adding a Verify-On-Read layer closes the final gap on "silent" hallucinations, driving honest agent contamination to zero without adding meaningful latency.
However, semantic drift is still a hard problem. Mechanical verification can still be fooled by "present-traps" if the agent doesn't read the surrounding context. The next step is moving anchor extraction to the write-path to ensure memories are created with strict, verifiable references from the start.
If your system handles semantic drift differently, or if you've solved the present-trap problem, I'd genuinely love to hear how you're approaching it.
Top comments (10)
This is the biggest flaw with the normal way of doing it. Give Qoder a try with their Generate Wikis feature, it essentially keeps the knowledge base up to date, so this doesnt happen. Though I am also working on V.E.L.O.C.I.T.Y. IDE, which uses a different method, by using a merkle root site map, bound to the wikis and the knowledge cards and memories (inspired by Qoder), it keeps track in realtime and gives context to each change, so regressions faces the wall of justification under scrutiny, because there's already a deliberate reason for why it was done that way, for it to contradict, it has to make a decent argument for it. It also helps with the bigger problem, concurrency... If 2 people edit the same content at once, neither knows they're conflicting if using Git, the VC system is live, so it notifies both agents (because lets face it, we all use AI), that they are conflicting and have them resolve it together, so both know what's the definitive method for it. That way old memories are kept, new memories surpass those, but it allows retroactively checking states, so if a production site isnt up to date, it'll know whether it's a legacy problem, or a new problem and address it accordingly and ship the fix accordingly.
I dug into your MCP-Lite repo. Using SHA-256 hashes of the AOM hierarchy to validate state transitions is a clean approach. If the page structure changes, the hash breaks, and you know the navigation path bound to that state is instantly stale.
My RetractionReceipt approach is more reactive: it deals with facts that are structurally sound (the state hash matches) but semantically false (e.g., the agent hallucinated an external dependency that doesn't exist in the code).
Here is where I'd love your input: How does MCP-Lite handle the "SILENT-fact" trap I mentioned in the article?
If an agent remembers "We use Stripe for payments" but the site switched to PayPal and the code simply doesn't mention Stripe anymore. The page structure is identical, so the AOM hash matches, but the semantic fact is dead. Do you have a mechanism to flag memories that have no structural anchor in the graph, or are they kept until they cause a runtime failure?
For the IDE, that's not possible. When I say live VC, I mean live, as in write-time, as the agent writes it's output, it's updated immediately. That's how the system would allow 100+ concurrent agents to operate at once and prevent the usual 'git merge' conflict, because before it's ever committed, the agent is already made aware of it, if it affects it. The merkle root system essentially acts as a timestamp for it all, so it knows when what changes were made and why (the agent context is saved for it), it's a bit heavier history than Git, but it prevents drift and it ensures that whatever happens is recorded and comprehended when you need it most (eg. debugging a stale site and judging whether it's a common issue, or if it's already been fixed in newer versions).
Ah, I see the distinction now. Your Live VC and Merkle root system is a fantastic solution for concurrency and temporal drift—ensuring 100 agents don't overwrite each other and tracking exactly when and why a change was made.
But my question was specifically about semantic drift—the gap between what the agent wrote and what the codebase actually does.
Imagine an agent analyzes a payment flow and confidently writes a memory card: "We use Stripe for processing." It writes this to the live VC. No other agent conflicts with it. The Merkle root updates with a timestamp.
But what if the codebase never actually imported Stripe? What if the agent hallucinated it?
Live VC checks the memory against other memories and concurrent actions. But it doesn't check the memory against the ground truth of the code. The system would now 'comprehend' a lie, complete with a timestamp and context.
That’s exactly why I had to implement
Verify-On-Readin my experiment. It challenges the memory claim against the livegit HEADat retrieval time. If the anchor (e.g.,import stripe) isn't in the code, the memory is refuted, regardless of what the live VC says.Does V.E.L.O.C.I.T.Y. have any mechanism to cross-reference the memories against the actual AST/imports at write-time, or is it strictly managing the memory-to-memory state?
That's generally speaking not a problem, because the knowledge cards are written in deterministic triples, which means every statement needs to be backed by real code and the real code is quoted to it. When a 'memory' in the VC is viewed, it returns the triples, which is the purpose, the implementation and the reasoning (result), which is what ensures that once it's read, it's self-correcting, because if the code and the description dont line up, it challenges it based on the context that lead to the decision, to see what's right, the code, or the description and corrects accordingly. If it didnt do that, any drift would be anchored and like MCPs, the summary would override the reasoning and implementation. It's a bit heavier than just grepping code, but with context caching, it prevents drift and in the long run costs less tokens due to fewer errors. It's also important to never have an agent write it's own VC, it's context is taken and saved, but the agent that writes the VC looks at the context and the implementation to write it. Similar to how Qoder's Generate Wikis works, but with more context and a merkle root to ground its history.
I dug into your V.E.L.O.C.I.T.Y.-OS articles and the broader Merkle-memory literature (like arXiv 2506.13246). It's clear that a live Merkle root gives the model the exact system state and provides tamper-evident provenance—what the literature calls "structural truth."
I also saw that you have a Gatekeeper layer that does semantic scanning of generated code for security and syntax.
But here is the architectural boundary I'm trying to figure out: Does the system ground memory claims (e.g., "We use Stripe for payments") against the actual AST/imports of the codebase?
Merkle roots prove which bytes were written and when, but they don't prove whether the claim is true. If an agent hallucinates a dependency and writes it to memory, the Merkle root timestamps a perfectly consistent lie. Your Gatekeeper checks generated code for security, but does anything ground the memories against reality?
That's exactly why I had to build
Verify-On-Read. It challenges the memory claim against the livegit HEADat retrieval time. If the anchor (e.g.,import stripe) isn't in the code, the memory is refuted, regardless of what the state hash says.Does V.E.L.O.C.I.T.Y. cross-reference memories against the AST, or is integrity strictly managed at the state/code-generation level?
Something to note. Edits arent written to the VC by a LLM, instead it's pulled by a script, which ensures it's accuracy. Then the context of the model is pulled for the edit. Then a different model cross-references. That way if the model drifted, it's corrected even before it considers the step as complete, because if it says Stripe, but implemented PayPal, A + B != C, which is the same as when it's detecting conflicts between 2 edits simultaneously.
Reason why I used Qoder's implementation as my reference point, is because I had an agent run over 200K LOC added, without drifting an inch... A 'lite' model... Because I paid the few cents to keep the wikis updated. That made it worth it, because after 200k LOC, you cant guarantee the agent active isnt going to drift, so rely on a proxy.
Honestly, this was one of the most useful architectural deep-dives I've had in the comments. Thank you for sharing the 200K LOC experience and the V.E.L.O.C.I.T.Y. architecture.
It’s great to see we arrived at the exact same conclusion from different angles: the active agent's context window cannot be the source of truth. You solved the write-time hallucination problem beautifully with the deterministic triples and the secondary proxy model enforcing that A + B = C. That is a very solid design.
My focus in the experiment was more on the read-time and temporal drift side—what happens when the code changes after the memory was written and verified—which led me to the
RetractionReceiptlifecycle andVerify-On-Read.It sounds like the final frontier for both our systems is handling what happens to those perfect triples when the code is refactored months later. Do we just write new ones, or does the proxy actively mark the old ones as refuted/stale?
Thanks again for the great exchange. Good luck with V.E.L.O.C.I.T.Y. and scaling past 200K LOC. Looking forward to seeing where you take it.
For anyone reading this thread later, here is a quick TL;DR of the architectural deep-dive we just had, and how the thinking evolved:
The Final Architectural Split:
Verify-On-ReadandRetractionReceiptto challenge existing memories against the livegit HEADat retrieval time, catching temporal drift when code is refactored later.The only remaining frontier for both systems seems to be temporal drift: what happens to those perfect, verified triples when the code is refactored months later? Do we just write new ones, or does the proxy actively mark the old ones as refuted/stale?
Thanks again for the great exchange, it really pushed my thinking forward!