Your coding agent's memory does not fail loudly. It fails the way a slow leak does: the agent just knows a little less every week, and you blame the model.
I've been running my company's products through Claude Code sessions for five months, with a memory of short markdown notes behind them — 417 notes today, about 440,000 words, every one written while actually shipping something. Over that time I kept a list of every way that memory broke. Not one of them raised an error.
Today we're releasing CortHeXis 2.0, the engine that came out of that list. It's open source (Apache-2.0), it runs alone next to any agent that speaks MCP, and it's the same code that runs our own memory. This post is about the failures first, because they are the product.
Four ways a memory dies quietly
1. The memory server vanished for nine days. A tool rewrote our project's .mcp.json to add a browser server, and dropped the memory server in the process. Every session kept working. None of them could search the memory. Agents started answering from an index file that had been truncated to fit the context window. It took nine days to notice — because nothing was wrong, things were just worse.
2. One colon broke recall for a note. A note whose description: contains an unquoted : is invalid YAML. The indexer skipped its header, the note lost its description, and with it the few words the search was matching on. It still existed. It was just never found again.
3. A ghost note. A session appended a fact with cat >> pricing_rules.md while the note was called pricing-rules.md. A new file appeared, with no header, carrying a corrected fact — while the old note kept the wrong one. Two truths, and the search happily returned the stale one.
4. A better model made recall worse. We switched embeddings from MiniLM to EmbeddingGemma and recall jumped on our bench. But the hook that injects notes at every message kept the score threshold I had tuned "by eye" for the old model. Measured afterwards: it was injecting the right note in 22 % of cases. Re-tuned on the bench, not on examples: 87 %. Scores don't mean the same thing from one model to the next.
None of these is exotic. Each was invisible from inside a session.
What CortHeXis does about it
Recall at every turn — including sub-agents. A Claude Code hook (UserPromptSubmit) puts the relevant notes in front of the agent before it answers, not only when the session starts. A second hook (PreToolUse on Task|Agent) does the same for every sub-agent it launches. Sub-agents start with an empty context, and that's where most of our memory used to get lost. The injected block is framed as data, not instructions: notes are written by agents, and a note can carry an order.
A review that re-reads everything, every hour. Twenty-five checks, each born from a failure like the ones above: the MCP server not declared or not answering, the embedding server down, the index lagging behind the files, broken links, notes with no description, a description that contradicts the body, near-duplicates, projects still "in progress" two months later, secrets pasted into a note, and a mass rewrite that made every note look new. You get a score, findings with their remedy, a history, and a daily digest to your webhook. On the public demo, it looks like this:
score 92
warn broken_links 1 [[modele-restitution-2025]] cited by methode-restitution-en-une-page
→ point the link to the note that replaced the old one
info orphans 3 notes with no link in or out: found by search only
That broken link is real, not staged: the note was renamed and one reference was forgotten. A demo that shows 100/100 proves nothing.
Repairs you approve. Naming is fixed losslessly at indexing time (and links written with the old name are reattached). Anything that needs judgement — relinking, merging two notes, closing a stale project — comes as a proposal with a diff, applied all-or-nothing once you approve it, with a backup.
Dates you can trust. Every recalled note carries its age and where that date comes from: written in the note, reconstructed by the indexer, or unknown. A script that touches every file no longer makes the whole memory look fresh.
Measured, not promised. corthexis eval builds a recall bench from your own sessions: a prompt followed by memory_get(note) is a question with a known answer, no annotation needed. A model or profile change is judged on the same questions, and a drop in recall blocks the switch.
The numbers, and what they don't say
On our private memory, 300 questions (65 % French, 25 % English, 10 % German), hybrid search:
| hit@1 | MRR | |
|---|---|---|
| 1.x engine (MiniLM) | 42 % | 0.55 |
| 2.0, CPU, 4 cores, no GPU | 73 % | 0.82 |
| 2.0 with a reranker on a GPU | 82 % | 0.88 |
The biggest gap is on questions whose answer lives only in a note's body: MiniLM reads the first 128 tokens of a passage and barely sees them.
You can't rerun that bench — the corpus is full of infrastructure and clients, and it stays private. So 2.0 ships a public bench you can: 60 questions in French, English and German over the fictional demo corpus, rerunnable in one command against your install. Honest caveat, written in the repo too: 38 tidy notes are an easy corpus. Every model lands at MRR 0.89–0.97 there; the bench shows the ordering (MiniLM drops to 0.73 on German questions over French notes) and lets you check your install gives the same figures. It doesn't replace the bench you build on your own notes.
Try it
git clone https://github.com/ninabot-ch/corthexis && cd corthexis
./setup.sh ~/.claude/projects/<your-project>/memory # or any folder of markdown notes
docker compose up -d # dashboard on http://localhost:8420
claude mcp add -s user --transport http corthexis http://localhost:8420/mcp \
--header "Authorization: Bearer $(sed -n 's/^CORTHEXIS_TOKEN=//p' .env)"
python3 corthexis/hook.py install # recall at every message + sub-agents
Docker, about 2 GB of RAM, four cores, no GPU. Postgres with pgvector for the index; the notes stay plain markdown files you own, and the index can be rebuilt from them at any time. The default embedding model is downloaded on first run after you accept its licence, with an MIT-licensed fallback if you don't. After that, nothing leaves your machine to index or recall.
- Repo: github.com/ninabot-ch/corthexis
- Live demo, fictional corpus, no account: demo.corthexis.com
- Site: corthexis.com
If you already use Mem0, Zep, Letta or your framework's RAG: CortHeXis isn't trying to retrieve better than they do. It starts from the next question — how do you know that what your agent retrieves is still there, still true, and still found? I'd love to hear how your own agent memory has failed. I suspect most of those failures never threw an error either.

Top comments (0)