DEV Community

Cover image for I tried most of the memory systems for AI agents. None did what I needed, so I built one.
BrethofAI
BrethofAI

Posted on

I tried most of the memory systems for AI agents. None did what I needed, so I built one.

You open a session with your agent the way you open a hundred others. Before any real work: the stack, the conventions, the decision from last Tuesday about why the auth flow is the way it is, the thing you tried in March that didn't work. You paste it in, or you point the agent at a markdown file you maintain by hand, or you type "as I discussed before" and hope. Then you work. Then the session ends and everything you spent establishing is gone.

If you run an agent daily, this is your routine, and you've stopped noticing it's a routine. "Save context" before closing. A CLAUDE.md you edit like a wiki nobody owns. Five md files of notes that are really just a memory you're operating by hand, for software that costs more than your editor.

Here's the strange part. The smartest tool on your machine is the only one that remembers nothing. Your editor remembers your settings. Your terminal remembers your history. Your agent — the thing that costs the most to brief — starts from zero every morning, and you've made peace with it.

We hadn't. We tried the fixes on offer, they all failed the same test, so we built one that doesn't. This is what it is, what it does on an ordinary Tuesday, and the one sentence that installs it.

Three things that were supposed to fix this

Judge every memory tool by one question: when the next prompt arrives, what does the agent actually know? Not what's stored. Not what's findable. What's in front of the model, right now.

Markdown files — CLAUDE.md and its cousins. Loaded once at session start, then frozen. Whatever you write there on Tuesday is what the agent believes all week, and nobody maintains the file because maintaining it is the same manual chore you were trying to escape. Within a month it's a junk drawer: three conflicting instructions, a stack list from two projects ago, a note-to-self that stopped being true. A snapshot, not a memory.

Memory tools the agent has to call. These store things well. But they wait to be asked: nothing arrives before the agent answers unless it decides, mid-turn, to search, and remembers to. And at session end the handover is whatever the next session thinks to ask about. A library you have to remember to use is not a memory.

Search hooks. The better ones do push something with every prompt: a few snippets of raw text that matched. That's real injection, and it's still not enough — snippets are retrieval, not briefing. They find fragments. They don't carry your standing rules, don't reconcile a fragment against what you decided later, and don't come with a reason attached. The agent gets a search result; it doesn't get told what's true.

All three fail the same way: the knowledge exists somewhere, and the moment of the prompt — the only moment that matters — gets nothing, or almost nothing.

The test, then

A working memory has to answer at three moments:

  • Session start. A brief, before you say anything: where we left off, the rules in force, what's next on the plan.
  • Every prompt. The records that bear on that specific prompt, arriving with it — without dumping the whole memory into your context window every turn.
  • "Why is it like this?" The decision, the date, the reason, and the option it beat — in the words of the conversation where it happened.

That's the spec. Everything else is implementation detail.

One working day with a brain instead of a memory

We run our own company's work on our own memory — it has run our estate since December 2025, archiving every conversation since April. This isn't a demo script; it's Tuesday.

Morning. Start the machine, open the session, say "continue." The brief is already there — the last session's handover note, the project rules, the next lines of the plan, an index of the playbooks for this project. No re-explaining. No "quick recap of the repo." The session picks up the top line of the plan and starts.

Midday. Each prompt arrives pre-briefed. Not with everything — with the records that bear on that prompt, and nothing else. When I ask about the rate limiter, the memory of the rate limiter conversation comes with the question. I stopped keeping md files in March and haven't missed them; there's no "save context" ritual because there's nothing to save — curation happened while we were talking.

Afternoon. Two agents on one project: one on the front end, one on the back end, each with its own key, working from one memory. Neither has to be told what the other decided — the back-end agent's next prompt carries what the front-end conversation settled. Later, the front-end agent suggests an approach for the settings page, and the memory answers for me: that option was rejected in March, here's the day, here's the reason, in the words of that conversation. Nobody re-litigates a dead option because nobody forgot it was dead.

Everything above happens by saying it. I never file anything, never call a save tool, never run a sync. The agent manages the memory because the memory shows up in its context before it needs to be managed.

One honest trade: injection costs context. A brief and per-prompt records occupy tokens — that's what injection is. What you get back is tool calls, turns and hunting. On balance, that trade has been lopsided in our favor for months.

Why this one is different under the hood

Three layers, with different jobs:

  • Curated records are the truth of now. The memory reads your conversations in windows of five exchanges and keeps what the conversation settled — reconciling new facts against old ones, quietly retiring records that describe something you replaced. You never file anything; there's nothing to file.
  • The decision graph is how it got there. Every decision a line: dated, with its reason as it was said, and the option it beat. Reversals are new lines — history is never rewritten, so "why is it like this" always has an answer.
  • The full history is everything, kept in full and searchable by keyword or meaning. Records are what's true; the graph is what happened; the archive is the floor under both.

Two details that matter in practice:

Projects are walled off. Each has its own records, rules, graph and plan; nothing bleeds between them unless you call a cross-project search yourself. On hosted memory, scoped keys let you hand a teammate, a client or a subcontractor access to exactly the projects you choose — read-write, read-only or invisible — and nothing else.

Rules can't rot into a junk drawer. Each one is capped, the pool is capped, you're warned at 80%, and a save that would overfill is refused with the fix. Nothing you wrote is ever quietly trimmed or dropped.

And the privacy line, because it's the reason we built it this way: your memory lives on your machine (a local container, stored there and nowhere else) or in our cloud, encrypted. Our hub processes pieces of it to make memory and compose briefs, and keeps none of those pieces.

Why believe us

Because we run on it. Our own company's work has gone through it every day since December 2025. Before that it was md files, and almost everything from those months is lost — which is the whole argument in one sentence. Every conversation since April is still there, searchable.

And because of how it started. I tried most of the memory systems for AI agents. Each did a part of this, and none did what I needed. So I built the one that does.

Say one sentence

You don't need IT knowledge for this. The install is one sentence to your agent: "install brethof-brain." It fetches the client, wires the hooks, and checks its own work. The only human steps are the ones every application on earth asks for: sign up, choose local or hosted, copy the key, and paste it once into a small window on your own screen, so your agent never sees it. Free to start, and it works with Claude Code, Codex, OpenClaw, Hermes and the other harnesses you're probably already running.

If you've read this far, you've accepted re-explaining as part of the job. It isn't.

Start at brethof.ai/brain/ — and tell your agent: install brethof-brain.

If you end up writing about it yourself, we run an affiliate programme — details at brethof.ai/brain/affiliate/.

Top comments (1)

Some comments may only be visible to logged-in visitors. Sign in to view all comments.