Every AI agent I deployed had the same problem: it started each conversation from zero. And as soon as there was more than one agent (one for support, one for bookings, one for billing), each of them learned things about the same customer and kept it to itself.
So I built Agent Brain Hub: one memory that many agents share. This post is about the design, the brain metaphor behind it, and the mistakes I made along the way.
Repo: https://github.com/leluong141996-dev/Agent-Brain-Hub (MIT, docker compose up)
The goal in one example
- The customer tells Kai, the repair agent: "My car is a Honda Civic, it broke down and will be in the shop for 3 days."
- Later they ask Atlas, the travel agent: "I need a flight to Da Nang next week, window seat please."
- Atlas answers: "I've got the context from Kai. I remember you have no car for 3 days, no need to repeat it. I'll look for a flight to Da Nang, window seat preferred. Also: a self-drive rental at the destination, would you like that?"
Atlas never talked to the customer about the car. The fact "car unavailable for 3 days" was written by Kai into shared memory, while the repair details stayed private to Kai.
Why a brain?
I needed a way to split a memory system into parts with clear jobs. The brain turned out to be a good map, and it made the live visualization explain itself. This is a metaphor, not neuroscience.
| Region | Job in the system |
|---|---|
| Thalamus | Routes the input: language, intent, which agent |
| Brainstem | Redacts PII, catches crisis messages before anything is stored |
| Amygdala | Salience: urgency and emotion raise priority |
| Prefrontal cortex | Working memory for the current conversation |
| RAS | Retrieval: scope, then TTL, hybrid scoring, re-rank, token budget |
| Neocortex | Long-term facts and episodes, hot/warm/cold tiers, contradiction resolution |
| Hippocampus | Encodes new memories from each turn |
| Cerebellum | Skills: a playbook that succeeds 3 times in 7 days is promoted, versioned, can be rolled back |
| Basal ganglia | Picks the next best action with Thompson sampling |
| Corpus callosum | Handoff between agents |
| Forgetting, DMN | The sleep loop: decay, consolidation, reflection |
Every request is traced through these regions over Server-Sent Events, so in the UI you can watch which region did what and which memories the answer used.
Three ways an agent plugs in
- Native: create it in the UI and the hub answers with its own LLM.
- REST or SDK: your agent keeps its own LLM and calls the brain around each turn:
const ctx = await brain.recall({ customerId, text, lang: 'en' });
const reply = await myLLM({ system: mySystemPrompt + ctx.promptBlock, user: text });
await brain.remember({ traceId: ctx.traceId, reply, lang: 'en' });
-
MCP:
brain_recall,brain_remember,brain_profileandbrain_feedbacktools for Claude Desktop, Claude Code, Cursor and agent frameworks.
What worked
Scopes and an audit log, from day one. Every memory is private, shared or global, and every agent has read and write permissions. Every read, write and blocked access goes to an append-only log. Without this, "shared memory" is just a leak with extra steps.
One writer per kind of fact. Only the repair domain writes asset_issue, only travel writes trip_destination. This removed a whole class of agents overwriting each other.
Retrieval as a pipeline you can see. Breaking recall into scope, TTL, scoring, re-rank and budget stages, and showing each stage in the trace, made ranking bugs obvious.
Skills that earn promotion. A playbook only becomes a reusable skill after it succeeds 3 times within 7 days, and every version is kept. Bad promotions can be rolled back.
What I got wrong
Private data leaked through side doors. Retrieval respected scopes, but private details still reached other agents through the recent conversation turns, the handoff summary and the sleep-time reflection. Every path that builds a prompt needs the same visibility filter, not just the main retrieval.
Keyword matching across languages is a trap. Vietnamese "cửa sổ" (window) matched the keyword "sợ" (afraid) once accents were stripped, so a seat preference looked like anxiety. Matching now keeps diacritics when the input has them and uses word boundaries; Japanese uses substring matching because it has no spaces.
The bandit was too eager. Thompson sampling sometimes let an irrelevant action win on a lucky draw. Narrowing the exploration range around relevance fixed it.
Rules and the LLM contradicted each other. When both extracted a fact from the same message, single-valued facts could end up with two values. Now they're merged and deduplicated before writing.
Honest limitations
- Embeddings are local feature hashing, so the hub runs with no external API. A real embedding model would be better.
- Fact extraction is rules plus an optional LLM pass.
- It's a single process on SQLite. PostgreSQL with pgvector is the next step for scale.
Try it
git clone https://github.com/leluong141996-dev/Agent-Brain-Hub
cd Agent-Brain-Hub
docker compose up -d # → http://localhost:4317
It runs offline with no configuration. Pick an LLM later in Settings: Claude, GPT, Gemini, DeepSeek, or local models through Ollama or vLLM. The UI speaks English, Vietnamese and Japanese.
If you run several agents in production, I'd love to hear what you need from shared memory. Issues labelled good first issue are open, and a ⭐ on the repo helps other people find it.

Top comments (0)