DEV Community

Cover image for I modeled AI agent memory on the human brain. Here's what worked and what didn't
Lượng Lê
Lượng Lê

Posted on

I modeled AI agent memory on the human brain. Here's what worked and what didn't

Agent Brain Hub live brain view

Every AI agent I deployed had the same problem: it started each conversation from zero. And as soon as there was more than one agent (one for support, one for bookings, one for billing), each of them learned things about the same customer and kept it to itself.

So I built Agent Brain Hub: one memory that many agents share. This post is about the design, the brain metaphor behind it, and the mistakes I made along the way.

Repo: https://github.com/leluong141996-dev/Agent-Brain-Hub (MIT, docker compose up)

The goal in one example

  1. The customer tells Kai, the repair agent: "My car is a Honda Civic, it broke down and will be in the shop for 3 days."
  2. Later they ask Atlas, the travel agent: "I need a flight to Da Nang next week, window seat please."
  3. Atlas answers: "I've got the context from Kai. I remember you have no car for 3 days, no need to repeat it. I'll look for a flight to Da Nang, window seat preferred. Also: a self-drive rental at the destination, would you like that?"

Atlas never talked to the customer about the car. The fact "car unavailable for 3 days" was written by Kai into shared memory, while the repair details stayed private to Kai.

Why a brain?

I needed a way to split a memory system into parts with clear jobs. The brain turned out to be a good map, and it made the live visualization explain itself. This is a metaphor, not neuroscience.

Region Job in the system
Thalamus Routes the input: language, intent, which agent
Brainstem Redacts PII, catches crisis messages before anything is stored
Amygdala Salience: urgency and emotion raise priority
Prefrontal cortex Working memory for the current conversation
RAS Retrieval: scope, then TTL, hybrid scoring, re-rank, token budget
Neocortex Long-term facts and episodes, hot/warm/cold tiers, contradiction resolution
Hippocampus Encodes new memories from each turn
Cerebellum Skills: a playbook that succeeds 3 times in 7 days is promoted, versioned, can be rolled back
Basal ganglia Picks the next best action with Thompson sampling
Corpus callosum Handoff between agents
Forgetting, DMN The sleep loop: decay, consolidation, reflection

Every request is traced through these regions over Server-Sent Events, so in the UI you can watch which region did what and which memories the answer used.

Three ways an agent plugs in

  • Native: create it in the UI and the hub answers with its own LLM.
  • REST or SDK: your agent keeps its own LLM and calls the brain around each turn:
const ctx = await brain.recall({ customerId, text, lang: 'en' });
const reply = await myLLM({ system: mySystemPrompt + ctx.promptBlock, user: text });
await brain.remember({ traceId: ctx.traceId, reply, lang: 'en' });
Enter fullscreen mode Exit fullscreen mode
  • MCP: brain_recall, brain_remember, brain_profile and brain_feedback tools for Claude Desktop, Claude Code, Cursor and agent frameworks.

What worked

Scopes and an audit log, from day one. Every memory is private, shared or global, and every agent has read and write permissions. Every read, write and blocked access goes to an append-only log. Without this, "shared memory" is just a leak with extra steps.

One writer per kind of fact. Only the repair domain writes asset_issue, only travel writes trip_destination. This removed a whole class of agents overwriting each other.

Retrieval as a pipeline you can see. Breaking recall into scope, TTL, scoring, re-rank and budget stages, and showing each stage in the trace, made ranking bugs obvious.

Skills that earn promotion. A playbook only becomes a reusable skill after it succeeds 3 times within 7 days, and every version is kept. Bad promotions can be rolled back.

What I got wrong

Private data leaked through side doors. Retrieval respected scopes, but private details still reached other agents through the recent conversation turns, the handoff summary and the sleep-time reflection. Every path that builds a prompt needs the same visibility filter, not just the main retrieval.

Keyword matching across languages is a trap. Vietnamese "cửa sổ" (window) matched the keyword "sợ" (afraid) once accents were stripped, so a seat preference looked like anxiety. Matching now keeps diacritics when the input has them and uses word boundaries; Japanese uses substring matching because it has no spaces.

The bandit was too eager. Thompson sampling sometimes let an irrelevant action win on a lucky draw. Narrowing the exploration range around relevance fixed it.

Rules and the LLM contradicted each other. When both extracted a fact from the same message, single-valued facts could end up with two values. Now they're merged and deduplicated before writing.

Honest limitations

  • Embeddings are local feature hashing, so the hub runs with no external API. A real embedding model would be better.
  • Fact extraction is rules plus an optional LLM pass.
  • It's a single process on SQLite. PostgreSQL with pgvector is the next step for scale.

Try it

git clone https://github.com/leluong141996-dev/Agent-Brain-Hub
cd Agent-Brain-Hub
docker compose up -d   # → http://localhost:4317
Enter fullscreen mode Exit fullscreen mode

It runs offline with no configuration. Pick an LLM later in Settings: Claude, GPT, Gemini, DeepSeek, or local models through Ollama or vLLM. The UI speaks English, Vietnamese and Japanese.

If you run several agents in production, I'd love to hear what you need from shared memory. Issues labelled good first issue are open, and a ⭐ on the repo helps other people find it.

Top comments (0)