DEV Community

Cover image for Three releases of my AI memory layer. The best bugs were found by other people.
Lượng Lê
Lượng Lê

Posted on

Three releases of my AI memory layer. The best bugs were found by other people.

I'm building Agent Brain Hub, an open-source shared memory layer for multi-agent systems.

Imagine a travel agent, a repair agent, and a finance agent all talking to the same user. Each should know what the others learned, without leaking sensitive data (like health records or income) between them. The architecture borrows loosely from cognitive psychology: a short-term working memory per session, a long-term store for facts and episodes, and an asynchronous "sleep cycle" that consolidates one into the other.

Over the last few days, I pushed three releases. Here's what broke, what the benchmarks showed, and why the best edge cases didn't come from my test suite. They came from people reading my posts.

v0.5: Point-in-time correctness (time-traveling memory)

Until v0.5, a replaced or expired fact was deleted at the next sleep cycle. Now facts are invalidated with validity intervals: "I moved to Saigon" closes "I live in Hanoi" with a timestamp and a pointer to the superseding record.

You can query the brain's state at any point in history:

await brain.recall({
  customerId,
  text: 'Book a taxi from my home',
  asOf: '2026-10-01T09:00:00Z',
});
// Returns the address known at that timestamp, not today's address.
Enter fullscreen mode Exit fullscreen mode

When debugging multi-agent loops, this is exactly the question you need to answer: "What did the travel agent actually believe when it made that call?"

The bug the benchmark caught: Adding temporal queries exposed a privacy leak before release. Querying past states skipped the privacy filter on recent conversation turns, so one agent's private conversation could reach another agent. The new temporal benchmark scenarios caught it before it ever shipped.


v0.6: An entity graph, and an honest benchmark

Facts are now grouped into per-customer entities ("my car" = "xe" = "車" = "Honda Civic"), and retrieval follows the links between them:

Mia (personal agent):  "My car is a Honda Civic"
Kai (repair agent):    "The car broke down, it'll be in the shop for 3 days"

Penny (finance agent): "Should I set money aside for my Civic this week?"
→ availability: no car for 3 days (from Kai, 2 hours ago, expires in 4 days, via: Honda Civic)
Enter fullscreen mode Exit fullscreen mode

Notice that Penny's question never mentions "car" or "shop". The graph links "Civic" to the car, and the car to its repair. Traversal only runs over memories the asking agent is already allowed to see, so it can't carry a private fact across domains (a medical allergy never reaches the booking agent).

The Graph tab: a trip linked to the car in the shop, the budget, the home city, and a private allergy only health agents can follow

I ran the benchmark with --no-graph to see if the graph actually pulls its weight:

Offline mode (hashing vectors) With graph Without graph
Multi-hop questions 100% 40%
All 37 scenarios 97.3% 81.1%
Leak rate 0% 0%
Mean prompt tokens 427 421

The part most benchmark posts leave out: if you switch from offline hashing vectors to an embedding model like bge-m3, the multi-hop scenarios pass even without the graph. The model already associates "the XPS" with "the laptop".

Right now, the graph earns its keep in:

  1. Zero-config offline mode, with no embedding model.
  2. Explainability: a deterministic trace of why a fact was pulled in.

My current scenarios aren't hard enough to show more than that against dense embeddings. Harder, adversarial scenarios are now in the backlog.


The edge cases I completely missed

I posted an earlier devlog on daily.dev and opened the repo to contributions. Three real insights came straight from readers.

1. Stepping on a contributor's toes

A contributor claimed an open issue and raised a PR. A few hours later, unaware of it, I shipped my own fix for the same issue.

Their PR turned out to catch a case mine missed: in Vietnamese, "không thích … nữa" ("don't like … anymore") was still treated as a conflict instead of an update to the old preference. I rebased their PR onto main, kept their authorship, added a small follow-up, merged it and tagged v0.6.1, the project's first outside contribution.

Takeaway: if your repo is public, check claimed issues and open PRs before you open your IDE.

2. "Headroom rather than a guarantee" (the race condition)

Working memory holds the last 40 turns. Once 24 unconsolidated turns pile up, the scheduler triggers a consolidation "sleep cycle".

A commenter, Ahmet Özel, pointed out that the 16 turns between 24 and 40 are "headroom rather than a guarantee": a slow summarizer or a queue backlog can still let turns fall off the end.

Following their suggestion, I wrote a burst test with a deliberately slow summarizer. Result: 10 out of 26 turns vanished silently. Turns arriving during the summarization call weren't part of the current batch, and the buffer trim after consolidation dropped them anyway.

Fixed in v0.6.2:

  • A turn is only evicted once it's inside a stored episode.
  • Hitting the 40-turn cap triggers a consolidation immediately instead of evicting turns.
  • The burst test is now part of the regression suite.

3. "What happens to a screenshot?"

Another reader asked how consolidation handles a turn that carries a screenshot.

Honest answer: it doesn't. The hub is text-only. And sending images to a remote multimodal model just to summarize them would expose private data and rack up API costs. Their suggestion, a local OCR pass before the consolidation loop, is now tracked in issue #25.


Key takeaways

Synthetic benchmarks catch the edge cases you anticipated when writing the code. Readers catch the assumptions you took for granted. Neither helps if you only share the numbers that make your project look good.

What's next for v0.7

  • Open-schema extraction: moving past the 19 hardcoded relation types.
  • Conflict resolution between agents: rules for when agent A and agent B assert contradictory facts about the same entity.

Question for anyone running multi-agent stacks:
How do you resolve conflicts when two independent agents write contradictory facts about the same entity? Last write wins? Confidence scores? CRDTs? Escalating to a human? I'd love to hear how you handle it.

Agent Brain Hub is MIT-licensed, has no telemetry, and runs fully local with docker compose up, or with any LLM provider (OpenAI-compatible, Anthropic, Gemini, Ollama…): github.com/leluong141996-dev/Agent-Brain-Hub.

Top comments (0)

Some comments may only be visible to logged-in visitors. Sign in to view all comments.