I maintain Agent Brain Hub, an open-source memory layer that several AI agents share. Last release I added real embedding models, and recall on my benchmark went from 88.5% to 100%. Then I wrote a harder test, and the model fell over on something any keyword search gets right.
The failure
A customer has twelve shopping conversations, one per order:
My order ORD-48200 for a blue backpack hasn't arrived yet
My order ORD-48201 for running shoes hasn't arrived yet
…
My order ORD-48211 for a water bottle hasn't arrived yet
Then another agent asks the memory: "what is going on with order 48207".
| Retrieval | Finds ORD-48207? |
|---|---|
| Local feature hashing (lexical) | ✅ |
bge-m3 embeddings |
❌ |
To the embedding model all twelve conversations mean the same thing ("a late order"), so they scored almost identically, and recency decided which four made it into the context. The one distinguishing token, 48207, is exactly what dense embeddings blur. Hashing vectors are crude, but they match exact tokens.
What didn't work: taking the max
My first fix took whichever signal was stronger: max(model, lexical). It still failed. The model's score for every order (~0.6) was above the lexical score of the right one, so the max was the model score everywhere and nothing changed.
What worked: adding them
// Hybrid (CombSUM): meaning + wording
const lexical = Math.max(0, cosine(qHash, item.hashVec));
return modelVecAvailable ? Math.max(0, dot(qVec, item.vec)) + lexical : lexical;
An item that matches in meaning and wording now beats items that are merely on topic. The sum is never below either signal, so paraphrase matches ("Can I drive to the airport?" → "no car for 3 days") keep the score that made them work. I added an exact_match category to the benchmark, with order numbers that differ by one digit, with and without the ORD- prefix, plus device model names. A unit test proves the fix: it fails if you revert it.
With bge-m3 on all 25 scenarios: 96% pass, 100% recall, 0% leaks.
Then I asked a public dataset
My own scenarios only compare versions of my own project. So v0.4.0 also runs LongMemEval-S (MIT): 470 questions, each hidden among ~50 past chat sessions. Is the session that holds the answer near the top?
| Local hashing, no network | top 1 | top 4 | top 10 |
|---|---|---|---|
| All questions | 30.0% | 52.8% | 71.3% |
| "The assistant told me…" | 66.1% | 87.5% | 92.9% |
| Preferences | 3.3% | 20.0% | 43.3% |
Honest takeaway: about half the time the right session is in the top 4, and preference questions are bad. That number is the reason to have a public benchmark: it tells me where to work next instead of telling me I'm done. (The embedding-model run needs a GPU; on CPU it would take days, so it's next.)
And it got fast
Retrieval used to scan every customer's memories, then filter. A per-customer index changed the tail:
| 100,000 episodes | before | after |
|---|---|---|
| recall p50 | 4.3 ms | 0.6 ms |
| recall p95 | 40.1 ms | 1.9 ms |
Try it
git clone https://github.com/leluong141996-dev/Agent-Brain-Hub && cd Agent-Brain-Hub
docker compose up -d # http://localhost:4317
npm run bench -- --url http://localhost:4317 # benchmark your own running hub
If your agents' memory ever failed in a way like this (a code, a name, a date the retriever couldn't see), I'd love it as a benchmark scenario. It's one JSON file, and every future release has to pass it.
Top comments (1)
Some comments may only be visible to logged-in visitors. Sign in to view all comments. Some comments have been hidden by the post's author - find out more