DEV Community

Lượng Lê
Lượng Lê

Posted on

My embedding model couldn't find "order 48207". Here's the one-line fix and what LongMemEval said next.

I maintain Agent Brain Hub, an open-source memory layer that several AI agents share. Last release I added real embedding models, and recall on my benchmark went from 88.5% to 100%. Then I wrote a harder test, and the model fell over on something any keyword search gets right.

The failure

A customer has twelve shopping conversations, one per order:

My order ORD-48200 for a blue backpack hasn't arrived yet
My order ORD-48201 for running shoes hasn't arrived yet
…
My order ORD-48211 for a water bottle hasn't arrived yet
Enter fullscreen mode Exit fullscreen mode

Then another agent asks the memory: "what is going on with order 48207".

Retrieval Finds ORD-48207?
Local feature hashing (lexical) ✅
bge-m3 embeddings ❌

To the embedding model all twelve conversations mean the same thing ("a late order"), so they scored almost identically, and recency decided which four made it into the context. The one distinguishing token, 48207, is exactly what dense embeddings blur. Hashing vectors are crude, but they match exact tokens.

What didn't work: taking the max

My first fix took whichever signal was stronger: max(model, lexical). It still failed. The model's score for every order (~0.6) was above the lexical score of the right one, so the max was the model score everywhere and nothing changed.

What worked: adding them

// Hybrid (CombSUM): meaning + wording
const lexical = Math.max(0, cosine(qHash, item.hashVec));
return modelVecAvailable ? Math.max(0, dot(qVec, item.vec)) + lexical : lexical;
Enter fullscreen mode Exit fullscreen mode

An item that matches in meaning and wording now beats items that are merely on topic. The sum is never below either signal, so paraphrase matches ("Can I drive to the airport?" → "no car for 3 days") keep the score that made them work. I added an exact_match category to the benchmark, with order numbers that differ by one digit, with and without the ORD- prefix, plus device model names. A unit test proves the fix: it fails if you revert it.

With bge-m3 on all 25 scenarios: 96% pass, 100% recall, 0% leaks.

Then I asked a public dataset

My own scenarios only compare versions of my own project. So v0.4.0 also runs LongMemEval-S (MIT): 470 questions, each hidden among ~50 past chat sessions. Is the session that holds the answer near the top?

Local hashing, no network top 1 top 4 top 10
All questions 30.0% 52.8% 71.3%
"The assistant told me…" 66.1% 87.5% 92.9%
Preferences 3.3% 20.0% 43.3%

Honest takeaway: about half the time the right session is in the top 4, and preference questions are bad. That number is the reason to have a public benchmark: it tells me where to work next instead of telling me I'm done. (The embedding-model run needs a GPU; on CPU it would take days, so it's next.)

And it got fast

Retrieval used to scan every customer's memories, then filter. A per-customer index changed the tail:

100,000 episodes before after
recall p50 4.3 ms 0.6 ms
recall p95 40.1 ms 1.9 ms

Try it

git clone https://github.com/leluong141996-dev/Agent-Brain-Hub && cd Agent-Brain-Hub
docker compose up -d                                   # http://localhost:4317
npm run bench -- --url http://localhost:4317           # benchmark your own running hub
Enter fullscreen mode Exit fullscreen mode

If your agents' memory ever failed in a way like this (a code, a name, a date the retriever couldn't see), I'd love it as a benchmark scenario. It's one JSON file, and every future release has to pass it.

👉 https://github.com/leluong141996-dev/Agent-Brain-Hub

Top comments (1)

Some comments may only be visible to logged-in visitors. Sign in to view all comments. Some comments have been hidden by the post's author - find out more