Over the past couple of years I've tried several times to teach an AI agent how to impersonate my voice. While I haven't been super successful, it's been a wonderful learning opportunity for me to dig into RAG, agent memory, and other advanced topics.
I've iterated quite a bit on my original design, which was mostly just stuffing entire blog posts into the system prompt as examples. I built an interesting version of agent memory that had its shortcomings. Then I swapped over to Oracle and designed a hybrid search with three layers of memory, which was objectively better. But it still wasn't perfect.
Actually, the memory part is pretty good. It's the retrieval part that was leaving something on the table. It's taken me a long time to figure out where I was going wrong with loading relevant few-shot examples. Let's see if you can spot the problem:
SELECT id, content, topic, VECTOR_DISTANCE(embedding, :q, COSINE) AS distance
FROM posts
WHERE user_id = :userId AND platform = :platform AND is_deleted = 0
ORDER BY distance
FETCH APPROX FIRST :k ROWS ONLY
Did you find it?
No? Well, that's because there's nothing technically wrong with it. It finds the memories most similar to the question. The problem is that the most similar memory isn't always the right memory.
Memories pile up. After a few years of writing, my store has a post where I argued for one approach and a newer one where I'd changed my mind. It has posts on event-driven architecture, event sourcing, and event streaming that are all very similar to each other in the embedding space. It has a post on serverless pricing and one on cold starts that use the same vocabulary but make completely different points. As far as similarity scores go, these groups are basically the same thing. It just comes down to a date, a specific term, or the actual point of the post.
So retrieval worked, but my context was garbage.
The standard fix is a reranker, which helps to improve relevance instead of just similarity.
When you embed a memory, the model has no idea what will be asked of it later. Your query gets embedded and a similarity score is calculated against the embeddings in the store. A cross-encoder (a specific type of reranking model) looks at the query and memory together and scores them by how well a memory answers the question. This means the relevance is higher, and it often leads to better results.
But there's a catch (of course).
A cross-encoder runs at query time, once per candidate, on every request. That means added latency to your query. 🙃
But if the relevance of the results is meaningfully better, that makes up for the increased latency, right? To a point, yes. If the latency is so high the user experience suffers, it's probably not worth the cost. But if it's relatively minor, then absolutely.
So I built a benchmark to evaluate the tradeoff. I started from the evaluation plan in Oracle's production RAG evaluation guide, which compares keyword, vector, and hybrid retrieval against judged queries. Then I pushed it one step further and put a reranker on top of each, to see how much relevance I got for the latency cost. For the reranker, I used the BGE reranker loaded into the database as an ONNX model, so scoring happens directly in the query with no extra service to call. 🔥
Let's dive in. The results are fascinating.
Before I continue, thank you to Oracle for sponsoring this post. Opinions are my own.
Filter first, then rank
RAG is cool because it involves multiple kinds of lookups. You have lexical retrieval, which is great for identifiers, codes, and names. You also have vector retrieval, which finds semantically similar data via embeddings. I benchmarked both on their own and together as a hybrid search, fusing the two ranked lists with reciprocal rank fusion (RRF).
Both paths run the same scope filter:
WHERE TENANT = :tenant
AND (OWNER_ID IS NULL OR OWNER_ID = :owner)
AND (EXPIRES_AT IS NULL OR EXPIRES_AT > SYSDATE)
I share this because the WHERE clause handles eligibility, which is different from relevance. I don't want an expired memory coming back in my results, like a 2025 take on building AI agents that I've since moved past. And I definitely don't want another tenant's data showing up (security nightmare). Oracle's guide calls these filters mandatory for multi-tenant data. They run before ranking, so everything downstream only deals with memories that are allowed to be there.
Setting the baseline
Before I could tell whether a reranker was worth anything, I needed to know the baseline values for these queries. In addition to latency, I also needed to know the normalized discounted cumulative gain (nDCG) and recall. Capturing these metrics lets us objectively measure relevance alongside latency.
The baseline measurements were taken with K=10, meaning I took the top 10 results.
| Retrieval | nDCG@10 | Recall@10 | Median latency |
|---|---|---|---|
| Vector | 0.753 | 0.865 | 20.0 ms |
| Lexical | 0.727 | 0.740 | 4.7 ms |
| Hybrid RRF | 0.832 | 0.885 | 20.5 ms |
nDCG measures how well a ranking puts relevant items near the top, with higher being better. It's not a measure of how many answers were right, though. Vector's 0.753 means that the rankings earned 75% of the score a perfect ordering would have, on average.
Recall ignores order. It tells you whether the correct memories made it into the top 10 at all. A 0.865 means that about 87% of the memories that should have come back did. Both of these metrics are evaluated against a known ground-truth set of data in order to be objective. Meaning I know ahead of time what the results should have been given the corpus and the query I was benchmarking.
You can see from these results that lexical is fast, but it's not super accurate. Vector catches more at the cost of latency, but still ranks them worse than I'd want. Fusing the two rankings had the best results, with minimal additional latency.
From the baseline results alone, it looks like a hybrid search will be the retrieval method to beat if latency doesn't drop off a cliff.
The enlightening reranker results
Now to run the same benchmark with the reranker. In Oracle, it's one more step at the end of the retrieval query. Retrieval narrows everything down to the top candidates, and PREDICTION() runs the cross-encoder against each one:
SELECT id,
PREDICTION(BGE_RERANKER USING
:query || '</s></s> ' || title || '. ' || content AS DATA) AS score
FROM candidates
ORDER BY score DESC
FETCH FIRST 10 ROWS ONLY;
The </s></s> is the delimiter the model expects between a query and a passage. I'm only pulling back the ids and the reranked scores for the benchmark. And you might notice that the reranking is done entirely in the database query - super cool.
| Configuration | nDCG@10 | Recall@10 | Median latency |
|---|---|---|---|
| Hybrid, no reranker | 0.832 | 0.885 | 20.5 ms |
| Hybrid + reranker, 10 candidates | 0.832 | 0.885 | 585 ms |
| Vector + reranker, 20 candidates | 0.830 | 0.896 | 1,258 ms |
Surprisingly, adding the reranker to hybrid search didn't do anything to nDCG at all. The best reranked score came from vector retrieval with 20 candidates, and it still scored lower than plain hybrid with no reranker, at 61 times the latency. It found slightly more of the right memories, but ranked them a little worse.
I also tried giving the reranker more to work with. For vector retrieval, latency grew linearly with candidate count (1,258ms at 20 candidates and 2,604ms at 40), and nothing got meaningfully better.
That doesn't mean the reranker did nothing, though. I dug a little deeper and took the change for each query and resampled them 2,000 times to see how much the average varies depending on which queries you run. 95% of the resampled averages landed between -0.079 and +0.073.
And if we use that range, +0.073 would take hybrid from 0.832 to 0.905, which is a meaningful improvement. But conversely, the other end of the spectrum would drop it to 0.753, which is where plain vector search started.
Given the range, the results are inconclusive, at least with my sample size of 16 queries. The real effect could be a solid improvement, nothing at all, or even a loss, and 16 queries can't tell those apart.
This ended up being a big realization for me. Relevance is measurable, but how confident you can be depends on how many judged queries you have. A couple thousand iterations of 16 queries is still only 16 queries, and that's why the range is so wide.
Meanwhile, the latency hit is pretty obvious. It was nearly 30 times slower for every request. Now I had to ask myself when that latency is worth it.
Where rerankers make sense
Stay with me. Rerankers have lots of value despite what we've discussed so far. The underwhelming results are because I improved retrieval quality in a cheaper way. Let's look again at the baseline compared to the reranked version:
| First stage | No reranker | 10 candidates | 20 candidates | 40 candidates |
|---|---|---|---|---|
| Lexical | 0.727 | 0.800 | 0.801 | 0.809 |
| Vector | 0.753 | 0.821 | 0.830 | 0.814 |
| Hybrid RRF | 0.832 | 0.832 | 0.829 | 0.817 |
Intuitively, the weaker retrieval methods got the most out of reranking. Lexical and vector each gained about 0.075, while hybrid gained nothing.
Which makes sense because a reranker doesn't retrieve anything. It reorders (or re-ranks 😄) what it's given. If your first stage is already pulling the correct memories into the pool in a sensible order, there's nothing left to fix.
Reranking vector retrieval with 20 candidates left me a little curious because it made the top five less alike. I measured it a couple of different ways (with embeddings and with plain word overlap) and verified the numbers. Thinking about it though, that is desired behavior. Remember that memory lookalike issue I talked about? The cross-encoder is separating the right memory from its lookalikes and pushing the lookalikes down. Which is exactly what I said I wanted.
The takeaway I want you to leave with is that you should rerank when you have one retrieval signal, and fuse when you have two. Doing both leaves you paying for additional latency you don't need.
My new approach
I started this wanting to know whether a reranker was worth the extra inference. After running through several rounds of benchmarks and tweaking queries, I'm not sure that was the right question. My approach to the problem is much more informed and now walks a different path than I expected.
When building a RAG pipeline for memory, I'll always filter first. Owner, expiration, whatever makes a memory eligible for you. Apply these in the query before anything gets ranked. Reducing the number of candidates means fewer memories to rank (plus, you know, security and quality and all that 😜).
Next, I'd add a second retrieval signal before adding a reranking model. If you have vector search, add lexical. If you have lexical, add vector. Fuse the rankings with RRF. Doing that with my tests took nDCG@10 from 0.75 with vector retrieval to 0.83 and only slowed my query down by half a millisecond.
After adding the second retrieval signal, I'd measure everything. Write down 15 to 20 real requests you'd give your agent, making sure to include some difficult ones and edge cases. Things like asking for exact keyword lookups, or for advice on something you've changed your mind on a couple of times. Keep those handy. The guide I was following says to rerun them every time you change chunking, the embedding model, or the reranker, and it's right.
Then you can try a reranker against your baseline. Try a few different candidate counts (:n in the query above). Check whether the quality difference is consistently higher across your queries and decide if the additional latency is worth the quality difference.
Luckily for me, what I built in my original post doesn't need to change much. I still have the same tables and the same reflection loop. I'm really just updating the database query that decides which memories the agent sees.
Try it yourself
The benchmark and instructions to try it all yourself are available in GitHub. It spins up Oracle AI Database Free in Docker, generates a synthetic corpus of 480 memories and 16 judged queries, loads the ONNX models, and runs the whole thing.
Everything we've discussed today ran on Free in Docker, capped at two CPUs and 2 GB. It's an intentionally small box, because we aren't running a capacity benchmark and none of these numbers should be used for sizing anything.
If you want to run this kind of evaluation on your own retrieval, start with Oracle's production RAG evaluation guide. It covers more than what we talked about here, like grading the generated answer and testing freshness, tenant isolation, and when the agent should refuse to answer. It's pretty solid.
Agent memory needs more than vector search. Turns out, just using vector search means you're asking a single similarity score to decide what's relevant, current, and yours. Fortunately for all of us, "more" is a lot cheaper than I assumed.
Happy coding!
Top comments (1)
The "most similar memory isn't always the right memory" line is the whole game. The example that stuck with me is the one you gave: an old post arguing X and a newer one where you'd changed your mind. A cross-encoder scores relevance to the query, but both of those are highly relevant — what it can't tell you is which one is current. Relevance and currency are different axes, and rerankers only fix the first one.
Two things that helped us before reaching for the cross-encoder tax on every request: (1) a cheap recency/decay prior baked into the candidate scoring, so newer memories win ties for free, and (2) explicit supersession edges — when a memory contradicts an earlier one, link them, so retrieval can prune the stale one instead of feeding the model both and hoping. Overfetch → rerank only the top ~20-30 also keeps the latency you flagged bounded. Have you experimented with letting the memories themselves carry a "supersedes" pointer rather than resolving conflicts purely at query time?