The assistant sat on top of about 9,000 internal documents: runbooks, design docs, meeting notes, and a wiki that had been migrated twice. You asked a question, it found relevant documents, put them in front of a model, and the model answered. Standard retrieval augmented generation. It had been built in a fortnight the year before, and by the team's own measurement it was right about two thirds of the time.
When I got involved, the proposal on the table was to switch to a bigger model. I asked for one thing first: for fifty questions where we knew the answer was wrong, show me which documents were retrieved. In 38 of the 50, the document with the correct answer wasn't in the retrieved set at all. The model had answered from documents that didn't contain the answer, and done it confidently, because that's what a model does when you give it context and ask for an answer.
So the problem was the search. Nobody had looked at it because RAG has a name that makes it sound like an AI technique, when most of it is the same search engineering we've been doing since 2005.
Measure the retrieval on its own
Before touching anything, stop measuring the system end to end and measure the retrieval step by itself. For that you need a set of questions, each with a known correct document, and a number for how often that document lands in the top k results.
We built the set from the 50 failures plus 150 questions the support team had answered by hand, along with the document they'd used for each answer. That gave us two hundred pairs. The metric was recall at 5, because the assistant put five documents in front of the model.
Recall at 5 on the original system was 61 percent, and that number explains everything. Two thirds of the time the correct document was in the five. A third of the time it wasn't, and the model was guessing from the wrong material. No matter which model sat on top, the system could never be more accurate than its retrieval recall.
Once we had the number, everything after it was routine. Change something, run the 200 questions, read the number.
Chunking was the first problem
The documents were split into 512 token chunks with no overlap, because that was the number in the tutorial the original builder had followed. Runbooks written as numbered steps got cut in the middle of a step. A design doc's "Decision" section, the paragraph everyone actually wanted, was often split across two chunks, and neither had the whole decision.
We switched to chunking on document structure, headings and paragraphs, with a maximum size instead of a fixed one, and gave each chunk its document title and the heading path above it as a prefix. That took recall at 5 from 61 to 70: nine points just from not cutting sentences in half.
The heading prefix matters more than it sounds. A chunk that says "Rotate the key in the vault, then restart the workers" is ambiguous. One that says "Payments service runbook / Key rotation / Rotate the key in the vault, then restart the workers" matches the query "how do I rotate the payments key" on the words that matter.
Embeddings alone lose on exact terms
The original system was pure vector search: embed the query, find the nearest chunks by cosine similarity, done. Vector search is good at meaning and bad at names. A query for PAYMENTS_WEBHOOK_SECRET finds chunks about webhook configuration in general, because the embedding of one environment variable name sits close to the embedding of any other. And internal documentation is full of specific names: services, variables, error codes, ticket numbers, people.
Adding a keyword index next to the vector index and combining the two paid off more than any other change in the project. That meant Postgres with pgvector for the embeddings and its built in full text search for keywords, combined with reciprocal rank fusion:
WITH vec AS (
SELECT id, row_number() OVER (ORDER BY embedding <=> $1) AS r
FROM chunks ORDER BY embedding <=> $1 LIMIT 40
),
kw AS (
SELECT id, row_number() OVER (ORDER BY ts_rank_cd(tsv, q) DESC) AS r
FROM chunks, plainto_tsquery('english', $2) q
WHERE tsv @@ q LIMIT 40
)
SELECT id, sum(1.0 / (60 + r)) AS score
FROM (SELECT * FROM vec UNION ALL SELECT * FROM kw) u
GROUP BY id
ORDER BY score DESC
LIMIT 20;
Recall at 5 went from 70 to 79. Queries with a specific name in them went from about 50 percent to about 90, and queries without one barely moved, which is what you'd expect from adding a keyword signal to a semantic one.
In short, the two signals fail in different ways and combining them costs almost nothing, and I still meet teams running vector only because that's what the diagram in the tutorial showed.1
Rerank the top twenty
After fusion the correct document was usually in the top twenty and often not in the top five. The fix for that is a reranker, a model that takes the query and each candidate chunk together and scores how well the chunk answers the query. Per pair it costs far more than comparing embeddings, which is why you run it on twenty candidates and not on nine thousand chunks.
Recall at 5 went from 79 to 88, the biggest single jump after hybrid search, for one extra call per question.2
You can also use the main model as the reranker and ask it to pick the five most relevant out of twenty. It works and it's slower, and on our set the dedicated reranker did better. Try both if you have a set to try them on. Without one you're guessing, which is where this project started.
The documents were the last problem
At 88 percent, most of the remaining failures weren't retrieval failures. The correct document didn't exist, or existed three times with conflicting content, or was a 2023 meeting note describing a process that had since changed.
Search and AI can't fix that. It's a documentation problem, and the assistant had been hiding it by answering confidently from whatever it found. The fix was organisational: an owner for every runbook, a "last verified" date on every page, and a rule that the assistant only retrieves from pages verified in the last year unless nothing verified matches, in which case it says so. That last rule cut the confident wrong answers more than anything technical did, because "I found a document from 2023 that may be out of date" beats a fluent paraphrase of stale instructions.
Recall at 5 ended at 91 percent on the set. End to end accuracy, measured by the support team on 100 fresh questions, went from 66 percent to 89. The model was the same one we started with.
Query rewriting, the piece I left for last
One more technique came after those four. I've left it for last because it's the one that looks most like an AI technique, and it should be the last thing you reach for.
Users don't write search queries. They write questions, or fragments, or whatever they remember from the last time they looked. "that page about the vault rotation thing from the payments outage" is a real query from the logs. Keyword search finds nothing useful in it because half the words are filler. Vector search finds pages about vaults and pages about outages and has no idea "rotation" is the word that matters.
Query rewriting puts a model in front of the search. Given the user's text, it produces two or three search queries that would find the answer. For that query it came up with "payments service key rotation runbook", "vault key rotation procedure" and "payments outage post-mortem key rotation", and each of those is a good query. Run all three through the hybrid search, fuse the results, rerank, and the right runbook came out first.
On the eval set this took recall at 5 from 88 to 91, the last three points I quoted, and it cost a model call before every search, about 400 milliseconds. That's why it's last. Chunking, hybrid search and the reranker each gave more for less, and yet query rewriting is what people build first, because it's the one with a prompt in it.
There's a cheaper version worth trying first: expand the query with synonyms from your own domain, in code, from a table.3 That table has 40 entries and was worth two points on its own.
Keeping it honest after launch
Retrieval quality decays. Documents get added, the questions people ask shift, a new product launches and the corpus has nothing on it. A system at 91 percent in March won't be at 91 percent in September unless someone is measuring.
Three things keep it measured. First, the eval set grows from production. Every answer a user marks as wrong, and every question a support engineer answers by hand because the assistant couldn't, becomes a case with the correct document attached.4
Second, the retrieval metric runs nightly against the full set and posts the number to a channel, and a drop of more than two points opens a ticket. That's happened four times in six months. Two were new document types with a structure the chunker didn't handle. One was the provider changing the embedding model version, which shifted every vector slightly and dropped recall three points until we re-embedded the corpus. One was a runbook that had been split into three pages with the same title, which the reranker couldn't tell apart.
Third, the assistant shows its sources. Every answer lists the documents it used, with links. That's partly so the user can check. Mostly it's for the team, because a wrong answer with visible sources tells you straight away whether retrieval or generation was at fault, and in six months of looking it's been retrieval every time but two.
What I would do on day one
Build the question set first. Fifty questions with known correct documents is enough to start, and it grows with every wrong answer. Without it you can't tell whether a change helped, and every argument about which model or which embedding is guesswork.
Measure retrieval recall on its own, separately from the answer.
Chunk on structure, and prefix chunks with their heading path.
Run keyword and vector search together and fuse them. Postgres does both.
Rerank the top twenty.
Only after that, look at the model. In our case there was nothing there to fix.
The name "RAG" makes the generation sound like the interesting part, but generation is the easy bit, and a model given the right document answers well. All the difficulty is in "retrieval", and people who've built search engines for twenty years already know how to do it. Borrow their methods. They aren't new or glamorous, and they're where the 25 points came from.
Originally published at zeybek.dev.
-
I've written about hybrid search in Postgres in more detail before. ↩
-
We used a small hosted cross-encoder reranker, at about 80 milliseconds for twenty pairs. ↩
-
In our docs "vault" also means "secrets manager", because the tool was renamed in 2024. ↩
-
The set is over 600 cases now, and the newest are the most valuable, because they're what people are asking now. ↩
Top comments (0)