You built a RAG app, asked it five questions in the demo, and every answer sounded right. That is not evaluation, that is vibes. The moment a real user asks something slightly off the beaten path, retrieval quietly returns the wrong chunks, the model answers confidently anyway, and nobody notices until a customer does.
RAG systems fail silently. A bad retrieval does not throw an error, it just hands the model the wrong context, and the model, doing exactly what it is built to do, writes a fluent, convincing, wrong answer on top of it. Without a real evaluation process, you find out about these failures from users, not from tests.
Why "it sounded right" is not evaluation
Reading five outputs and nodding is not measurement, it is spot-checking, and spot-checking only catches the failures that happen to land in your five questions. Retrieval quality is not a single number you eyeball, it is a set of properties that need to be checked against a real, repeatable set of queries with known correct answers.
The two systems working together, retriever and generator, fail differently and need to be evaluated separately before you evaluate them together. A generator can write a great answer from bad context. A retriever can find the perfect chunk and still get a bad answer if the generator ignores it. Blending both into one end-to-end "did it sound good" check hides which half is actually broken.
Build an eval set before you build more features
The single highest leverage thing you can do is create a small, real set of query, expected-answer, expected-source triples. Something like:
{
"query": "What is the default max_connections value in Postgres?",
"expected_chunk_ids": ["postgres-pooling-03"],
"expected_answer_contains": ["100"]
}
Twenty to fifty of these, pulled from real questions your users actually ask or would ask, beats a thousand synthetic ones nobody will trust. Write them once, run them on every change to chunking, embedding model, or prompt, and you have a real signal instead of a feeling.
Retrieval metrics: is the right chunk even in the list
Before touching generation quality, check whether the retriever is finding the right source material at all.
Recall@k asks: of the chunks you actually needed, how many showed up in the top k results. If the right chunk is buried at position 15 and you only pass the top 5 to the model, recall@5 is 0, and no amount of prompt engineering fixes that. This is usually the first thing to check when answers go wrong, because a generation problem downstream of a retrieval miss looks identical to a bad prompt.
Precision@k asks the opposite: of what you retrieved, how much was actually relevant. High recall with low precision means you are burying the right chunk in a pile of noise, which still hurts the generator, since irrelevant context competes for the model's attention and can get woven into a wrong answer.
MRR (Mean Reciprocal Rank) captures how high up the correct result lands, not just whether it appears. A system that ranks the right chunk first every time scores very differently from one that eventually finds it at position 8, even if raw recall looks similar.
A simple way to compute recall@k against your eval set:
def recall_at_k(retrieved_ids, expected_ids, k):
top_k = set(retrieved_ids[:k])
hits = top_k & set(expected_ids)
return len(hits) / len(expected_ids)
Run this across your eval set every time you change chunking strategy, embedding model, or index settings, and you get a real before/after number instead of a guess.
Generation metrics: did the model actually use what it was given
Once retrieval is solid, check whether the generator is grounded in what it received, not hallucinating on top of it.
Faithfulness checks whether every claim in the answer is actually supported by the retrieved context. An answer can be well written and completely unsupported by the sources it was supposedly built from, this is the specific failure mode that makes RAG dangerous in production, since it looks exactly like a correct answer.
Answer relevance checks whether the response actually addresses the question asked, independent of whether it's grounded, since a model can retrieve the right chunk and still wander off and answer a related but different question.
A common, practical pattern is using a second LLM call as a judge, asking it to check faithfulness against the retrieved context directly:
judge_prompt = f"""
Context: {retrieved_context}
Answer: {generated_answer}
Does the answer contain any claim not supported by the context?
Respond with only: SUPPORTED or UNSUPPORTED
"""
This is not perfect, an LLM judge has its own blind spots, but it is far better than nothing and it is cheap enough to run against your whole eval set on every change.
Where teams actually get this wrong
Testing only the happy path. Your eval set should include queries with no good answer in the corpus, ambiguous questions, and queries phrased nothing like your documents. A retriever that only works when the question echoes the document's wording will fail constantly in production.
Never re-running the eval set. Building it once and never running it again defeats the purpose. Every change to chunking, embedding model, or prompt needs to run against the same eval set, otherwise you have no idea if you just made things better or worse.
Optimizing retrieval and generation together, blindly. If an answer is wrong, check retrieval first. Fixing a prompt to compensate for consistently bad retrieval just hides the real problem and makes it harder to fix later.
Takeaway
A RAG app that sounds convincing in a five-question demo is not a RAG app you've evaluated, it's one you've gotten lucky with. Build a real eval set from actual questions, measure retrieval with recall@k and precision@k before you touch the prompt, and check faithfulness on generation once retrieval is solid. The failures are silent by default, the only way to catch them before a user does is to actually measure for them.
Top comments (0)