Build a small evaluation set before you ship a RAG feature. That means 50 to 100 real questions, each with an expected answer and the source document that should support it. Score retrieval and answer quality separately, and run the set on every change to chunking, embeddings, prompts or models.
Without it, every tweak is a guess. You change the chunk size and try three questions in a playground. It "feels better," so you ship it, and the regression stays hidden until a customer finds it.
Why a demo is not an evaluation
Most RAG features get tested the same way: someone types questions they already know the answer to. That set is tiny, leans toward easy cases and is different every time.
An evaluation set fixes three things:
- It is stable. You ask the same questions on every run, so you can compare versions.
- It is honest. Real questions come with typos, vague wording and topics your docs do not cover.
- It shows where the failure is. When an answer is wrong, you can tell whether retrieval missed the right chunk or the model mishandled a good one.
Step 1: Collect 50 to 100 real questions
Do not write the questions yourself from the docs. You will use the docs' own wording, and that makes retrieval look better than it really is.
Good sources:
- Support tickets and live chat transcripts
- Sales call notes ("does it integrate with X?")
- Internal channels where people ask the ops or product team things
- Search logs from your existing help center
- Beta users, if you have them
Aim for a mix:
- Simple lookups with one clear source
- Questions that need two documents combined
- Questions worded very differently from the docs
- Questions your corpus cannot answer, where the correct behavior is "I don't know"
- A few adversarial cases: an outdated policy, conflicting docs, a prompt injection attempt
The unanswerable questions matter more than most teams expect. A RAG system that confidently makes up an answer is worse than one that admits a gap.
Why 50 to 100? Below 50, one flaky question moves your score too much. Above 100, labeling becomes a chore and nobody keeps the set up to date. Start small and add to it from production failures.
Step 2: Write expected answers and sources
For each question, record:
-
question: copied word for word, typos included -
expected_answer: a short reference answer from someone who knows the domain -
source_ids: the document or section IDs that contain the answer -
must_include: two or three key facts the answer must contain -
answerable: true or false
Keep it in a JSONL file in the repo, next to the code, with one case per line:
{"id": "q017", "question": "how long do u keep deleted files", "expected_answer": "Deleted files stay in the trash for 30 days, then they are removed permanently.", "source_ids": ["data-retention#trash"], "must_include": ["30 days", "permanently"], "answerable": true}
Use stable source IDs. If your chunker creates new chunk IDs every time you re-index, map each case to the document plus section heading instead. Otherwise your retrieval scores break on every re-index.
Step 3: Score retrieval and answers separately
Teams often skip this step, but it is the most useful one.
Retrieval metrics (no LLM needed, so they are cheap and give the same result every run):
- Hit rate @k: did any expected source show up in the top k results?
- Recall @k: what share of the expected sources showed up?
- MRR: how high did the first correct source rank?
Answer metrics:
-
Fact coverage: does the answer contain each
must_includefact? Start with simple string matching. - Faithfulness: is every claim backed by the retrieved context? This usually needs an LLM judge with a strict rubric.
- Refusal correctness: for unanswerable questions, did the system decline instead of making something up?
Reading the two together tells you what to fix:
| Retrieval | Answer | What it means |
|---|---|---|
| Good | Good | Working |
| Good | Bad | Problem with the prompt, the model or how the context is formatted |
| Bad | Good | Luck, or the model is answering from its training data (risky) |
| Bad | Bad | Fix chunking, embeddings or query rewriting first |
Step 4: A small Python harness
You do not need a framework to start. This harness covers the core loop. Replace retrieve and generate with your own pipeline.
import json, sys
from pathlib import Path
K = 5
def load_cases(path):
lines = Path(path).read_text().splitlines()
return [json.loads(line) for line in lines if line.strip()]
def retrieval_scores(expected, retrieved):
top = retrieved[:K]
hits = [s for s in expected if s in top]
rr = 0.0
for rank, doc_id in enumerate(top, start=1):
if doc_id in expected:
rr = 1.0 / rank
break
return {
"hit": bool(hits),
"recall": len(hits) / len(expected) if expected else 1.0,
"rr": rr,
}
def fact_coverage(answer, facts):
text = answer.lower()
found = [f for f in facts if f.lower() in text]
return len(found) / len(facts) if facts else 1.0
def refused(answer):
markers = ["i don't know", "not in the documentation", "couldn't find"]
return any(m in answer.lower() for m in markers)
def run(path, retrieve, generate):
results = []
for case in load_cases(path):
chunks = retrieve(case["question"], k=K) # [{"id": ..., "text": ...}]
answer = generate(case["question"], chunks)
row = {"id": case["id"], "answer": answer}
if case["answerable"]:
ids = [c["id"] for c in chunks]
row.update(retrieval_scores(case["source_ids"], ids))
row["coverage"] = fact_coverage(answer, case["must_include"])
else:
row["refused_ok"] = refused(answer)
results.append(row)
return results
def summarize(results):
pos = [r for r in results if "hit" in r]
neg = [r for r in results if "refused_ok" in r]
avg = lambda rows, key: round(sum(r[key] for r in rows) / max(len(rows), 1), 3)
return {
"hit_rate": avg(pos, "hit"),
"recall": avg(pos, "recall"),
"mrr": avg(pos, "rr"),
"coverage": avg(pos, "coverage"),
"refusal_accuracy": avg(neg, "refused_ok"),
}
if __name__ == "__main__":
from my_rag import retrieve, generate # your pipeline
results = run(sys.argv[1], retrieve, generate)
Path("eval_results.json").write_text(json.dumps(results, indent=2))
print(json.dumps(summarize(results), indent=2))
To gate a build, compare the new run against the last run on main:
def regressions(current, baseline, tolerance=0.02):
return {k: (baseline[k], v) for k, v in current.items()
if v < baseline[k] - tolerance}
If that returns anything, exit with a non-zero code.
When you add an LLM judge for faithfulness, keep the rubric narrow. For example: "Is every factual claim in the answer supported by the context? Reply PASS or FAIL with one reason." Pin the judge model version, and check a sample of its verdicts by hand every few runs. A judge you never check is just one more untested component.
Step 5: Run it on every change
The set only earns its keep if it runs automatically.
- Run the retrieval metrics on every pull request that touches chunking, embeddings, indexing or query logic. They are fast and cheap.
- Run the full answer evaluation on prompt or model changes, and once a night.
- Save each run's results as a build artifact so you can compare case by case, not only the totals.
- Fail the build when scores drop against a baseline, not when they miss a fixed number. "Hit rate dropped since the last main run" tells you exactly what to look at.
The case-by-case comparison is where most of the value is. A flat total score can hide five questions that got fixed and five that broke.
Keep the set alive
- Turn every production bug report into a new case.
- Review the set when the docs change, because outdated expected answers cause false failures.
- Tag cases by category so you can spot weak areas, such as questions that need several documents.
- Hold back a slice of cases you rarely look at. If you tune prompts against every case, you will overfit and not notice.
Common mistakes
- Writing questions from the docs instead of collecting them from users
- Measuring only the final answer, so you cannot tell where failures come from
- Giving the LLM judge a vague rubric like "rate quality 1 to 10"
- Leaving out unanswerable questions
- Running the evaluation once before launch and never again
Decision rule
Can you answer "did this change make the feature better or worse?" with a score and a list of the affected questions? If not, you are not ready to ship. A weekend of collecting and labeling questions is usually enough to get there.
At Geminate Solutions we treat this set as part of the feature itself, because it is what makes later changes safe to ship.
For how this evaluation loop fits with chunking, embeddings and retrieval tuning, see the RAG pipeline guide.
Top comments (0)