DEV Community

Geminate Solutions
Geminate Solutions

Posted on

Build a RAG Evaluation Set Before You Ship Your AI Feature

Build a small evaluation set before you ship a RAG feature. That means 50 to 100 real questions, each with an expected answer and the source document that should support it. Score retrieval and answer quality separately, and run the set on every change to chunking, embeddings, prompts or models.

Without it, every tweak is a guess. You change the chunk size and try three questions in a playground. It "feels better," so you ship it, and the regression stays hidden until a customer finds it.

Why a demo is not an evaluation

Most RAG features get tested the same way: someone types questions they already know the answer to. That set is tiny, leans toward easy cases and is different every time.

An evaluation set fixes three things:

  • It is stable. You ask the same questions on every run, so you can compare versions.
  • It is honest. Real questions come with typos, vague wording and topics your docs do not cover.
  • It shows where the failure is. When an answer is wrong, you can tell whether retrieval missed the right chunk or the model mishandled a good one.

Step 1: Collect 50 to 100 real questions

Do not write the questions yourself from the docs. You will use the docs' own wording, and that makes retrieval look better than it really is.

Good sources:

  • Support tickets and live chat transcripts
  • Sales call notes ("does it integrate with X?")
  • Internal channels where people ask the ops or product team things
  • Search logs from your existing help center
  • Beta users, if you have them

Aim for a mix:

  • Simple lookups with one clear source
  • Questions that need two documents combined
  • Questions worded very differently from the docs
  • Questions your corpus cannot answer, where the correct behavior is "I don't know"
  • A few adversarial cases: an outdated policy, conflicting docs, a prompt injection attempt

The unanswerable questions matter more than most teams expect. A RAG system that confidently makes up an answer is worse than one that admits a gap.

Why 50 to 100? Below 50, one flaky question moves your score too much. Above 100, labeling becomes a chore and nobody keeps the set up to date. Start small and add to it from production failures.

Step 2: Write expected answers and sources

For each question, record:

  • question: copied word for word, typos included
  • expected_answer: a short reference answer from someone who knows the domain
  • source_ids: the document or section IDs that contain the answer
  • must_include: two or three key facts the answer must contain
  • answerable: true or false

Keep it in a JSONL file in the repo, next to the code, with one case per line:

{"id": "q017", "question": "how long do u keep deleted files", "expected_answer": "Deleted files stay in the trash for 30 days, then they are removed permanently.", "source_ids": ["data-retention#trash"], "must_include": ["30 days", "permanently"], "answerable": true}
Enter fullscreen mode Exit fullscreen mode

Use stable source IDs. If your chunker creates new chunk IDs every time you re-index, map each case to the document plus section heading instead. Otherwise your retrieval scores break on every re-index.

Step 3: Score retrieval and answers separately

Teams often skip this step, but it is the most useful one.

Retrieval metrics (no LLM needed, so they are cheap and give the same result every run):

  • Hit rate @k: did any expected source show up in the top k results?
  • Recall @k: what share of the expected sources showed up?
  • MRR: how high did the first correct source rank?

Answer metrics:

  • Fact coverage: does the answer contain each must_include fact? Start with simple string matching.
  • Faithfulness: is every claim backed by the retrieved context? This usually needs an LLM judge with a strict rubric.
  • Refusal correctness: for unanswerable questions, did the system decline instead of making something up?

Reading the two together tells you what to fix:

Retrieval Answer What it means
Good Good Working
Good Bad Problem with the prompt, the model or how the context is formatted
Bad Good Luck, or the model is answering from its training data (risky)
Bad Bad Fix chunking, embeddings or query rewriting first

Step 4: A small Python harness

You do not need a framework to start. This harness covers the core loop. Replace retrieve and generate with your own pipeline.

import json, sys
from pathlib import Path

K = 5

def load_cases(path):
    lines = Path(path).read_text().splitlines()
    return [json.loads(line) for line in lines if line.strip()]

def retrieval_scores(expected, retrieved):
    top = retrieved[:K]
    hits = [s for s in expected if s in top]
    rr = 0.0
    for rank, doc_id in enumerate(top, start=1):
        if doc_id in expected:
            rr = 1.0 / rank
            break
    return {
        "hit": bool(hits),
        "recall": len(hits) / len(expected) if expected else 1.0,
        "rr": rr,
    }

def fact_coverage(answer, facts):
    text = answer.lower()
    found = [f for f in facts if f.lower() in text]
    return len(found) / len(facts) if facts else 1.0

def refused(answer):
    markers = ["i don't know", "not in the documentation", "couldn't find"]
    return any(m in answer.lower() for m in markers)

def run(path, retrieve, generate):
    results = []
    for case in load_cases(path):
        chunks = retrieve(case["question"], k=K)  # [{"id": ..., "text": ...}]
        answer = generate(case["question"], chunks)
        row = {"id": case["id"], "answer": answer}
        if case["answerable"]:
            ids = [c["id"] for c in chunks]
            row.update(retrieval_scores(case["source_ids"], ids))
            row["coverage"] = fact_coverage(answer, case["must_include"])
        else:
            row["refused_ok"] = refused(answer)
        results.append(row)
    return results

def summarize(results):
    pos = [r for r in results if "hit" in r]
    neg = [r for r in results if "refused_ok" in r]
    avg = lambda rows, key: round(sum(r[key] for r in rows) / max(len(rows), 1), 3)
    return {
        "hit_rate": avg(pos, "hit"),
        "recall": avg(pos, "recall"),
        "mrr": avg(pos, "rr"),
        "coverage": avg(pos, "coverage"),
        "refusal_accuracy": avg(neg, "refused_ok"),
    }

if __name__ == "__main__":
    from my_rag import retrieve, generate  # your pipeline
    results = run(sys.argv[1], retrieve, generate)
    Path("eval_results.json").write_text(json.dumps(results, indent=2))
    print(json.dumps(summarize(results), indent=2))
Enter fullscreen mode Exit fullscreen mode

To gate a build, compare the new run against the last run on main:

def regressions(current, baseline, tolerance=0.02):
    return {k: (baseline[k], v) for k, v in current.items()
            if v < baseline[k] - tolerance}
Enter fullscreen mode Exit fullscreen mode

If that returns anything, exit with a non-zero code.

When you add an LLM judge for faithfulness, keep the rubric narrow. For example: "Is every factual claim in the answer supported by the context? Reply PASS or FAIL with one reason." Pin the judge model version, and check a sample of its verdicts by hand every few runs. A judge you never check is just one more untested component.

Step 5: Run it on every change

The set only earns its keep if it runs automatically.

  • Run the retrieval metrics on every pull request that touches chunking, embeddings, indexing or query logic. They are fast and cheap.
  • Run the full answer evaluation on prompt or model changes, and once a night.
  • Save each run's results as a build artifact so you can compare case by case, not only the totals.
  • Fail the build when scores drop against a baseline, not when they miss a fixed number. "Hit rate dropped since the last main run" tells you exactly what to look at.

The case-by-case comparison is where most of the value is. A flat total score can hide five questions that got fixed and five that broke.

Keep the set alive

  • Turn every production bug report into a new case.
  • Review the set when the docs change, because outdated expected answers cause false failures.
  • Tag cases by category so you can spot weak areas, such as questions that need several documents.
  • Hold back a slice of cases you rarely look at. If you tune prompts against every case, you will overfit and not notice.

Common mistakes

  • Writing questions from the docs instead of collecting them from users
  • Measuring only the final answer, so you cannot tell where failures come from
  • Giving the LLM judge a vague rubric like "rate quality 1 to 10"
  • Leaving out unanswerable questions
  • Running the evaluation once before launch and never again

Decision rule

Can you answer "did this change make the feature better or worse?" with a score and a list of the affected questions? If not, you are not ready to ship. A weekend of collecting and labeling questions is usually enough to get there.

At Geminate Solutions we treat this set as part of the feature itself, because it is what makes later changes safe to ship.

For how this evaluation loop fits with chunking, embeddings and retrieval tuning, see the RAG pipeline guide.

Top comments (0)