DEV Community

Cover image for How to Evaluate a RAG System Before Launch: Retrieval, Groundedness, and a Test Set
AI Feed
AI Feed

Posted on Originally published at aifeed.space

How to Evaluate a RAG System Before Launch: Retrieval, Groundedness, and a Test Set

A good answer from a RAG assistant does not prove that the system is ready to launch. Retrieval may have found the right passage by chance, the model may have supplemented it with knowledge from training, and the next question may expose a gap in the corpus or a confident fabrication. A reliable evaluation separates retrieval quality from generation quality, uses a versioned test set, and separately assesses situations in which the system should not answer.

Start by Defining the Task and the Boundaries of an Acceptable Answer

RAG is an architecture in which a generative model receives information from an external index before producing an answer. The original paper on retrieval-augmented generation combined the parametric memory of a pretrained model with non-parametric memory in the form of an external Wikipedia index. However, the paper's experiments concerned specific tasks and models and do not constitute a universal acceptance standard for modern enterprise RAG systems.

Before selecting metrics, define the product's purpose. An assistant may search internal policies, answer questions about documentation, compare contract terms, or prepare a draft for a specialist. These scenarios differ in the cost of errors, acceptable autonomy, and answer boundaries.

AI Feed editorial recommendation — practical checklist: document the following parameters before beginning the evaluation. This is editorial guidance, not a proven universal protocol:

  • which sources are considered authoritative and as of what date;
  • whether information from multiple documents may be combined;
  • whether the answer must include links or quotations;
  • what to do when versions conflict;
  • when the system must refuse to answer or ask for clarification;
  • which claims require human review;
  • what request latency and cost are acceptable for the product.

For example, for the question “How many days are allowed for a return?”, the only acceptable answer may be one based on the current policy for the relevant region, with a link to the corresponding section. An old policy does not become applicable merely because it contains the same words.

Separate the Main Failure Points

A RAG pipeline can fail at no fewer than two main stages. Retrieval may fail to find evidence, return an outdated document, or add too much noise. Generation may receive the correct context but distort a number, omit a condition, or attribute a claim to a source that does not contain it. OpenAI's guide to optimizing LLM accuracy recommends distinguishing between context problems and model-behavior problems.

This framework does not cover every cause of failure. Errors may also arise during corpus preparation, filter application, access control, citation generation, external tool calls, and result presentation in the interface. Dividing failures into retrieval and generation is useful for initial diagnosis, but it is not a complete taxonomy.

If the required passage is absent from the first K results, changing the prompt will usually not correct the retrieval failure. If the passage was found but the answer contradicts it, replacing the embedding model alone will not eliminate the generation error. A single subjective assessment that “the answer looks good” is therefore poorly suited to diagnosis.

How to Build a Test Set

Each row in a test set usually corresponds to one user query. The set should be versioned together with identifiers for the corpus and configuration. Possible fields include ID, question, case type, expected behavior, relevant documents or passages, reference facts, answerability, and risk level.

AI Feed editorial recommendation — practical list of example sources: combine several data types. This list is not a validated requirement for dataset composition:

  • real questions from users or a pilot group after sensitive data has been removed;
  • questions from subject-matter experts, including professional terminology;
  • typical queries, misspellings, abbreviations, and conversational variants;
  • synthetic examples for rare combinations of conditions;
  • errors discovered during manual testing and the pilot;
  • deliberately difficult and adversarial queries.

Synthetic data can broaden coverage, but the generator may repeat its own assumptions and produce questions that are too neat. It should therefore not entirely replace real phrasing and expert annotation. OpenAI's guide to evaluation design recommends considering production, historical, expert, and synthetic data, as well as adding discovered failures to the set.

AI Feed editorial recommendation on reporting slices: do not limit reporting to a single average score; show results for product-relevant categories. For example, simple facts, questions with multiple conditions, multi-document queries, negative examples, conflicting sources, and critical scenarios can be analyzed separately. This is an editorial approach to diagnosis, not a universal reporting framework.

Create a Gold Set for Retrieval

Annotating the entire corpus is expensive. A practical compromise is to select a priority subset of questions and, for each one, identify the documents or passages sufficient to answer it. This set can be used as a gold subset. Document-level annotation is easier, but it can conceal a chunking problem: the document was found, yet the required paragraph was not included in the context. Passage-level annotation is more informative when tuning chunk sizes.

Relevance must be defined in advance. A passage may be sufficient for a complete answer, partially useful while requiring a second source, topically similar without providing evidence, or outdated and inapplicable to the relevant jurisdiction.

Some questions have several equally valid passages. If only one is annotated, a correct alternative result will be mistakenly counted as a miss. Disagreements among experts are useful to retain and analyze. The annotation version is no less important than the index version: after a policy update, a previous reference may no longer be current.

Retrieval Metrics: What They Show

Precision@K is the proportion of relevant items among the first K results. The metric shows how well the context has been cleared of noise. In a hypothetical arithmetic example where one of five passages is useful, precision@5 is 0.2.

Recall@K is the proportion of known relevant items found among the first K results. If three passages are hypothetically annotated as necessary for a complete answer and the system returns two, recall@K is 2/3. Increasing K may improve recall while also adding noise, latency, and token usage.

MRR, or mean reciprocal rank, accounts for the position of the first relevant result. If it ranks first, its reciprocal rank is 1; if it ranks fourth, it is 1/4. The values are then averaged across queries. MRR is useful when one good source is sufficient, but it represents the completeness of multi-document retrieval less effectively.

AI Feed note on the numbers: the values 5, 0.2, 3, 2/3, 1, and 1/4 above are only arithmetic illustrations of the definitions. They are not research results and are not recommended acceptance thresholds.

Hit rate, NDCG, and metrics with graded relevance are also used. Microsoft separates retrieval evaluation from final-answer evaluation in its documentation on RAG evaluators. This documentation describes, among other things, evaluations of document retrieval, groundedness, relevance, and response completeness. Metric selection should correspond to the task and the available annotations.

How to Evaluate the Generated Answer

For reproducibility, retain the exact context given to the model, the order of passages, document identifiers, prompt version, and run parameters. The answer can be evaluated against several independent criteria:

  • Groundedness: whether every verifiable claim is supported by the supplied context.
  • Relevance: whether the text answers the question rather than merely summarizing the documents.
  • Correctness: whether the conclusion matches the reference facts and domain rules.
  • Completeness: whether required conditions, exceptions, and parts of the answer are included.
  • Citation accuracy: whether links point to passages that support the corresponding claims.

Groundedness is not the same as truth. An answer may accurately restate an outdated document and be grounded relative to the supplied context, yet still be wrong for the user. The reverse is also possible: the model may state a correct fact that is absent from the supplied sources. For a closed enterprise assistant, this may violate the established answer boundaries.

Automated checks and LLM-as-a-judge approaches accelerate evaluation runs. The authors of RAGAS proposed a reference-free approach to evaluating several dimensions of RAG, including the quality of the retrieved context and the faithfulness with which it is used. However, a judge model remains a proxy evaluator. OpenAI's evaluation guide recommends comparing automated assessments with human feedback and notes that LLM judges may be inclined to prefer longer answers.

Test Situations in Which There Should Be No Answer

Negative tests complement answerable questions. The set can include queries for which the corpus contains no relevant information. Depending on the product, the expected behavior may be to report insufficient data, ask a clarifying question, or escalate to a human rather than make an unsupported guess.

AI Feed editorial recommendation — checklist of negative and adversarial tests: test the situations below. This is neither an exhaustive list nor a proven universal security standard:

  • two sources with different dates and conflicting rules;
  • partial context containing only one of several mandatory conditions;
  • a question with a false premise;
  • an instruction inside a retrieved document that proposes ignoring previous rules;
  • a user attempt to override system restrictions;
  • a request outside the user's permissions;
  • a document with similar terminology but a different country, version, or product.

AI Feed editorial recommendation on security: when designing safeguards, treat instructions from retrieved documents as untrusted data rather than commands to be executed automatically. AI Feed also recommends separately testing authorization, data access controls and isolation, source filtering, and logging. These measures are presented as practical editorial security checks, not as requirements proven by the research cited in this article. Their specific composition should be determined by the architecture, threat model, and mandatory organizational rules. Groundedness evaluation is not a substitute for a specialized security audit.

Practical Decision Table — an AI Feed Editorial Recommendation

Status of the table: this is an AI Feed diagnostic reference, not a validated legal, industry, or other universal standard. It helps identify the first hypothesis to investigate but does not prove the cause of a failure.

Observation Working hypothesis What to check first
The required passage is absent from the top K Retrieval, index, or filters Embeddings, hybrid search, metadata, query rewriting
The document was found, but the required paragraph was cut off Chunking Size, overlap, and structural splitting
Recall increases while precision decreases Noise is being added to the context Reranker, filters, and selected K
The context is correct, but a number in the answer is wrong Generation or data transformation Prompt, extraction format, and model
The answer is correct, but an exception is omitted Incomplete retrieval or generation Whether the exception is present in the supplied context
The system answers without evidence Weak no-answer policy Refusal rules and claim-support verification
The result changes between runs Variability in generation or evaluation Repeated runs, parameters, and a manual sample

Compare Configurations Reproducibly

When comparing candidates, it is useful to change one significant component at a time and run the same frozen test set. Record the corpus version, chunking, embedding model, search algorithm, top K, reranker, generative model, prompt, and evaluator settings. Without this information, an improvement is difficult to reproduce or attribute to a specific change.

AI Feed editorial recommendation: supplement average scores with a manual review of failures and pairwise blind comparisons of answers. A new configuration may improve overall recall while making critical negative cases worse. This is a practical analytical recommendation, not a mandatory universal procedure.

Practical Pre-Launch Checklist — AI Feed Editorial Recommendations

Important: the steps listed below and their number are an AI Feed editorial recommendation. They are not a research-validated minimum for readiness, a guarantee of security, or a universal acceptance standard.

  1. Define answer boundaries and the consequences of different types of error.
  2. Build a versioned set of ordinary, difficult, negative, and adversarial queries.
  3. Create a gold subset with relevant documents or passages.
  4. Calculate retrieval metrics separately from answer quality.
  5. Evaluate groundedness, relevance, correctness, completeness, and citation accuracy.
  6. Manually review critical errors and a sample of successful cases.
  7. Compare configurations on the same set and document the decision.
  8. Conduct security checks appropriate to the product's threat model.
  9. Run regression tests after changes to documents, the index, chunking, embeddings, retrieval, the model, the prompt, or the evaluator.

Evaluation does not end at release. Production failures, after sensitive data has been properly removed, can become new tests. In this way, the set gradually becomes more representative of real user behavior, although it always remains a sample rather than a complete model of future traffic.

How to Choose Numerical Thresholds

There is no universal precision@K, recall@K, MRR, or groundedness value above which every RAG system can be considered ready. Results depend on annotation completeness, question difficulty, the definition of relevance, the value of K, the cost of errors, and human oversight. A high average score can conceal an unacceptable failure in a rare but critical category.

Microsoft's documentation provides its own scales and default settings for individual evaluators, while OpenAI's guide includes examples of numerical targets for specific demonstration scenarios. Such values apply to the corresponding tools or examples and do not prove that the same thresholds are suitable for another product.

AI Feed editorial recommendation on risk-based thresholds: derive acceptance criteria from an analysis of error consequences and pilot data for the specific product. The editorial team recommends setting separate limits for critical error classes, incorrect refusals, access violations, and unsupported claims. This is AI Feed's practical risk-based approach, not a proven universal threshold or a substitute for the organization's industry requirements.

Methodology and Limitations

This article was updated on October 8, 2026, based on research papers about RAG and RAGAS, as well as documentation from Microsoft and OpenAI. Definitions and descriptions of methods directly linked to sources are separated from the editorial team's practical recommendations. All checklists, the diagnostic table, security checks, and the approach to risk-based thresholds are explicitly identified as AI Feed editorial recommendations.

Evaluation results apply only to the tested corpus, questions, annotations, models, and pipeline version. Scores from different projects cannot be compared directly if their datasets, relevance definitions, judges, prompts, or scales differ. Automated evaluation is not objective truth; important examples require review by a subject-matter expert.

A measured failure on a particular test set is an observation. Expected behavior after launch remains a forecast until confirmed by production data. The existence of a technical risk also does not mean that failure is inevitable: the outcome depends on product constraints, the interface, escalation, access controls, and human review.

Frequently Asked Questions

How many questions should a test set contain?

There is no universal number. The size is determined by coverage of the main intents, risks, and rare error classes. A small, carefully annotated gold subset may be more useful than a large collection of repetitive synthetic questions. No specific quantity should be treated as a proven threshold without context.

Can evaluation be performed without reference answers?

Partly. Reference-free evaluation accelerates development, but reference facts or answers are useful for checking correctness, completeness, and critical rules, at least for the priority portion of the set.

Which is more important: precision@K or recall@K?

It depends on the task. Clean top-ranked results matter when retrieving one exact rule. Completeness is more critical when an answer must list every requirement. The metrics are usually considered together, with the final answer evaluated separately as well.

Does high groundedness mean the system is ready?

No. It characterizes how well the answer is supported by the supplied context, but it does not guarantee that the source is current, the corpus is complete, access rights are correct, the system is secure, or users are satisfied.

When should regression tests be run?

AI Feed editorial recommendation: run them after changes that may affect the retrieved context or answer, including updates to documents, cleaning, chunking, embeddings, filters, the reranker, top K, the model, the prompt, or the evaluator. This is a practical recommendation, not a validated universal schedule.

Sources and Research

Planning Tools

Use the local model directory to estimate available memory, and the AI API price comparison to compare cloud deployment costs.

AI-assisted article, checked against the primary sources cited above.

Originally published by AI Feed: How to Evaluate a RAG System Before Launch: Retrieval, Groundedness, and a Test Set.

Top comments (0)