DEV Community

Cover image for I found the optimal top-k for my RAG system. Then I ran it again.
Hua Li
Hua Li

Posted on AI-assisted

I found the optimal top-k for my RAG system. Then I ran it again.

I built a small retrieval-augmented generation system over a corpus of research papers and wanted to answer an ordinary question: how many chunks should I retrieve? So I did what the tutorials suggest — swept top-k across 3, 5, 8 and 12, scored each configuration with RAGAS, and read off the winner.

The answer was clean. k=8 led on all four metrics. Context recall was 0.53 at k=3, peaked at 0.83 at k=8, and fell back to 0.64 at k=12 — a textbook inverted U, with a story to match. Too few chunks and the generator starves; too many and distractors dilute the prompt. I had a number, a mechanism, and a plot.

Then I ran k=8 twice more, changing nothing.

Run Faithfulness Answer relevancy Context precision Context recall
k=8, run 1 0.767 0.900 0.655 0.833
k=8, run 2 0.590 0.804 0.585 0.611
k=8, run 3 0.640 0.835 0.687 0.625

Recall ranged from 0.611 to 0.833 across three runs of an identical configuration. The 0.833 that made k=8 look like a peak was the high draw of three.

Set the variation within one configuration against the differences between configurations:

Metric Spread between configs (means) Spread within k=8
faithfulness 0.084 0.177 noise exceeds signal
answer relevancy 0.059 0.096 noise exceeds signal
context precision 0.023 0.102 noise exceeds signal
context recall 0.162 0.222 noise exceeds signal

On every metric, one configuration varies more against itself than the configurations vary against each other. k=5, k=8 and k=12 are not distinguishable on this evidence. The only ordering that survives is that k=3 under-retrieves — it has the lowest recall in every run.

The inverted U was an artifact of running each configuration once.

The judge moves the numbers more than the parameter does

While building the harness I had changed judge models — the free tier I started on ran out of quota. That turned out to be a second experiment I hadn't meant to run.

Same corpus, same generator, same k=5. Only the judge differs:

Metric Judge A Judge B Difference
faithfulness 0.840 0.582 0.258
context recall 0.917 0.604 0.313
context precision 0.752 0.628 0.125
answer relevancy 0.777 0.861 0.084

Changing the judge moved faithfulness by 0.26 and recall by 0.31. Changing top-k from 3 to 8 — the parameter I was actually studying — moved faithfulness by 0.12.

Part of that gap is different grading. Part is something subtler: the two judges fail differently. One of them intermittently returned output the parser couldn't read, so those samples scored as missing and dropped out of the average. The two numbers were therefore computed over different subsets of questions. Both effects are properties of the judge, and both are invisible in a reported score.

I want to be careful here: I have one run per judge, so I can't cleanly separate a judge effect from an unlucky draw. The honest statement is that the judge shift is about the same size as the run-to-run variation I measured, which is already enough to make absolute RAGAS numbers meaningless without naming the judge that produced them.

The significance test couldn't have found anything

Comparing configuration means throws away the pairing, so I did it properly: per-question differences between two configurations, Wilcoxon signed-rank, each question acting as its own control. Nothing reached significance. For a while I read that as "no difference."

It wasn't. Wilcoxon discards tied pairs, and with n non-zero pairs the smallest two-sided p-value it can return is 2^(1−n):

Non-zero pairs Smallest achievable p
1 1.000
2 0.500
3 0.250
4 0.125
5 0.0625
6 0.031

Below six non-zero pairs, significance at α = 0.05 is arithmetically impossible regardless of the data.

My context-precision comparison had three non-zero pairs and returned p = 0.250. My recall comparison had one and returned p = 1.000. Both are exactly the floor — every non-zero pair pointed the same way, the most extreme result obtainable, and it still couldn't clear the threshold.

My sample collapsed twice over: 18 questions, then 8–9 matched pairs after judge failures, then 1–3 non-zero pairs after ties. Each stage looked survivable. Together they left a test with no power at all.

The tie rate is the more interesting half, though. Eight of nine pairs scored identically between k=8 and k=12. On most questions the chunks ranked 9 through 12 changed nothing the judge could see. That's a finding about retrieval, and it explains the null better than low power does.

The first thing my harness found was a bug in my own corpus

Before any of this, the eval surfaced something I hadn't gone looking for. My "21-paper corpus" was 15 papers. Each one's PDF and its metadata record were being indexed as separate works, so retrieving both counted as retrieving two relevant documents. Every precision number I'd computed was inflated.

Deduplicating on a canonical work ID fixed it. But I'd been quoting 21 for weeks, and nothing except an evaluation harness would have caught it.

What I'd do differently, and what I'd keep

Replicate before believing. At minimum three runs per configuration. If within-configuration spread approaches the between-configuration differences, report that and stop — it's the finding.

Report the judge and generator with every metric. A RAGAS score without them isn't comparable to anyone else's, or to your own from last month.

Check the power before spending the quota. Computing 2^(1−n) is one line and it tells you in advance whether a comparison can possibly reach significance. Mine couldn't, and I found out afterwards.

Report ties and sample attrition. "Not significant" and "no measurable difference" are different claims, and the gap between them lives in the counts that usually go unreported.

Classify the missing values, don't just count them. A missing score can be a generation failure, a judge parse failure, a metric undefined by design, or a genuine zero. Pooled together, a parser fix changes your reported mean by changing which questions are in it. [Suggested by Ahmet Özel (@ahmetozel) in the comments.]

Separate what a judge grades from what code can check. My suite has two tracks. Queries about corpus statistics and listings return zero chunks by design, so context-based metrics are undefined for them; those get deterministic assertions instead. Those assertions, and the intent-routing accuracy, came back at 100% across every run. Code-driven checks carried no variance at all. The LLM-judged metrics carried all of it. That contrast is worth designing for.

Limits

The replication spread conflates two sources of variance. My three identical-configuration runs re-generated answers and re-scored them, so the 0.22 recall range mixes generator sampling with judge variance, and I reported it as a single quantity. Separating them needs two runs the harness can support but I didn't do: replay the judge over saved answers to isolate judge variance, and re-generate under a fixed judge to isolate the other half. Credit to Ahmet Özel (@ahmetozel) for pointing this out in the comments.

This is one system, one 15-paper corpus, 18 scored content questions, and two judges. I'm not claiming a general law about RAGAS or about LLM-as-judge evaluation — I'm claiming that the standard procedure, run on a realistic small setup, produced a confident answer that replication dissolved, and that the failure modes were all invisible from the reported numbers.

If your evaluation is larger, it may be fine. If it's about this size, and you ran each configuration once, you have the result I had before I ran it again.


Harness, golden set, ablation and results: github.com/huali10044/research-rag. The system itself is live at huali-research-rag.streamlit.app.

Top comments (2)

Collapse
 
ahmetozel profile image
Ahmet Özel •

The attrition counts make the judge comparison more revealing than the headline scores. I would keep a fixed per-question table across every configuration, with separate fields for generation failure, judge parse failure, undefined metric and a valid zero score. Otherwise a parser improvement can change the average by changing its population.

For the repeated top-k sweep, saving each generated answer and rescoring it separately would help partition generator variance from judge variance. Then the same-answer judge replay and the full pipeline rerun answer different questions, while both preserve the question pairing and show which additional chunks were actually available.

Collapse
 
hualiphd profile image
Hua Li •

Both of these are better than what I had, and the second one found a confound I'd missed.

On the per-question table — agreed, and the parser point is the part I hadn't thought through. Right now a missing score could be a generation failure, a judge parse failure, a metric that's undefined by design, or a genuine zero, and they're all pooled. So improving the parser would move my reported mean without anything about the system changing. I've added that to the checklist in the post.

On the sweep — you're right that my three identical-config runs re-generated and re-judged, so the 0.22 recall spread I reported mixes generator sampling with judge variance and I presented it as one number. The harness does store each generated answer, so replaying the judge over saved answers isolates the judge half without burning generation quota, and a re-generation under a fixed judge isolates the other. That's running next, and I'll post the split if it's interesting. I've added the caveat to the Limits section and credited you in both places.

Thanks — this was more useful than the post.