<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Hua Li</title>
    <description>The latest articles on DEV Community by Hua Li (@hualiphd).</description>
    <link>https://dev.to/hualiphd</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4167655%2Fc683c000-3f28-4266-b15f-5055f9a35c90.jpeg</url>
      <title>DEV Community: Hua Li</title>
      <link>https://dev.to/hualiphd</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/hualiphd"/>
    <language>en</language>
    <item>
      <title>I found the optimal top-k for my RAG system. Then I ran it again.</title>
      <dc:creator>Hua Li</dc:creator>
      <pubDate>Wed, 07 Oct 2026 13:00:00 +0000</pubDate>
      <link>https://dev.to/hualiphd/i-found-the-optimal-top-k-for-my-rag-system-then-i-ran-it-again-3oi3</link>
      <guid>https://dev.to/hualiphd/i-found-the-optimal-top-k-for-my-rag-system-then-i-ran-it-again-3oi3</guid>
      <description>&lt;p&gt;I built a small retrieval-augmented generation system over a corpus of research papers and wanted to answer an ordinary question: how many chunks should I retrieve? So I did what the tutorials suggest — swept top-k across 3, 5, 8 and 12, scored each configuration with RAGAS, and read off the winner.&lt;/p&gt;

&lt;p&gt;The answer was clean. k=8 led on all four metrics. Context recall was 0.53 at k=3, peaked at 0.83 at k=8, and fell back to 0.64 at k=12 — a textbook inverted U, with a story to match. Too few chunks and the generator starves; too many and distractors dilute the prompt. I had a number, a mechanism, and a plot.&lt;/p&gt;

&lt;p&gt;Then I ran k=8 twice more, changing nothing.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Run&lt;/th&gt;
&lt;th&gt;Faithfulness&lt;/th&gt;
&lt;th&gt;Answer relevancy&lt;/th&gt;
&lt;th&gt;Context precision&lt;/th&gt;
&lt;th&gt;Context recall&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;k=8, run 1&lt;/td&gt;
&lt;td&gt;0.767&lt;/td&gt;
&lt;td&gt;0.900&lt;/td&gt;
&lt;td&gt;0.655&lt;/td&gt;
&lt;td&gt;0.833&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;k=8, run 2&lt;/td&gt;
&lt;td&gt;0.590&lt;/td&gt;
&lt;td&gt;0.804&lt;/td&gt;
&lt;td&gt;0.585&lt;/td&gt;
&lt;td&gt;0.611&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;k=8, run 3&lt;/td&gt;
&lt;td&gt;0.640&lt;/td&gt;
&lt;td&gt;0.835&lt;/td&gt;
&lt;td&gt;0.687&lt;/td&gt;
&lt;td&gt;0.625&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Recall ranged from 0.611 to 0.833 across three runs of an identical configuration. The 0.833 that made k=8 look like a peak was the high draw of three.&lt;/p&gt;

&lt;p&gt;Set the variation &lt;em&gt;within&lt;/em&gt; one configuration against the differences &lt;em&gt;between&lt;/em&gt; configurations:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Spread between configs (means)&lt;/th&gt;
&lt;th&gt;Spread within k=8&lt;/th&gt;
&lt;th&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;faithfulness&lt;/td&gt;
&lt;td&gt;0.084&lt;/td&gt;
&lt;td&gt;0.177&lt;/td&gt;
&lt;td&gt;noise exceeds signal&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;answer relevancy&lt;/td&gt;
&lt;td&gt;0.059&lt;/td&gt;
&lt;td&gt;0.096&lt;/td&gt;
&lt;td&gt;noise exceeds signal&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;context precision&lt;/td&gt;
&lt;td&gt;0.023&lt;/td&gt;
&lt;td&gt;0.102&lt;/td&gt;
&lt;td&gt;noise exceeds signal&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;context recall&lt;/td&gt;
&lt;td&gt;0.162&lt;/td&gt;
&lt;td&gt;0.222&lt;/td&gt;
&lt;td&gt;noise exceeds signal&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;On every metric, one configuration varies more against itself than the configurations vary against each other. k=5, k=8 and k=12 are not distinguishable on this evidence. The only ordering that survives is that k=3 under-retrieves — it has the lowest recall in every run.&lt;/p&gt;

&lt;p&gt;The inverted U was an artifact of running each configuration once.&lt;/p&gt;

&lt;h2&gt;
  
  
  The judge moves the numbers more than the parameter does
&lt;/h2&gt;

&lt;p&gt;While building the harness I had changed judge models — the free tier I started on ran out of quota. That turned out to be a second experiment I hadn't meant to run.&lt;/p&gt;

&lt;p&gt;Same corpus, same generator, same k=5. Only the judge differs:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Judge A&lt;/th&gt;
&lt;th&gt;Judge B&lt;/th&gt;
&lt;th&gt;Difference&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;faithfulness&lt;/td&gt;
&lt;td&gt;0.840&lt;/td&gt;
&lt;td&gt;0.582&lt;/td&gt;
&lt;td&gt;0.258&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;context recall&lt;/td&gt;
&lt;td&gt;0.917&lt;/td&gt;
&lt;td&gt;0.604&lt;/td&gt;
&lt;td&gt;0.313&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;context precision&lt;/td&gt;
&lt;td&gt;0.752&lt;/td&gt;
&lt;td&gt;0.628&lt;/td&gt;
&lt;td&gt;0.125&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;answer relevancy&lt;/td&gt;
&lt;td&gt;0.777&lt;/td&gt;
&lt;td&gt;0.861&lt;/td&gt;
&lt;td&gt;0.084&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Changing the judge moved faithfulness by 0.26 and recall by 0.31. Changing top-k from 3 to 8 — the parameter I was actually studying — moved faithfulness by 0.12.&lt;/p&gt;

&lt;p&gt;Part of that gap is different grading. Part is something subtler: the two judges fail differently. One of them intermittently returned output the parser couldn't read, so those samples scored as missing and dropped out of the average. The two numbers were therefore computed over different subsets of questions. Both effects are properties of the judge, and both are invisible in a reported score.&lt;/p&gt;

&lt;p&gt;I want to be careful here: I have one run per judge, so I can't cleanly separate a judge effect from an unlucky draw. The honest statement is that the judge shift is about the same size as the run-to-run variation I measured, which is already enough to make absolute RAGAS numbers meaningless without naming the judge that produced them.&lt;/p&gt;

&lt;h2&gt;
  
  
  The significance test couldn't have found anything
&lt;/h2&gt;

&lt;p&gt;Comparing configuration means throws away the pairing, so I did it properly: per-question differences between two configurations, Wilcoxon signed-rank, each question acting as its own control. Nothing reached significance. For a while I read that as "no difference."&lt;/p&gt;

&lt;p&gt;It wasn't. Wilcoxon discards tied pairs, and with &lt;em&gt;n&lt;/em&gt; non-zero pairs the smallest two-sided p-value it can return is 2^(1−n):&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Non-zero pairs&lt;/th&gt;
&lt;th&gt;Smallest achievable p&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;1.000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;0.500&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;0.250&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;0.125&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;0.0625&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;0.031&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Below six non-zero pairs, significance at α = 0.05 is arithmetically impossible regardless of the data.&lt;/p&gt;

&lt;p&gt;My context-precision comparison had three non-zero pairs and returned p = 0.250. My recall comparison had one and returned p = 1.000. Both are exactly the floor — every non-zero pair pointed the same way, the most extreme result obtainable, and it still couldn't clear the threshold.&lt;/p&gt;

&lt;p&gt;My sample collapsed twice over: 18 questions, then 8–9 matched pairs after judge failures, then 1–3 non-zero pairs after ties. Each stage looked survivable. Together they left a test with no power at all.&lt;/p&gt;

&lt;p&gt;The tie rate is the more interesting half, though. Eight of nine pairs scored &lt;em&gt;identically&lt;/em&gt; between k=8 and k=12. On most questions the chunks ranked 9 through 12 changed nothing the judge could see. That's a finding about retrieval, and it explains the null better than low power does.&lt;/p&gt;

&lt;h2&gt;
  
  
  The first thing my harness found was a bug in my own corpus
&lt;/h2&gt;

&lt;p&gt;Before any of this, the eval surfaced something I hadn't gone looking for. My "21-paper corpus" was 15 papers. Each one's PDF and its metadata record were being indexed as separate works, so retrieving both counted as retrieving two relevant documents. Every precision number I'd computed was inflated.&lt;/p&gt;

&lt;p&gt;Deduplicating on a canonical work ID fixed it. But I'd been quoting 21 for weeks, and nothing except an evaluation harness would have caught it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'd do differently, and what I'd keep
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Replicate before believing.&lt;/strong&gt; At minimum three runs per configuration. If within-configuration spread approaches the between-configuration differences, report that and stop — it's the finding.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Report the judge and generator with every metric.&lt;/strong&gt; A RAGAS score without them isn't comparable to anyone else's, or to your own from last month.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Check the power before spending the quota.&lt;/strong&gt; Computing 2^(1−n) is one line and it tells you in advance whether a comparison can possibly reach significance. Mine couldn't, and I found out afterwards.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Report ties and sample attrition.&lt;/strong&gt; "Not significant" and "no measurable difference" are different claims, and the gap between them lives in the counts that usually go unreported.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Classify the missing values, don't just count them.&lt;/strong&gt; A missing score can be a generation failure, a judge parse failure, a metric undefined by design, or a genuine zero. Pooled together, a parser fix changes your reported mean by changing which questions are in it. [Suggested by Ahmet Özel (&lt;a class="mentioned-user" href="https://dev.to/ahmetozel"&gt;@ahmetozel&lt;/a&gt;) in the comments.]&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Separate what a judge grades from what code can check.&lt;/strong&gt; My suite has two tracks. Queries about corpus statistics and listings return zero chunks by design, so context-based metrics are undefined for them; those get deterministic assertions instead. Those assertions, and the intent-routing accuracy, came back at 100% across every run. Code-driven checks carried no variance at all. The LLM-judged metrics carried all of it. That contrast is worth designing for.&lt;/p&gt;

&lt;h2&gt;
  
  
  Limits
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The replication spread conflates two sources of variance.&lt;/strong&gt; My three identical-configuration runs re-generated answers &lt;em&gt;and&lt;/em&gt; re-scored them, so the 0.22 recall range mixes generator sampling with judge variance, and I reported it as a single quantity. Separating them needs two runs the harness can support but I didn't do: replay the judge over saved answers to isolate judge variance, and re-generate under a fixed judge to isolate the other half. Credit to Ahmet Özel (&lt;a class="mentioned-user" href="https://dev.to/ahmetozel"&gt;@ahmetozel&lt;/a&gt;) for pointing this out in the comments.&lt;/p&gt;

&lt;p&gt;This is one system, one 15-paper corpus, 18 scored content questions, and two judges. I'm not claiming a general law about RAGAS or about LLM-as-judge evaluation — I'm claiming that the standard procedure, run on a realistic small setup, produced a confident answer that replication dissolved, and that the failure modes were all invisible from the reported numbers.&lt;/p&gt;

&lt;p&gt;If your evaluation is larger, it may be fine. If it's about this size, and you ran each configuration once, you have the result I had before I ran it again.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Harness, golden set, ablation and results: &lt;a href="https://github.com/huali10044/research-rag" rel="noopener noreferrer"&gt;github.com/huali10044/research-rag&lt;/a&gt;. The system itself is live at &lt;a href="https://huali-research-rag.streamlit.app/" rel="noopener noreferrer"&gt;huali-research-rag.streamlit.app&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>rag</category>
      <category>machinelearning</category>
      <category>ai</category>
      <category>llm</category>
    </item>
  </channel>
</rss>
