Two retrieval runs for the same query. Relevant document is D.
Run A: [D, x, x, x, x, x, x, x]
Run B: [x, x, x, x, x, x, x, D]
Precision@8 is identical: one relevant document out of eight, 0.125 both times. Recall@8 is identical: you found the one that existed, 1.0 both times. Hit rate is 1.0 both times.
Every set-based metric says these runs are the same. For a RAG system they are not.
Why position is most of the quality
Retrieval feeds a context window, and a context window is ordered and finite.
Take the top 4 chunks and run A includes D while run B does not. Recall@8 was 1.0 and your actual recall at the size you use is 0. The metric measured a depth you never read from.
Even when everything fits, position matters. Models attend unevenly across long contexts, and material buried in the middle of a large window gets used less reliably than material at the top. The practical effect today is that rank order is part of your quality, and a metric that ignores it ignores the part you can most easily improve.
What nDCG does
Discounted cumulative gain gives each hit a value that shrinks with position, using a logarithm so the penalty is steep near the top and gentle further down:
top.forEach((id, i) => {
if (relevantSet.has(id)) dcg += 1 / Math.log2(i + 2);
});
Position 1 contributes 1 / log2(2) = 1.0. Position 2 contributes 1 / log2(3) ≈ 0.63. Position 8 contributes 1 / log2(9) ≈ 0.32.
That shape matches how people and models use ranked results. The gap between first and second is worth more than the gap between seventh and eighth.
The n is the normalisation. Raw DCG is not comparable across queries, because a query with five relevant documents can score higher than one with a single relevant document without being better retrieval. So you divide by the best achievable ordering for that query:
let idcg = 0;
const ideal = Math.min(relevant.length, k);
for (let i = 0; i < ideal; i++) idcg += 1 / Math.log2(i + 2);
return idcg === 0 ? 0 : dcg / idcg;
Every relevant document first, as many as could fit in k. The result is 1.0 for perfect ordering and comparable across queries with different numbers of relevant documents.
For the two runs above, A scores 1.0 and B scores about 0.32, which is the difference precision and recall could not see.
Which metric answers which question
They are not competing. A report should have all of them.
| Metric | The question it answers |
|---|---|
hit_rate |
Did we find anything at all? The floor below which nothing else matters. |
recall@k |
Could the model possibly have got it right? It cannot cite what was never fetched. |
precision@k |
How much noise is in the context window? This is what predicts hallucination. |
mrr |
How far down is the first useful thing? |
ndcg@k |
Is the good material near the top, or merely present? |
Read them in that order, and fix them in that order. Hit rate first, since ranking is irrelevant if you are not retrieving anything relevant. Then recall, which bounds everything downstream. Then nDCG, usually where the cheapest wins are, because reranking is a smaller change than reindexing.
Two traps in reporting them
Report k and use the k you actually read. recall@100 is a fine diagnostic and a bad headline if your prompt takes 5 chunks.
Do not average away a metric with missing inputs. A case with no relevance labels has no meaningful recall, and scoring it zero drags the average down for a labelling gap.
All five are in ragbench, computed deterministically with no model in the loop, so the same dataset and predictions give the same numbers on any machine. The nDCG test asserts exactly the property this article is about: a document found first scores higher than the same document found last, while precision and recall report both as identical.
Top comments (5)
"Recall@8 was 1.0 and your actual recall at the size you use is 0" deserves to be on a poster. I have shipped that exact mistake: evaluated at k=20 because the dashboard looked healthier, then fed the model top 5. If someone changes one thing after reading this, make k in the metric equal to k in the prompt.
One caveat on top of nDCG: it still depends on relevance labels you had to produce, and in my experience the labels drift faster than the ranker does. Every time the corpus changes shape I re-audit a sample of the qrels, otherwise the discount curve is being applied to a stale notion of relevant.
Orphaned label as the weaker statement is the right call. From outside the scorer you genuinely cannot separate "deleted" from "now ranks below k for every question", and since the bound on recall is identical either way, claiming the stronger version would only be wrong more often for no extra information.
The re-chunk case is the argument that sells corpus turnover: every id changes, retrieval quality is untouched, every metric goes to zero. Print-only rather than gating is also the right default, because a check that fails a build on legitimate corpus work gets switched off within a week and then you have neither the gate nor the signal.
One thing I would watch: orphaned labels accumulate for benign reasons as the corpus grows, so the raw count drifts upward on its own. A rate against total labels, or the delta versus the baseline run, carries more information than the absolute number, and you are already storing the retrieved ids per run to support exactly that.
Shipped in 0.3.0. orphanRate is surfaced, and the warning leads with the change rather than the level:
Two things came out of building it that I had not thought through when I answered you.
The recovered line is not symmetry for its own sake. A dataset repair that reached nothing looks identical to one that reached everything if only the total is printed, and the total does not move at all when one label goes unreachable while another comes back. That is the run where you most want to be told.
And an empty stored orphan set has to read differently from a missing one. Empty means the previous run had clean labels, so every orphan now is new, which is the single case most worth reporting. Null means the run predates the column and a delta against it would be invented. The obvious implementation, an early return when the baseline array is empty, swallows exactly the case the feature exists for. There is a test whose whole job is that distinction.
run compares against the previous run under the same label, gate against the baseline, so the line reads the way the metric table above it does.
The rate is half there and I buried it. The warning reads "3 of 47 labelled documents", which is the rate written as a fraction, and the JSON hands you both numbers so you can divide. auditLabels computes orphanRate internally and then does not surface it, which is the worst of the two options.
The delta is not there at all, and you are right that it carries more. One correction on the mechanics though: what I store per run is the retrieved corpus, not the orphan set, and the labels move between runs too, so the delta cannot be recomputed from what history holds. Two runs with the same orphan count can have completely different sets behind them. It needs the orphan ids stored per run alongside the corpus, which is a small addition on a migration path that already exists.
What that buys is the sentence worth printing. Not "3 labelled documents are unreachable", which is a standing condition you learn to scroll past, but "2 more than at the baseline, and here they are", which is an event.
Built it, thanks. Shipped as two deterministic signals that run on every run and gate nothing.
An orphaned label is a labelled document that came back for no question in the whole run. That is deliberately weaker than "this document was deleted", because from outside a scorer cannot tell a deleted document from one that now ranks below k everywhere. It is also the more useful of the two statements: either way recall for the cases that need it cannot reach 1, whatever you do to the ranker, which is the same bound the post is about.
The second is corpus turnover against the baseline run, which meant storing the set of retrieved ids per run rather than only the aggregates.
Neither one fails a build. A corpus that changed shape is usually somebody doing their job, and a check that failed the build for that would be switched off within a week, same reasoning as the regression tolerance. They print under their own heading and never touch the verdict.
The case that convinced me was a re-chunk that changes every document id. Retrieval quality is untouched and every retrieval metric goes to zero. Without those two lines that is a day spent bisecting a ranker that never changed.
What it still cannot do is your actual job. It catches labels that stopped pointing at anything. It cannot catch a label that still points at a document which is no longer the right answer, and that is the half of drift only a person reading a sample will find. The tool can say when to go and look. It cannot look.