Almost every RAG evaluation metric on offer needs a language model to produce it. Faithfulness, answer relevance, context precision: a model reads the answer and scores it.
Those are good metrics. They measure things that are hard to measure otherwise, and for a research sweep or a quarterly quality review, use them.
They cannot gate a build, and the reason people usually give is only half of it.
The half everyone says
A judged evaluation costs money per run. A hundred cases across a few metrics is a few thousand model calls, which is real money on every commit and every branch.
The consequence is that you move the gate. It runs nightly instead of per commit, then weekly, then on a button someone remembers to press. A check that costs a dollar gets run less, and a check that runs less catches things later, which is the property you were buying.
The half that matters more
A judge is not deterministic.
Run the same dataset against the same predictions twice and you get slightly different scores. Not wildly different, but different. That is workable for a report and disqualifying for a gate, because a gate exists to answer one question: did this change make things worse?
Answering it means comparing two numbers. If both carry noise of unknown size, you cannot separate a regression from the measurement. You get a check that fails sometimes for no reason, and the standard response to that is to disable it, usually within a week, usually by someone under deadline pressure.
So the judged metric fails twice: too expensive to run often, and untrustworthy on the difference when it does run.
What is left when you remove the model
More than you would think, and all of it deterministic.
From labelled relevant documents: precision@k, recall@k, MRR, nDCG@k, hit rate. These need a golden dataset and nothing else.
From expected answers: exact match after SQuAD style normalisation, token F1 for partial credit, required phrase presence for when an amount or a date must appear.
Recall bounds everything downstream, so it is the one to be loudest about. The model cannot cite what retrieval never fetched. At recall@k of 0.6, forty percent of your questions were unanswerable before generation began, and no amount of prompt engineering touches that.
Precision measures how much noise is in the context window, which is what predicts hallucination.
The honest limit
One measure in my own tool is a proxy, and the documentation says so.
groundedness is the share of answer content words that appear in the retrieved context:
const answerTokens = normalise(predicted).split(' ').filter((t) => t.length > 2);
const contextTokens = new Set(normalise(contexts.join(' ')).split(' '));
return answerTokens.filter((t) => contextTokens.has(t)).length / answerTokens.length;
That is lexical overlap. It will miss a fluent misreading of a passage that was correctly retrieved, which a judge would catch. It does catch an answer invented wholesale, which is the failure that gets shipped, and it costs nothing and returns the same number every time.
The trade is a cheap proxy running on every commit against an accurate measure running monthly. Which one is the better metric and which one is still switched on in March are different questions.
Missing inputs are not zeroes
One detail decides whether people trust the output.
If a case has no labelled relevant documents, the retrieval metrics for that case are omitted rather than scored zero. Aggregation skips missing values instead of averaging them in:
if (typeof value !== 'number' || Number.isNaN(value)) continue;
A partially labelled dataset should report what it can measure. Scoring the gaps as zero produces a column of failures that describes your labelling rather than your system, and nobody reads a report full of zeroes twice.
Why I wrote one
The platforms are priced for teams. Confident AI runs Free, Starter at $200 a month and Team at $2,000, with Enterprise above that. Braintrust Pro is $249. Galileo Pro is $100. Each has a free tier and each meters it: Confident AI's is two seats, one project and five test runs a week, which a per-commit gate exhausts by Tuesday. (Checked on the vendors' own pages, 9 September 2026.)
The open source libraries, RAGAS and DeepEval, compute good metrics and leave you to build the storage, the comparison and the CI gate yourself. Most solo projects end up with three tools wired together and no gate at all.
npx ragbench gate --baseline main --threshold recall@k=0.8
Local SQLite for history, zero dependencies, nothing leaves your machine. ragbench.
The gate has a second half that matters more than the thresholds, and that is the next article.
Top comments (4)
The "fails twice" point lands — but I'd argue there's a third failure that's even quieter: the judge's own distribution drifts under you. Pin temperature to 0 and seed everything and you still get a gate that silently recalibrates the day the provider ships a new model version, so your "did this change make it worse" question is now confounded by a change you didn't make. That's what pushed us to the same place you landed: deterministic metrics for the gate, judged metrics for the report. Recall@k being the loudest is right and under-appreciated — it's the one number that bounds everything downstream, and no prompt engineering buys it back. One thing I'd add: even a judged metric can survive as a monitored signal if you treat it like an SPC control chart (track the mean and variance over time, alert on out-of-band shifts) rather than a binary pass/fail. It won't gate a commit, but it catches slow regressions the deterministic set misses. Is your groundedness proxy something you'd ever promote to a gate, or is it permanently report-only by design?
Gateable on direction, not on level, and that split does not line up with report-only versus gate.
groundedness is deterministic, so it does not have the property that keeps a judge out of a gate. Same predictions, same number, every time. It already participates in the regression check, since an empty gateOn means every metric shared with the baseline. What it cannot carry is a threshold. A threshold asserts that 0.8 is good enough, and there is no reading of a lexical overlap ratio that supports that claim: 0.72 means one thing for an extractive answer and something else for an abstractive one over the same context. The number has no calibration, so only its movement is information.
Which is why your control chart framing fits it better than either label. Mean and variance over history, alert on out of band. The history is already in local SQLite, so that is a reporting question rather than a plumbing one.
On drift I think it goes a layer deeper than the judge. Taking the model out of the metric does not take it out of the system under test. If the pipeline being measured calls a hosted model, the predictions move when the provider ships a new version, and a deterministic metric will faithfully report that as a change you made. Comparing against a stored baseline run makes it worse rather than better, because the two numbers can come from different upstream versions. The deterministic set makes the measurement trustworthy without making the comparison a controlled experiment. Recording the model version alongside the commit is the only thing I have for that, and it only helps if the version is something you can pin.
Some comments may only be visible to logged-in visitors. Sign in to view all comments.