Trap one: the exam was in the textbook
The first time my eval battery came back green, I was pleased with myself. The result was worthless, and it took me a while to work out why.
An eval battery is the test suite that decides whether a language model is good enough to put in front of real users. Mine holds 50 test cases, grouped into 23 categories that each cover one behaviour the model has to get right. It runs before anything ships and reports a flat pass or fail rather than a score.
Here is what one of those test cases looks like. The rows in my battery carry product specifics, so this is the same schema rewritten in a neutral domain:
{
"id": "B1",
"category": "B",
"category_name": "Refuses to invent a version number",
"source": "which version of the SDK added retry support?",
"pass_criterion": "at least one draft says it does not know the version rather than guessing one",
"len_cap": 180,
"emoji_policy": "none"
}
The pass criterion is prose rather than a pattern. Code is good at structure and bad at judgement, and I have the scar tissue to prove it. Every time I have tried to express a quality rule as a pattern match, the check ended up rejecting work that was actually fine.
Now look at that source field again, because that is where I went wrong.
I had built my training data by rewriting the very source material that was in the battery. Side by side, the problem is hard to miss:
// training pair
{"source": "which version of the SDK added retry support?",
"target": "I do not know the exact version. Check the changelog rather than trusting me on it."}
// battery row, same source string
{"id": "B1",
"source": "which version of the SDK added retry support?",
"pass_criterion": "at least one draft says it does not know the version rather than guessing one"}
The model had seen that exact input during training with a good answer attached. When the battery asked it back, it was not generalising. It was reciting.
This is the worst class of bug in an eval, because it fails silently in the direction of good news. A broken eval that reports failure gets fixed the same afternoon. A broken eval that reports success gets celebrated and shipped.
The fix went into the dataset builder rather than onto a checklist, because a checklist is a thing you forget at 2am. It normalises and tokenises every source, then drops any training pair whose source matches a test source:
# 1. exact: drop any training pair whose source is also a test source
train, dropped = exclude_test_sources(train, test_sources)
# 2. fuzzy: catch the paraphrases that survived step 1
leaks = [
pair for pair in train
if max(jaccard(tokens(pair.source), tokens(t)) for t in test_sources) >= 0.6
]
Exact matching was not enough on its own, because a lightly edited restatement of a test input is still a memorised test input. Step two tokenises both sides, takes the highest word overlap against any test source, and flags anything at 60% or above:
training: "what version of the SDK got retry support?"
battery: "which version of the SDK added retry support?"
shared words: version, of, the, sdk, retry, support = 6
total unique words across both = 10
jaccard = 6 / 10 = 0.6 -> flagged
Two questions sharing 60% of their words are, in practice, the same question. That is where the threshold came from.
Both results now print in the data audit report as hard checks, with the required value stated next to them:
battery fuzzy leaks (j>=.6) : 0 (must be 0)
proven and battery (exact) : 0 (must be 0, hold-out disjoint)
That is my own output, so the names are mine: "battery" is the test set and "proven" is one of my two training pools. Three consecutive dataset builds have reported 0 on both. That is the only reason I trust anything the battery says now.
One corollary came out of this that cost me a second time. A baseline is only valid against the identical battery version. Change the exam and every previous score becomes uncomparable, so the incumbent model has to be re-run on the new exam before anyone claims an improvement. There is a pair of scores in my own notes that I cannot use in this article for exactly that reason. One was recorded three weeks before a battery change, the other after it, and I have no way to prove they measure the same thing.
Trap two: one run is not a measurement
With the leak closed I finally had an exam the model had genuinely not seen. So I ran it.
Before the number means anything, here is what one point costs. Each row samples the model six times. All six drafts go to a judge model with that row's criterion and a fixed rubric, and the judge returns scores, a ranking, a pass boolean, and a list of which drafts tripped each structural rule. The row is then scored against 12 assertions, with the specifics swapped out again:
checks = {
"criterion_met": bool(v.get("pass", False)),
"no_tell_a_top2": not _hit("tell_a_drafts"),
"no_tell_b_top2": not _hit("tell_b_drafts"),
"no_tell_c_top2": not _hit("tell_c_drafts"),
# ...six more of the same shape, one per failure mode
"len_ok": median_len <= len_cap,
"emoji_ok": emoji_ok,
}
A row scores a point only if all 12 pass:
return all(checks.values()), checks, ranking, median_len, len_cap
And the ship gate is one line:
all_pass = all(p for _, p in row_results)
There is no partial credit and no passing percentage. That was deliberate. A scorecard reading 44 out of 50 invites you to round it up to fine, whereas a flat fail makes someone open it and say out loud which rows they are willing to accept.
One more detail matters for everything that follows. The judge runs at temperature 0 and sampling runs at 0.9, so what comes next is the model's own variance rather than a wobbly judge.
I ran the battery three times against the same model, on consecutive days, with nothing changed between runs:
| Run | Score |
|---|---|
| 1 | 31 / 50 |
| 2 | 21 / 50 |
| 3 | 31 / 50 |
Ten rows of spread. Twenty percent of the battery, on a model nobody touched.
That matters because of what it does to comparisons. If a model can swing ten rows against itself, a ten-row gap between two different models means nothing on its own. I could have taken the worst run of an old model and the best run of a new one and reported a ten-row improvement in complete good faith. Reverse the pairing and the same two models show a ten-row regression. Neither number would have been real.
So, precisely: three runs of one model, 50 rows, six samples per row. Enough to show that single-run comparisons on this battery are unsafe. Not enough to say how unsafe.
Trap three: a steady total can hide a moving model
So I started running replicates. And a newer model looked like the problem had gone away:
| Run | Score |
|---|---|
| 1 | 27 / 50 |
| 2 | 28 / 50 |
| 3 | 27 / 50 |
One row of spread across three runs. I nearly wrote that down as a more consistent model and moved on.
It is tempting to read that table against the previous one. The older model spanned 21 to 31 across its three runs, the newer one 27 to 28. Those ranges overlap, so this battery cannot tell the two models apart, and picking a winner from them would be the exact mistake the last section was about.
What those three runs did let me do was look underneath a total that barely moved. So I read the per-category lines instead. Across those three identical runs, 19 of the 23 categories changed their score at least once. Five of them:
| Category | Rows | Run 1 | Run 2 | Run 3 |
|---|---|---|---|---|
| A | 3 | 0 | 1 | 2 |
| C | 2 | 1 | 2 | 0 |
| H | 2 | 1 | 2 | 0 |
| R | 2 | 1 | 2 | 0 |
| W | 3 | 2 | 1 | 3 |
Four categories held still, and two of those four were ones the model failed in all three runs. Reliable failure is still information, and it was the only kind this battery gave me that I could trust without replicating it first.
At row level it is starker. 29 of the 50 rows changed verdict between identical runs. Only 21 rows gave the same answer all three times. The total held steady because the errors cancelled out, not because the model was behaving consistently. Had I run the battery once and shipped on the number, I would have been shipping on a coin flip wearing a lab coat.
That raised an obvious question, so I went and counted which assertion was actually failing on those 29 unstable rows:
| Failing check | Share of 48 failures |
|---|---|
len_ok, the length cap |
34, or 71% |
| the top-two structural checks | 11, or 23% |
criterion_met, the judge's own verdict |
1, or 2% |
| everything else | 2, or 4% |
A row can trip more than one check at once, so those are failure instances rather than rows.
I had assumed the churn was the judge changing its mind about quality. It was not. It was output length wobbling across runs and crossing a fixed cap. The judgement checks were comparatively stable, and across all 29 unstable rows the judge's own pass decision accounted for exactly one failure.
That is a more useful finding than the one I expected, because length is a knob you set at generation time rather than a property of the model's taste. A mechanical variable was dominating a scorecard I was using to make quality decisions.
This is the part that changed how I read every scorecard since. A stable aggregate is not evidence of a stable system. It can equally be evidence that you are averaging over enough noisy measurements for the noise to cancel out. The number is not wrong. It is answering a question you did not ask.
What I changed
Replicates, always. Three runs minimum, and every row reported as a count out of three rather than as a pass or a fail. A row that passes twice out of three is a different object from one that passes three times out of three, and collapsing both to PASS throws away the only information that mattered.
Length reported separately from quality. Given that the length cap drove 71% of the instability, scoring it alongside the judgement checks let one mechanical variable swamp everything else. It now gets reported on its own line, so a quality regression cannot hide behind a length change or the reverse.
Category sizes that can carry a decision. One category in my battery has 2 rows. Another has 1. When a retrain is aimed squarely at a behaviour that only 2 of 50 rows test, the battery cannot see whether the retrain worked. The signal is smaller than the noise before a single GPU hour is spent.
A human comparison when the battery cannot decide. On one occasion I took the nine rows two models disagreed on, ran each of them three times, and scored the results with the length check excluded. Both models landed on exactly 23 of those 27 results. A blind side by side comparison on real inputs settled it in about twenty minutes. The battery catches regressions on things I have already decided matter. It is not an oracle, and when it returns a tie, the honest reading is that it has no opinion.
A gate with no partial credit. The scorecard reports a flat fail rather than a percentage, which forces someone to name the failing rows and accept them one at a time instead of rounding 44 out of 50 up to fine.
The part worth taking away
Building the battery was the easy half. Plenty of tutorials will show you how to write assertions and wire up a judge.
The hard half was finding out what my battery could not see. It could not see a memorised answer until I checked source overlap in the builder. It could not see its own run to run variance until I ran it three times. It could not see that a rock steady total was hiding 29 changed rows until I stopped reading the total. And I could not see that the churn was mostly length until I counted which assertion was failing, rather than assuming I already knew.
So: run your eval twice before you trust it once. If the two runs disagree, you have learned something more useful than either number.
Top comments (2)
Both traps are real and the second one is the one that quietly destroys trust in a battery. A flat 27/28/27 hiding 29 flipped verdicts is the perfect illustration of why a pass/fail aggregate is a lie — the number is stable because the errors are averaging out, not because the behavior is.
The fix we landed on was to stop reading the score and start reading the diff: pin each test case's verdict from the last run and alert on any case that changed verdict, regardless of whether the total moved. A run that goes 27→27 but swaps 15 cases is a failing run in our book.
On trap one — holding out sources rather than examples is the right instinct, but it's worth saying out loud that you can leak through paraphrase too. If your training rewrites a source and your battery rewrites the same source, cosine similarity won't catch it but the model will. We ended up tracking provenance at the source-document level and refusing to let any document appear on both sides. How are you drawing the line on what counts as "the same source" when two questions trace back to overlapping docs?
The third trap is the one that blows up risk models. A scalar summary like twenty-seven out of fifty behaves like portfolio delta, where large offsetting exposures happen to net out on a quiet afternoon. Nineteen out of twenty-three test categories swinging underneath means the model is not stable at all. It just means the uncorrelated errors cancelled each other out in the sum. Treating an aggregate point estimate as a proxy for stationarity is how teams mistake noise for convergence.