This is how I used to build an LLM feature. Write a prompt. Try it on the four or five inputs I had to hand, and adjust it until those looked right. Ship it behind a flag. Wait for the first bug report, which came within a day and was about an input nothing like my five. Adjust the prompt and check that the bug report's input worked now, without checking whether the original five still did. Ship. Repeat.
I did that for a year, and it felt like engineering because the prompts were in git and there was a flag in the config. It wasn't. It was what we all did with ordinary code before tests, where every fix was a guess and every guess could break something you'd already fixed.
The way out is the one we've already found once: write the test first. For an LLM feature the test is called an eval. The word is different and the idea is the same, and whether a team has one explains almost all of the difference between teams that ship these features reliably and teams that flail.
What an eval is, concretely
An eval is a set of inputs, an expected outcome for each, a way to score an actual output against the expected one, and a script that runs it all and prints a number. That's all there is to it, whatever mystique has grown up around the word.
In the ticket classifier I'll use as the example throughout, the inputs are support tickets, the expected outcome is a category from a fixed list, the scorer is exact match, and the number is accuracy.1
Where it differs from a unit test is that the number isn't 100, and doesn't need to be. A classifier at 94 percent might be fine to ship and one at 89 percent might not, and the eval's job is to tell you which side of the line you're on and whether a change moved you. It's a regression test with a threshold in place of pass or fail.
Build it from failures, not from imagination
The first eval set I built was 40 tickets I wrote myself, and it was useless. I wrote tickets the way I imagined them: clear, one paragraph, one problem. Real tickets cram three problems into one message. Half are replies to an earlier ticket, a quarter have the actual question on the last line after four paragraphs of context, and some are in a language the customer apologises for.
The set that worked was built from production. Every time a human corrected something the feature produced, the input and the correction went into the set. Every bug report became a case. After a month there were 200 cases, and they were the 200 inputs that had really broken the feature, which is the distribution you care about.
// evals/tickets.jsonl, one case per line
// {"id":"t-0193","input":"...","expected":"billing","source":"correction","added":"2026-03-14"}
// evals/run.ts
import { classify } from "../src/classify";
const cases = readJsonl("evals/tickets.jsonl");
let correct = 0;
const failures: string[] = [];
for (const c of cases) {
const got = await classify(c.input);
if (got === c.expected) correct++;
else failures.push(`${c.id}: expected ${c.expected}, got ${got}`);
}
console.log(`${correct}/${cases.length} = ${(100 * correct / cases.length).toFixed(1)}%`);
console.log(failures.join("\n"));
The source field matters more than it looks. A case from a human correction is ground truth. A case from a bug report is nearly ground truth. A case from me writing what I thought a ticket looked like is a guess, and after a few months I deleted all of those, because the model scored well on them and they taught me nothing.
The prompt is the thing under test
Once the set exists, the loop changes. You don't tweak the prompt and try five inputs. You tweak it and run 200, the script prints the number and the list of failures, and you read the failures.
That flips the relationship. Before, the prompt was the artefact, and the examples were how I talked myself into thinking it was fine. Now the eval set is the artefact, and the prompt is whatever currently scores best against it, which makes prompts disposable. I've rewritten the classifier's prompt from scratch four times, and each time the question was whether it scored higher than the one in main, never whether it was a good prompt.
It also makes model changes boring, and that's the most valuable part. When a new model version comes out, you change one string, run the eval and read the number. Sonnet 5 took the ticket set from 93.5 to 95.0, and the change took eleven minutes including the deploy. Without the eval that upgrade would have been a week of "it seems better?" and then a rollback when someone found the one category it had got worse at. It did get worse at one category. The eval showed which one, and the fix was two lines in the prompt.
Scoring free text
Exact match works for classification and for anything with structured output. For a reply, a summary or a rewrite there's no single correct output, so you need a different scorer. I use three, in order of preference.
Assertions on structure. A reply must mention the ticket number, must stay under 120 words, must not contain the phrase "as an AI", and must include a link if the expected one has one. These are cheap and deterministic, and they catch a surprising share of failures. A 400 word summary is wrong whatever it says.
A reference comparison. For each case, keep the reply the human actually sent, and score the model's reply against it with a similarity measure, or with a small model asked whether the two replies would lead the customer to do the same thing. This is the "LLM as judge" pattern, and it works when the judge prompt is narrow. A model answers "Do these two replies give the customer the same instructions, yes or no" reliably. It doesn't answer "Rate this reply from 1 to 10" reliably.
Human grading, sampled. Twenty cases a week, graded by the person who would have written the reply, on a two point scale: would send, or wouldn't. This keeps the other two honest, because a judge model can drift and structural assertions can't see tone.
So the number for a free text eval ends up being three numbers, and the report shows all three. It's less tidy than accuracy, and it's still a regression test.
Running it where it matters
The eval runs in CI on every change to the prompt file, the model version or the code around the call. It runs against the current model with real API calls, which costs money. For 200 cases on the classifier that's under a dollar, which is nothing next to one bad afternoon.
It fails the build if the number drops more than a threshold below main. The threshold isn't zero, because model outputs aren't perfectly deterministic even at temperature zero, and a 0.5 percent wobble on 200 cases is one case. For the classifier it's 2 percent. A bigger drop is a regression, and the author has to explain it or fix it.
It also runs nightly against a sample of production traffic, with no expected outcome, just to record what the outputs look like. One week in February the share of tickets classified as "other" doubled, and the nightly run noticed before anyone else did. A new product had launched and its tickets didn't fit any category. That needed a new category, not a bug fix, and it went into the eval set as 20 new cases.
What this does not solve
An eval tells you the feature got worse. It doesn't tell you why, and it can't tell you the feature is good in any absolute sense, only what it scores on the inputs you have. If those inputs stop looking like production, the eval is measuring the past. That's why the set has to keep growing from corrections, and why I check the added dates and get nervous when the newest case is more than a couple of weeks old.
It doesn't remove judgement either. Someone still decides whether 94 percent is good enough, and whether the six percent that fail are failing in a way that matters. The classifier sends some tickets to a neighbouring category, which takes a human ten seconds to fix. It almost never sends one to a distant category. Both count the same in the number and matter very differently, and the failure list is where you see that.
Cost, since someone will ask
The classifier's eval costs under a dollar a run and runs maybe twenty times a day across branches. The reply feature's eval, with the judge model in the loop, costs about six dollars a run and runs on merges to main and nightly. Call it 300 dollars a month for both.
Before the eval existed, there was an afternoon when the ticket classifier sent every billing ticket to the wrong queue. It cost two support engineers a day each, and a handful of customers got their refunds late.2 I'm confident it was more than 300 dollars, and it wasn't the only afternoon like it in the year before the eval set existed.
The order of operations
If you're starting an LLM feature tomorrow, begin before the prompt. Collect 30 real inputs and decide what a correct output looks like for each. Write the scorer and the runner. Get a number for a trivial prompt so you know where the floor is.
Then write the prompt, and treat it like any other code that has to pass tests: change it, run it, read the failures. Ship when the number clears the line you set beforehand.
Then feed every correction back into the set, forever. The eval doesn't end. It's the feature's test suite, and the prompt is just its current implementation.
Originally published at zeybek.dev.
Top comments (0)