Headline: An LLM feature cannot be unit-tested with string equality, because the same prompt returns different text on every call, so I assert on properties of the output instead of the output itself. The three changes that made my AI features safe to refactor were a version-controlled golden set of real inputs, deterministic assertions running before any LLM-as-judge call, and a CI gate on pass-rate regression rather than on a perfect score.
Key takeaways
- An eval is a test that runs a prompt against fixed inputs and scores the output on properties rather than on exact text. Evals replace
toEqualassertions for any code path whose output comes from a model. - A golden set is a version-controlled file of real inputs paired with the properties their outputs must satisfy. Every production bug I fixed in an AI feature became one more row in that file.
- Deterministic assertions — Zod schema validation, required substrings, forbidden strings, numeric ranges — catch most regressions, run in milliseconds, and cost nothing. Run them before any judge call.
- LLM-as-judge is a second model call that scores the first model's output against a written rubric. A judge is only trustworthy when its model ID is pinned, its rubric asks binary questions, and its verdicts have been compared against human labels.
- Gate CI on regression, not perfection: fail the build when the pass rate drops below a baseline committed to the repository, and treat raising that baseline as a reviewed commit.
I did not plan to build an eval harness. I shipped a route that turns a free-text support message into a structured ticket, and it worked. Weeks later a colleague edited the system prompt for an unrelated reason — a wording tweak in the tone instructions — and the extraction started inventing order IDs for messages that contained none. Every test in the repository still passed, because the only tests were for the code around the model. Nothing in CI was looking at what the model said. That is the gap evals fill.
Why can't I unit-test an LLM feature the normal way?
A large language model is a non-deterministic function, so expect(output).toBe('...') fails for reasons unrelated to correctness. Setting temperature: 0 makes output far more stable but does not make it byte-identical: providers batch requests and floating-point accumulation differs between batches, so no major provider guarantees reproducible tokens. Asserting on exact strings produces a suite that is red on Tuesday and green on Wednesday, which teams learn to ignore within a sprint.
The useful move is to split the code at the model boundary. Prompt assembly, retry logic, output parsing, tool dispatch, and every branch that consumes the parsed result are ordinary deterministic functions and deserve ordinary unit tests. Only the single call that crosses into the model needs an eval. Once the seam exists, the eval asserts on properties that are true of any correct answer: does it parse against the schema, does it stay in the requested language, does it cite only identifiers that appear in the input, does it refuse when the context contains no answer, does it never echo the system prompt.
What is a golden set and where do the cases come from?
A golden set is a version-controlled file of real inputs paired with the properties their outputs must satisfy. Mine started at roughly a dozen cases — one per input shape I could name — and grew every time production surprised me. I do not write synthetic cases when real ones exist, because the cases that catch regressions are the awkward real ones: the empty message, the message written in Arabic, the message that is only an order number, the message that is a prompt-injection attempt pasted out of an email.
Two storage rules have paid for themselves. First, keep the fixtures as JSON or JSONL next to the prompt they test, so the case list is reviewable in a pull request. Second, keep the prompt in its own file rather than in a template literal inside a route handler, so a prompt change and its eval results show up in the same diff.
// evals/ticket-extraction.eval.test.ts
import { describe, it, expect } from 'vitest';
import { generateObject } from 'ai';
import { z } from 'zod';
import cases from './fixtures/tickets.json';
import { TICKET_PROMPT } from '../src/prompts/ticket';
const Ticket = z.object({
category: z.enum(['billing', 'bug', 'feature', 'other']),
urgency: z.number().int().min(1).max(5),
summary: z.string().max(200),
orderIds: z.array(z.string()),
});
describe('ticket extraction', () => {
for (const c of cases) {
it.concurrent(c.name, async () => {
const { object } = await generateObject({
model: 'anthropic/claude-sonnet-5', // pinned ID, routed via AI Gateway
schema: Ticket,
temperature: 0,
system: TICKET_PROMPT,
prompt: c.input,
});
expect(object.category).toBe(c.expected.category);
// grounding: never invent an order ID that is not in the message
for (const id of object.orderIds) expect(c.input).toContain(id);
});
}
});
generateObject from the Vercel AI SDK is doing double duty here: it validates the model output against the Zod schema before returning, so a shape regression throws instead of flowing into an assertion. That is the cheapest eval in the file and it is the one that would have caught my invented-order-ID bug.
When should I use a deterministic check instead of LLM-as-judge?
Use a deterministic check whenever the failure can be expressed as a predicate over the string, and reach for a judge only for properties that require reading comprehension. Deterministic checks are free, instant, and never disagree with themselves; a judge is a second model call with its own error rate and its own bill.
| Property to check | Method | Why |
|---|---|---|
| Output shape and JSON validity | Deterministic (Zod schema) | Schema validation is exact; a judge adds cost and can only be worse. |
| Grounding: every ID, price, or quote appears in the input | Deterministic (substring or set check) | Hallucinated identifiers are string-detectable, and this is the most damaging class of error. |
| Forbidden content: prompt leakage, internal URLs, competitor names | Deterministic (regex denylist) | Safety properties must be exact, not probabilistic. |
| Language, length, and format constraints | Deterministic | Word counts and script ranges are trivially computable. |
| Did the answer actually address the question? | LLM-as-judge | Requires comprehension; no predicate expresses it. |
| Is the answer semantically equivalent to a reference? | LLM-as-judge or embedding similarity | Many correct phrasings exist, so exact matching under-reports success. |
How do I keep the LLM judge from drifting?
Treat the judge as production code with a version, not as an oracle. Pin the exact model ID — claude-sonnet-5, never a floating alias — because a silent provider-side model change reads as a quality regression in your product. Ask binary questions instead of 1–10 scores: "does the answer state a refund deadline, yes or no" is reproducible, while "rate helpfulness out of ten" drifts and turns your threshold into an arbitrary number. Keep the rubric in a committed file that is reviewed like any other source. And calibrate: hand-label twenty to thirty outputs yourself, run the judge over the same outputs, and look at where the two disagree. A judge you have never compared to a human is a number, not a measurement. Where budget allows, use a different model family for the judge than for the generator, because models tend to rate their own phrasing style favourably.
const Verdict = z.object({
grounded: z.boolean(),
answersQuestion: z.boolean(),
reason: z.string().max(300),
});
export async function judge(question: string, answer: string, context: string) {
const { object } = await generateObject({
model: 'anthropic/claude-sonnet-5', // pinned; bump in a reviewed commit
schema: Verdict,
temperature: 0,
system: JUDGE_RUBRIC, // committed file, binary criteria only
prompt: `<context>${context}</context>\n<question>${question}</question>\n<answer>${answer}</answer>`,
});
return object;
}
The reason field is not decoration. When an eval fails at 2am, a one-line explanation from the judge is the difference between a triageable failure and a red square nobody can interpret.
How do I run evals in CI without paying for every pull request?
Run evals in two tiers. Tier one is a fast subset on every pull request: all deterministic checks plus a handful of representative cases, no judge calls. Tier two is the full set with judging, triggered nightly and on any pull request that touches the prompt or fixture directories — that is a paths filter in the workflow, not a human convention. A documentation-only pull request should cost nothing.
Three details keep tier one honest. Model calls are IO-bound, so run cases concurrently with it.concurrent and cap concurrency below the provider rate limit. Give CI its own API key routed through a single gateway, so eval spend is a legible line item instead of a mystery inside the product budget. And expect genuine flakiness: a case that fails once and passes on retry is a signal about output variance, so I retry a failed case once and only fail the build when it fails twice.
# .github/workflows/evals.yml
name: evals
on:
pull_request:
paths: ['src/prompts/**', 'evals/**', 'src/lib/ai/**']
schedule:
- cron: '0 3 * * *'
jobs:
full:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with: { node-version: '24' }
- run: npm ci
- run: npx vitest run evals --maxConcurrency=4
env:
AI_GATEWAY_API_KEY: ${{ secrets.AI_GATEWAY_EVAL_KEY }}
What pass rate should fail the build?
Gate on regression rather than on a fixed absolute score. A rule of "100% of evals must pass" fails constantly on borderline cases and gets disabled within a month. Instead I commit an evals/baseline.json holding the current pass rate per suite, and CI fails when the measured rate drops below the baseline by more than a small tolerance. Raising the baseline after a genuine improvement is a normal reviewed commit, which makes quality a ratcheting number that lives in git history rather than a claim in a standup.
Two categories are exempt from the statistics. Schema validity and safety checks — prompt leakage, invented identifiers, denylisted content — hard-fail at zero tolerance, because a single occurrence is a bug rather than a sample. Keeping those separate is what lets the fuzzy score stay fuzzy without anyone arguing that a leak is "within tolerance".
I still do not have a perfect measurement of whether the ticket extractor is good. What I have is a file of the ways it has been wrong, a build that goes red when it gets worse in one of those ways, and permission to refactor the prompt without holding my breath. For a non-deterministic feature, that has turned out to be the whole win.
FAQ
Q: Do I need a dedicated eval framework to start?
A: No. Vitest or the built-in node:test runner plus a JSON fixture file covers the first several months. Tools such as Promptfoo, Evalite, and Braintrust become worth adopting once you want dataset management, run-over-run tracing, and a UI for reviewing judge disagreements.
Q: How many cases should a golden set contain?
A: Start with one case per input shape you can name, which is usually ten to twenty, and grow it from production failures. Coverage of distinct failure modes matters far more than the case count.
Q: Does temperature: 0 make model output deterministic?
A: No. It sharply reduces variance but no major provider guarantees byte-identical output, because request batching changes floating-point accumulation. Assert on properties of the output regardless of temperature.
Q: Should the judge model be the same model that generated the answer?
A: Prefer a different model family when budget allows, since models tend to favour their own phrasing style. If you must reuse the same model, calibrate the judge against human labels before trusting its scores.
Q: How do I stop eval costs from growing with the team?
A: Run deterministic checks first and skip judging any case that already failed one, restrict full runs to prompt-touching pull requests and a nightly schedule, and cache model responses keyed by a hash of prompt file, model ID, and input.
Originally published on devya.dev. Also on eng-ahmed.com. Built by Devya Solutions.
Top comments (0)