DEV Community

Chen Yuan
Chen Yuan

Posted on Originally published at dispatch-blog.hashnode.dev

How to Build a Post-Launch Eval Canary That Tells a Real LLM Regression From Sampling Noise

Is the model actually getting worse, or did I just get unlucky on a handful of prompts?

That question is why threads like "is it just me or is it dumber today" keep recurring, and it is the question a post-launch eval canary has to answer with a number instead of a feeling. The reference implementation here is livenerf, a long-running, deterministic-as-possible benchmark tracking whether Claude Opus 5.5 (released 2026-09-22) gets quietly worse after launch. This article teaches the reusable method behind it: how to freeze prompts, pin the harness, calibrate a panel of "sometimes right" questions, compute paired per-item statistics with clustered standard errors, run a control arm, watch output-token counts as an early signal, and pre-register the rule that decides when you are allowed to say "regression."

You do not need a frontier model or a cluster. You need a few hundred labeled questions, one machine, and the discipline to fix everything except the model.

Step 1: Accept that the harness is part of the treatment

A canary compares a model to itself across time. Every other moving part becomes a confound. The livenerf v0 runner executes a Claude Max subscription through headless Claude Code (claude -p), pins the CLI version, disables the auto-updater, and refuses to run when the version no longer matches — because a changed harness looks exactly like a changed model. The same logic applies to your system prompt, your sampling parameters, your tool list, and your working directory.

The cheapest enforcement is a guard that fails loudly. The code below is illustrative and was not executed for this article; treat the version string and command shape as placeholders for your own stack.

# illustrative, not executed
import re
import subprocess
import sys

PINNED_CLI = "claude-code 2.4.1"

def assert_pinned_harness() -> None:
    out = subprocess.run(
        ["claude", "--version"], capture_output=True, text=True, check=True
    ).stdout.strip()
    if out != PINNED_CLI:
        sys.exit(
            f"harness drift: expected {PINNED_CLI!r}, got {out!r}. "
            "Refusing to run; a harness change is indistinguishable from a model change."
        )
Enter fullscreen mode Exit fullscreen mode

Version pinning only works if the run itself is hermetic. livenerf uses a short frozen system prompt, no tools, no MCP, no CLAUDE.md or memory, one turn, and a fixed empty working directory. Code answers are graded by hidden tests in a sandbox outside the model. That last decision matters more than it sounds: if an LLM judge scores your outputs, the judge can drift too, and you will be measuring two moving objects at once.

Step 2: Freeze a prompt file, not a prompt idea

A canary needs an artifact you can hash. Keep the system prompt, the item text, and the scoring function in version control, and make the runner read them from disk rather than from a database that someone can quietly edit.

# illustrative, not executed
import hashlib
import json
from pathlib import Path

PANEL = Path("panel/frozen_panel_v1.jsonl")
SYSTEM = Path("panel/system_v1.txt")

def load_frozen():
    system = SYSTEM.read_text(encoding="utf-8")
    items = [json.loads(line) for line in PANEL.read_text(encoding="utf-8").splitlines()]
    digest = hashlib.sha256(
        SYSTEM.read_bytes() + PANEL.read_bytes()
    ).hexdigest()[:12]
    return system, items, digest
Enter fullscreen mode Exit fullscreen mode

Two files, one digest, no edits without a new panel version. The digest is what you log next to every run so a later comparison can prove it used the same questions.

Step 3: Calibrate toward "sometimes right" questions

Here is the part most homegrown evals skip, and it is the part that buys you statistical power. Under a logit-shift model, a question with pass rate p carries information p(1-p) per sample. Questions the model always answers correctly contribute nearly nothing, and questions it always fails contribute nearly nothing. Only questions near the middle of the difficulty curve move when capability moves.

livenerf screened 2,336 GPQA Diamond, MMLU-Pro, competition-math and AIME 2025–26 questions with 4 samples each. Opus 5.5 got roughly 93% right on the first try, and 97% of questions turned out to be always right or always wrong. The surviving 78 "sometimes right" questions became the frozen panel. A second effect showed up: questions selected for being sometimes-right regress toward the mean. On fresh samples their pass rate rose from 54.7% to 62.0%, so the power calculation has to use the fresh rates rather than the selection rates.

If you calibrate your own panel, do the same: sample each candidate question several times, keep the ones with intermediate pass rates, then re-estimate their rates on data you did not use for selection. Budget the panel around that middle band, because that is where the information is.

Step 4: Compare items, not just averages

A raw accuracy comparison between two time windows carries question difficulty inside it. If day 30 happens to sample harder items, the score drops for reasons that have nothing to do with the model. The fix is to compare each item to itself: for every question, take the score in the current window minus the score in the launch-week baseline, then average those per-item deltas.

Because each item is measured repeatedly, the per-item deltas are not independent, and naive standard errors will be too small. That is what clustered standard errors are for, and it is the approach livenerf takes, following Evan Miller's "Adding Error Bars to Evals". Cluster by question; with one observation per item per window, the per-item delta is the cluster.

# illustrative, not executed
import numpy as np

def paired_delta_with_clustered_se(baseline: np.ndarray, current: np.ndarray) -> dict:
    """baseline/current are item-level scores in [0,1], same item order."""
    d = current - baseline
    n = len(d)
    mean = d.mean()
    # cluster-robust variance with one cluster per item:
    # var(mean) = sum_i (d_i - mean)^2 / (n * (n - 1))
    var = ((d - mean) ** 2).sum() / (n * (n - 1))
    se = np.sqrt(var)
    return {
        "n_items": n,
        "delta": float(mean),
        "se": float(se),
        "ci99": (float(mean - 2.576 * se), float(mean + 2.576 * se)),
    }
Enter fullscreen mode Exit fullscreen mode

The choice of 99% rather than 95% is deliberate: a canary fires repeatedly, so it needs a higher bar per look to keep false alarms rare. With one run per day of the whole panel, livenerf's own power calculation puts the detectable accuracy change at roughly 7.5 points per 10-day window. That is the instrument's resolution, and knowing it is what stops you from over-reading a 2-point wobble.

Step 5: Add a control arm and watch tokens first

The control arm is the cheapest way to avoid blaming the model for something else. livenerf runs claude-opus-5 on the same harness alongside the tracked model. If the control arm moves by the same amount, the change is probably infrastructure, a shared dependency, or the harness — not a silent downgrade of the tracked model.

Output-token counts are the early signal. In livenerf's validation, lowering effort to low cut output tokens by 62% and accuracy by 8.3 ± 4.5 points; effort medium cut tokens by 26% and accuracy by 4.2 ± 3.9 points. Tokens moved before accuracy did, which makes them a leading indicator worth logging on every run even when accuracy looks flat.

Step 6: Pre-register the decision rule

A canary that decides what counts as a regression after seeing the data is a vibes machine with extra steps. The livenerf decision rule was committed to public git before any series data existed, so the timestamp is meaningful. It calls a change only if three conditions hold together: the 99% interval excludes zero in two consecutive 10-day windows, the effect is at least 3 points, and the control arm does not show the same move. Null results and improvements get published just as loudly as regressions.

The two-window requirement is what protects against a single unlucky stretch. Here is a compact check you can adapt; it assumes you already computed intervals per window for the tracked model and the control.

# illustrative, not executed
MIN_EFFECT = 0.03

def call_regression(windows: list[dict], control: list[dict]) -> str:
    """windows: per-10-day dicts with 'ci99' and 'delta' for the tracked model."""
    if len(windows) < 2:
        return "insufficient data"
    last_two = windows[-2:]
    excludes_zero = all(w["ci99"][0] > 0 or w["ci99"][1] < 0 for w in last_two)
    big_enough = all(abs(w["delta"]) >= MIN_EFFECT for w in last_two)
    control_moved = any(
        c["ci99"][0] > 0 or c["ci99"][1] < 0 for c in control[-2:]
    )
    if excludes_zero and big_enough and not control_moved:
        return "regression called"
    return "no call"
Enter fullscreen mode Exit fullscreen mode

Note what the rule does not do. It does not treat the launch-week baseline as ground truth. livenerf's README is explicit that launch week is a reference point, not a ceiling: launch week could be the worst week because of capacity strain, and past quality incidents turned out to be infrastructure bugs rather than deliberate downgrades. The baseline is the thing you compare against, not the thing you trust.

Step 7: Publish the limits alongside the numbers

An instrument that hides its blind spots will be misused. livenerf documents a concrete one: a same-family swap (Opus 5 in place of Opus 5.5) was not distinguishable in a validation's worth of samples, at -3.8 ± 6.3 points and -23% tokens. That is a known limit of the instrument, written down in advance so a null result is not later spun as a clean bill of health.

Every canary has a floor. The honest move is to state yours: how many points you can detect, over what window, with what false-alarm rate, and which kinds of changes you would miss entirely. Then keep the raw logs append-only, as livenerf does with Inspect .eval files, so anyone can re-analyze without re-running.

The method in one paragraph

Freeze the prompts and the harness, calibrate toward questions that are sometimes right, compare each item against its own baseline, cluster your standard errors by question, run a control arm, log tokens as a leading indicator, and pre-register the rule that authorizes the word "regression." None of that requires a lab. It requires deciding, before you look, what would count as evidence — and then publishing the null results with the same energy as the alarming ones. The alternative is the loop we already know: vibes versus vibes, one launch week at a time.


Originally published on Dispatch.

Top comments (0)