Your eval harness trusts a witness it never cross-examined
Every agent benchmark I have built ends with the same line of code: a model grades another model's output, and I treat the resulting number as a measurement. Rubric, strong model, score out of five.
Ship it.
The judge is a language model. It has documented failure modes — it likes long answers, it likes the first answer, it likes confident answers, and it drifts when you change its temperature. We know all of this. And then we grade with it anyway and never once ask whether it was any good at grading.
So I built the inversion: a rig where the thing under test is the grader. It's called Arbiter — a gold set of human labels, four experiments, and one blunt verdict on whether the judge deserves to be trusted for your rubric.
What it measured surprised me more than I expected, and the most interesting result isn't about the judge at all.
The setup
Gold set (20 human-labelled items + 8 pairwise items)
|
v
Judge agent (Pydantic AI, output_type=Verdict) <-- every call recorded
|
+--> 1. Agreement raw %, Cohen's kappa, 5x5 confusion matrix, bootstrap CIs
+--> 2. Position (A,B) then (B,A) -> flip rate, mean score delta, Cohen's d
+--> 3. Verbosity padded ~2x vs terse <=0.6x -> mean delta, effect size
+--> 4. Drift second temperature, second model -> kappa per condition, KS shift
|
v
Corrected score + VERDICT: trustworthy / NOT trustworthy (failing criterion named)
The judge's output type is deliberately tiny:
class Verdict(BaseModel):
score: int # 1..5
criteria: dict[str, bool] # one key per rubric criterion
rationale: str # <= 240 chars
confidence: float # 0..1
and the system prompt is fixed, so the thing being measured is the judge and not a moving instruction:
You are a strict evaluator. Rubric criteria:
- correctness: the answer is factually right
- completeness: it covers every part of the question
- clarity: it is unambiguous and well-structured
Score 1-5. Reply ONLY with JSON:
{"score":<1-5>,"criteria":{"correctness":true|false,"completeness":true|false,"clarity":true|false},
"rationale":"<=240 chars","confidence":<0-1>}
Judge only what is present. Do not reward length.
That last line is doing a lot of work. It is the hypothesis under test.
One table is the whole design
The architecture decision that shaped everything: every judge call is a row, and every statistic is computed by reading rows back. No experiment ever calculates from an in-memory response.
CREATE TABLE judge_calls(
id INTEGER PRIMARY KEY AUTOINCREMENT,
run_id TEXT NOT NULL, experiment TEXT NOT NULL, -- agreement | position | verbosity | drift
item_ref TEXT NOT NULL, variant TEXT NOT NULL, -- original | padded | terse | forward | reverse
base_url TEXT NOT NULL, model TEXT NOT NULL, temperature REAL NOT NULL,
system_prompt TEXT NOT NULL, user_prompt TEXT NOT NULL, prompt_sent TEXT NOT NULL,
raw_output TEXT,
parse_ok INTEGER NOT NULL, -- 0 => failure, and score IS NULL
score INTEGER, criteria_json TEXT, rationale TEXT, confidence REAL,
human_score INTEGER, preferred TEXT, slot_scores_json TEXT,
latency_ms INTEGER, error TEXT, created_at TEXT NOT NULL
);
Three things fall out of this that I did not expect to need:
-
A report is rebuildable.
GET /api/report?run_id=…recomputes every number from stored rows, months later, without calling a model again. The UI's report step loads from a URL alone. - A crash loses nothing. Half-finished runs still report whatever completed.
-
A failure cannot be quietly dropped. An unparseable verdict is a row like any other, with
parse_ok = 0. It has to be actively excluded, and the exclusion is visible.
The prompt that went on the wire is stored too, which is more fiddly than it sounds. Pydantic AI can append its own instructions to what you passed in, so the honest "exact prompt" is the assembled request, not your string. In 2.54.0 that lives on the instance, not the class:
request = next(m for m in result.all_messages() if type(m).__name__ == "ModelRequest")
texts = [p.content for p in request.parts if isinstance(getattr(p, "content", None), str)]
prompt_sent = "\n\n".join(texts) # SystemPromptPart + UserPromptPart, as actually sent
Failures are rows, not defaults
The rule when a judge returns something unusable: record it, keep the raw text, score nothing. Never clamp, never retry-into-a-number, never substitute 0 — because 0 is a verdict, and a silent 0 turns a broken measurement into a bad result.
I expected this path to be rare. It is the single busiest code path in the project.
parse failures counted, never coerced — causes: rationale over the 240-char limit x3
failure g06: rationale is 244 chars, over the 240 limit
failure g15: rationale is 255 chars, over the 240 limit
failure g16: rationale is 255 chars, over the 240 limit
The model cannot hold a 240-character limit. Across a 116-call run, 17–19 calls fail to parse, and the dominant cause is a rationale that is 241–290 characters long. Nothing about that is visible in a normal eval harness — you would just see slightly noisier scores and quietly attribute the noise to the model under test.
To make this reachable I had to leave the typed model bare (score: int, not Field(ge=1, le=5)) and enforce ranges in my own validator. If I constrained the schema, Pydantic AI would coerce or retry at the provider, and the failure I most wanted to measure would disappear into a retry loop.
Proving the accounting also needed an escape hatch from structured output. With output_type=Verdict, a prompt begging for prose still returns valid JSON — the provider enforces the schema. So the verification path swaps both the type and the prompt:
system_prompt = PROSE_PROBE_SYSTEM_PROMPT if prose_mode else JUDGE_SYSTEM_PROMPT(rubric)
agent = Agent(model=model, system_prompt=system_prompt,
output_type=str if prose_mode else Verdict)
Three prose calls in, three rows with parse_ok = 0, score IS NULL, and the raw paragraph attached. Not one of them scored 0.
The sign trap in position bias
Position bias looks trivial: judge (A,B), judge (B,A), count flips. The bug is in the score delta, and it is a good one because the wrong version is self-cancelling and therefore invisible.
The judge names a slot, never an answer. Mapping back to identity is the whole trick:
if winner_slot == "slot1":
preferred_answer = "A" if order == "forward" else "B"
elif winner_slot == "slot2":
preferred_answer = "B" if order == "forward" else "A"
A flip is then "the content-winner changed", which is not the same as "the slot choice changed" — a judge that picks slot 1 both times is perfectly consistent by slot and completely biased by content.
For the score delta, each answer must be compared first-position minus second-position. A is first in the forward ordering; B is first in the reverse one. So the two subtract in opposite directions:
deltas.append(slots_f["answerA"] - slots_r["answerA"]) # A: first in forward
deltas.append(slots_r["answerB"] - slots_f["answerB"]) # B: first in reverse
I wrote the second line the obvious way first — slots_f["answerB"] - slots_r["answerB"] — and a test caught it, because the test fixture was built to produce a known +3 inflation and got 0.0. The wrong sign makes a real positional bias cancel itself out and report "no bias found". That is the kind of bug that survives code review and quietly launders a benchmark.
The denominator trap
Same category, caught later by looking at a screenshot instead of a test.
Verbosity reports three numbers: padded − terse, padded − original, terse − original. My first implementation averaged each side over whatever had parsed on that side. Since 15–19 of 20 items survive each pass, padded − original was comparing a mean over one subset against a mean over a different subset. A "difference" with no shared items is not a difference.
Now the comparisons against the unmodified answer only use items where all three parsed:
for item_id in sorted(by_item):
baseline = baseline_rows.get(item_id) # the agreement pass, not this experiment
p, t = by_item[item_id].get("padded"), by_item[item_id].get("terse")
if p and t and baseline and p.parse_ok and t.parse_ok and baseline.parse_ok:
triples.append((baseline.score, p.score, t.score))
and the report publishes its own denominators so you can see the sample each figure rests on:
"denominators": { "paddedScored": 15, "terseScored": 19, "paired": 15,
"triplesWithOriginalBaseline": 15, "originalScored": 18 }
Padding has to be generated in code
The verbosity experiment needs two variants of the same content at different lengths. Asking a model to produce them is a trap: it will add content, and sometimes add facts, and then you are measuring "does the judge like more information" instead of "does the judge like more words".
So padding is a fixed table of information-free sentences rotated by a hash of the item id:
FILLER = (
"It is worth noting that this topic has several dimensions that merit careful consideration.",
"Before continuing, it may help to step back and consider the framing of the question itself.",
...
)
def pad(text: str, item_id: str) -> str:
base, n = text.rstrip(), len(text.rstrip())
chosen, current, i = [], n, _seed_offset(item_id, len(FILLER))
while current < TARGET_RATIO * n: # aim for 2.0x
candidate = FILLER[i % len(FILLER)]
addition = len(candidate) + (1 if chosen else 0)
if current + addition > MAX_RATIO * n and current >= MIN_RATIO * n: # 2.4x ceiling / 1.7x floor
break
chosen.append(candidate); current += addition; i += 1
return base + "\n\n" + " ".join(chosen)
The tests are the interesting part: the original bytes must survive untouched, no filler sentence may contain a digit or a currency symbol (no new facts), every appended sentence must be a member of FILLER, and regenerating must be byte-identical.
terse() is the opposite operation and it is honest about being ugly: delete discourse-only sentences, delete scaffolding phrases, drop stop words, until the item fits under 0.6×. It never adds a word, so every surviving token is a substring of the original. Terse answers read like notes. That's correct — the controlled variable is length, and if the terse variant were better written I would be measuring two things at once.
Result 1: the judge is generous, not wrong
Here is the confusion matrix from a live run of deepseek-v4.1-flash against 20 human labels:
j1 j2 j3 j4 j5
h1 4 0 0 0 0
h2 1 0 0 0 0
h3 1 3 0 0 0
h4 0 0 0 0 3 <- every human-4 answer got a 5
h5 0 0 0 0 5
Raw agreement 52.9%. Unweighted Cohen's κ = 0.387 — fails the 0.6 threshold outright. But quadratic-weighted κ = 0.888, which reads as "excellent judge".
Both numbers are correct and they disagree because they answer different questions. Weighted κ treats a 4-vs-5 miss as a near-hit. Unweighted κ treats it as a miss. If your downstream use is "did this answer reach the top band", the offset is the entire signal, and the plain κ verdict is the honest one.
This is the argument for shipping both numbers next to each other rather than picking a metric and moving on.
Result 2: position bias was absent — genuinely
0 flips. Mean score delta 0.000. First-slot preference 45.5%, which is a coin flip. 100% agreement with the human preference across 8 pairs.
Worth stating plainly because a null result here is a finding: this judge's pairwise ordering was not moved by presentation order. Not every experiment has to find something.
Result 3: the one that changed how I think about this
I ran the same configuration four times. Same 20 items, same 8 pairs, same model, judge temperature 0.0 — the setting where a model is supposed to be most deterministic.
| Quantity | Run A | Run B | Run C | Run D | Threshold |
|---|---|---|---|---|---|
| Cohen's κ | 0.297 FAIL | 0.387 FAIL | 0.321 FAIL | 0.323 FAIL | ≥ 0.60 |
| verbosity delta | 0.533 FAIL | 0.200 PASS | 0.462 FAIL | 0.353 FAIL | ≤ 0.30 |
| verbosity 95% CI | [−0.067, 1.333] | [−0.467, 0.933] | [−0.154, 1.308] | [−0.177, 1.000] | all four straddle 0.30 |
| items scored in agreement pass | 18/20 | 17/20 | 19/20 | 17/20 | — |
| pairs completing both orderings | 7/8 | 4/8 | 4/8 | 3/8 | — |
The verbosity check passed once and failed three times on identical runs.
Look at the confidence intervals: each one is roughly ±0.7 wide, and the threshold is 0.3. The threshold sits inside the noise. This gold set cannot resolve that check, and no amount of re-running will make it able to.
Two things follow from that, and they apply well beyond this project.
A pass on a quantity whose interval straddles its threshold is not a pass. It is undecided, and a harness should say so. That is why every number in Arbiter's bias table is printed next to its CI rather than on its own.
Sample size is a first-class property of an eval, not a config value. Twenty items is the default in every quickstart I have ever written. For detecting a 0.3-point effect at the observed variance, twenty is the wrong number, and the fix is not more seeds — it is more items, or a wider threshold stated honestly.
I nearly averaged this away. The version of the README I started writing quoted one run and reported "verbosity bias: 0.53, judge inflates long answers", which is a real effect from a real measurement and also not stable enough to be a finding. Running it three more times is the only reason I know the difference — and note that the one passing run is the run whose full output is published, so a reader seeing only that table would conclude the judge has no verbosity problem.
War story: Pydantic AI cannot talk to the default provider
The spec pinned the default provider to an OpenAI-compatible endpoint and the framework to Pydantic AI. Both were reasonable. Together they did not work.
pydantic_ai.exceptions.UnexpectedModelBehavior: Invalid response from openai chat completions endpoint:
1 validation error for ChatCompletion
metadata.weight_versions
Input should be a valid string [type=string_type, input_value=[{'version': 'default', 'start': 0, 'end': 239}], input_type=list]
The gateway returns metadata as an object whenever tools or structured output are involved, and the OpenAI SDK declares ChatCompletion.metadata as str | None. Plain unstructured calls return
metadata: null and work fine, so the first probe succeeds and the real one fails, which is a confusing way to discover it.
Unknown extra fields are tolerated by the SDK; only this mistyped known field breaks it. The fix is a small httpx response hook that repairs the body before the SDK validates it:
async def sanitize(resp: httpx.Response) -> None:
if resp.status_code != 200: return
if "application/json" not in resp.headers.get("content-type", ""):
return # never touch SSE streams
body = await resp.aread()
try: payload = json.loads(body)
except (json.JSONDecodeError, UnicodeDecodeError): return
if not isinstance(payload, dict) or not isinstance(payload.get("metadata"), (dict, list)):
return
patched = json.dumps(sanitize_metadata(payload)).encode()
resp._content = patched
resp.headers["content-length"] = str(len(patched))
client = AsyncOpenAI(base_url=base_url, api_key=key,
http_client=httpx.AsyncClient(event_hooks={"response": [sanitize]}))
model = OpenAIChatModel(model_id, provider=OpenAIProvider(openai_client=client))
After that, structured calls run in 1.2–2.3 seconds. Every "the framework is broken" investigation deserves five minutes of reading the raw response body with curl first — the answer was in the metadata field the whole time.
The verdict is code, and it is on screen
Thresholds buried in a config file are how you end up arguing about a number six months later. They are constants, and the UI renders them next to the measurement:
KAPPA_MIN = 0.6
VERBOSITY_DELTA_MAX = 0.3
FLIP_RATE_MAX = 0.15
VERDICT: judge is NOT trustworthy for this rubric
- Cohen's kappa vs human labels >= 0.60 value 0.387 FAIL
- |padded - terse| mean score delta <= 0.30 value 0.200 PASS
- position-swap flip rate <= 15% value 0.000 PASS
! Cohen's kappa 0.387 < 0.6 — agreement with human labels is too weak to trust this judge for this rubric
All three must pass. A judge that agrees well but inflates long answers is not trustworthy for a rubric where length is a confound. And a degenerate result is a failure, never a pass: if a judge collapses every item onto one score, κ returns NaN rather than 0, because "unmeasurable" and "no agreement" are different claims and only one of them is reassuring.
What I'd add next
The obvious one, and it follows directly from Result 3: run the whole rig N times and report the verdict distribution. Right now the instability is visible only because someone ran it three times. A verdict_stability panel — "this judge passes the verbosity check in 1 of 5 runs" — is a more honest interface than a single PASS chip.
Then: a power analysis that warns when the gold set is too small to resolve its own thresholds; more probes in the same shape (authority bias, self-preference, a "nothing to grade" control that should score 1); judge ensembles with inter-judge κ; and calibration that transfers — fit an isotonic mapping from judge score to human score and report held-out error, replacing the hand-weighted correction I shipped.
Try it
git clone https://github.com/harishkotra/arbiter.git && cd arbiter
uv venv --python 3.13 .venv
uv pip install -p .venv/bin/python pydantic-ai fastapi "uvicorn[standard]" httpx python-dotenv pytest
cp .env.example .env # any OpenAI-compatible endpoint works
cd client && npm install && cd ..
npm run dev # api :3001, client :5173
npm run verify # 116 live judge calls, six checks, exits non-zero on failure
159 pytest tests, 9 client tests, no network required for either. npm run verify:offline exercises the whole pipeline through a scripted judge and labels itself a fixture on every line, because passing structural checks is not the same claim as measuring a judge.
Code & more: https://www.dailybuild.xyz/project/275-arbiter





Top comments (0)