Run the same LLM-as-judge eval twice and you can get pass, then fail, on identical input. Now try to build anything on top of that verdict. This is...
For further actions, you may consider blocking this person and/or reporting abuse
Two things to add, because your closing question has a cheaper answer than "repeat more".
How many repeats you need is only a number once you have the flip rate. Under the strict-majority rule as written, the chance the majority is wrong is a binomial tail: at a 10% per-run flip rate, 3 runs give 2.8% and 5 runs 0.86%; at 20% you need 7 runs to get under 5% and 13 for 1%; at 30% it is 17 and 31. So repeat-and-vote is nearly free while the judge is stable, and goes vertical exactly in the regime where you can least afford it.
Which suggests the right move at high flip rates is not more N but a third outcome. If a mutant's per-run verdicts straddle 50%, you have not earned a binary killed/survived - you have an unresolved mutant. Carrying
unresolvedas first-class (excluded from the score numerator and denominator, reported as its own rate, CI on the reduced set) costs nothing extra and is honest about what was actually measured. Worth noting the tie rule pushes the other way:killed = fails * 2 > len(runs)(src/muteval/runner.py:382) means ties survive, so an even N and a genuinely undecided judge both default to "survived". For a gate whose job is to block regressions that asymmetry points the wrong way - an uninformative judge at N=2 produces a kill 25% of the time. Escalating N on a tie, and blocking when a high-severity mutant is still tied after escalation, gets you the severity stratification and the adaptive N in one move.Last thing: the judge is not the only thing that flips between runs. We re-ran a 60-cell benchmark from a fresh clone on a second machine - same code, same seeds, same pinned versions - and 42/60 cells matched exactly while the heavy cells drifted by up to about 10. That drift is a systematic offset for the whole session, so unlike judge noise it does not average out with N; every repeat inherits it. The practical split that survived contact with the second machine was to run the control arm inside the same invocation so the comparison is paired, then assert invariants exactly (this region spikes, that one does not) and magnitudes inside a band with the tolerance written next to the number.
Checked the line before answering since you cited it: runner.py:382 is exactly killed = fails * 2 > len(runs), ties survive. You're right on all three counts.
The binomial tail reframes "repeat more" as only honest in the stable regime — nearly free at a 10% flip rate, vertical exactly where you can least afford it. Which means a fixed N is the wrong knob; it should be a function of the measured flip rate, not a constant.
The third outcome is the real fix, and it's the strongest point. A mutant whose per-run verdicts straddle 50% hasn't earned a binary verdict — forcing it into killed/survived is manufacturing a result out of noise. Carrying unresolved as first-class — out of both numerator and denominator, reported as its own rate, CI on the reduced set — is honest about what was actually measured and costs nothing. muteval doesn't have it; it should. And you've caught that the tie rule makes it worse than neutral: with ties surviving, an undecided judge at N=2 still resolves to a definite verdict a quarter of the time, and for a gate you never want an undecided high-severity mutant silently defaulting either way. The move you describe — escalate N on a straddle, block a high-severity mutant that's still unresolved after escalation — quietly delivers two things other readers asked for in one mechanism: adaptive N (spend passes only where the verdict is close) and severity stratification (the block decision is per-tier). unresolved + escalate-on-tie is going on the list as a single change.
Session drift is the part I hadn't priced in. You're right that judge flip isn't the only source, and a per-session systematic offset is worse because it doesn't average out with N — every repeat inherits it. The one thing in muteval's favor: it scores a differential, not an absolute. The verdict is whether the eval flips baseline→mutant, and both are graded in the same invocation, so a session offset that hits both sides largely cancels in the contrast that gets scored — in a way it wouldn't if the baseline came from another machine. That's not a free pass, though: the individual eval verdicts are still absolute, so a judge sitting near its threshold can flip both, and a user eval that asserts an exact magnitude stays brittle no matter how muteval votes. Your discipline — assert invariants exactly, magnitudes in a band with the tolerance written next to the number — is the fix on the eval-authoring side. muteval can keep the comparison paired; it can't make a brittle assertion robust. That belongs in the docs.
Thanks — the middle point in particular ties together a couple of the other threads on this post.
Went and measured the pairing claim, because "largely cancels" turns out to have a boundary worth writing down.
I modelled it rather than ran it, since I don't have your judge: cases with a latent score, binary verdicts against a threshold, per-call noise sd 0.03 on a 0-1 scale with the threshold at 0.5. Read the ratios, not the absolutes. Pairing does buy the level - at a session offset sd of 0.10, which is the row that matters given the ~10-point cross-machine drift I measured on heavy cells, the error in the estimated flip rate is 2.2pp paired against 12.1pp unpaired, with a 1.4pp floor from per-call noise alone. So the point holds, and it's worth about 5x.
But the cancellation is a property of the score, not of the verdict. In the same paired runs 17% of individual case verdicts differed from what a zero-offset session would have produced, and that is entirely concentrated by distance to the boundary: 0.01% for margins above 0.20, 1.6% at 0.10-0.20, 12.7% at 0.05-0.10, and 47% below 0.05. Which is the awkward part, because near-neutral mutants are exactly the ones whose flip question is real: a case whose true delta is ~0 is decided by the offset, and pairing cannot help there, since both sides shift together but each is still compared to the threshold separately. Pairing cancels the offset in the average, and the average is the least informative line in the report.
That gives unresolved a sharper definition plus one cheap measurement. Definition: unresolved = margin smaller than the measured session offset, rather than "per-run verdicts straddle 50%". Measurement: a null-mutant control - run the same eval twice in one invocation with a no-op mutation for the within-session offset, and the same pair across machines for the cross-session one. If that offset is a fraction of the typical margin, escalation-on-straddle fixes the regime, because extra passes average down the per-call noise. If it exceeds the typical margin, more passes cannot fix it, because every pass inherits it. Same conclusion reached from the other end, and the null-mutant number tells you which regime you are in before you spend the passes.
Agreed on the split in responsibility: muteval keeps the comparison paired, the eval author keeps the assertion inside a band. The doc line I'd want is that the contrast is paired and the constituents are not - so report the margin distribution next to the flip rate, since the runner already has both numbers.
This is the version of the argument I wanted and didn't have — thanks for actually modelling it.
The ratio is the thing: ~2.2pp paired vs ~12.1pp unpaired at a 0.10 offset, ~5x, with a per-call-noise floor you can't pair away. That matches the intuition that pairing buys the level and nothing at the boundary. And your sharper point lands — the cancellation is a property of the score, not the verdict. 17% of case verdicts flipping, concentrated below a 0.05 margin, is exactly the population whose flip question is real. "Pairing cancels the offset in the average, and the average is the least informative line in the report" is the sentence I'll be quoting back to myself. It's right, and it's the honest limit of the paired design.
Two things I'm taking as concrete:
Redefining unresolved as margin smaller than the measured session offset, not "per-run verdicts straddle 50%." That ties undecidedness to a measured noise floor instead of a fixed band — the only version that's honest about near-boundary mutants.
The null-mutant control as the cheap measurement that tells you the regime before you spend passes: run the eval twice in one invocation under a no-op mutation for the within-session offset, the same pair across machines for the cross-session one. If the offset is a fraction of the typical margin, escalate-on-straddle works because passes average down per-call noise; if it exceeds the typical margin, no N helps. That's the decision rule I was missing — measure the floor first, then decide whether repeating is even the right lever. It also composes with the canary/inert-detection machinery already there rather than fighting it.
And the doc line writes itself: the contrast is paired, the constituents aren't — so report the margin distribution next to the flip rate, since the runner already carries both numbers. Cheapest honest change on the list.
One of the most useful threads I've had on any of these posts. If you're up for it, this would make a great issue on the repo — happy to credit the analysis.
If you take that rule, test its form first, because I tried it and the obvious version under-counts - and it under-counts by the cases it exists to flag.
The mechanism: the offset moves the margin and the band together. A case whose true margin is small but non-zero, same sign as the session offset, gets carried further out of the band - so it reads as resolved even though a session with a different offset would have called it the other way. Measured against the population it is trying to describe (share of cases whose true margin is under the session offset), with offset sd 0.10 and per-call noise sd 0.05, minimum 400 trials per cell:
So: N does not fix it (the first row is flat), and the two bands built from a standard error get worse as you collect data - they become more confident about a quantity that is beside the point, and end up declaring resolved the cases the offset decides. The magnitude-anchored form tracks the target at every precision and tightens toward it from below.
The part that changes your decision rule: fix the anchor before spending passes. With a magnitude band, escalation sharpens the answer; with an error band, escalation makes the report look better while more cases are silently settled by the offset. That is the one place I would not trust the null control on its own.
Punchline from the other direction: with an offset at the measured drift magnitude, even a perfect precision budget leaves roughly 18% of verdicts wrong relative to a zero-offset session, against a 4-5% floor at zero offset. The control tells you which regime you are in; it does not get you out of it.
On the issue: yes, please. I would put in the four-rule table above as the concrete content - the magnitude-anchored definition of unresolved, the note that escalation and a precise control do not rescue a mis-anchored band, and the margin-distribution-next-to-flip-rate line. Credit the analysis however you think is accurate; if you draft it I will read it before it goes up and say so if anything in it overstates what the numbers support.
You're right, and this is a correction, not a refinement — the rule I accepted under-counts, and worse, it under-counts exactly the offset-decided cases it exists to catch. The mechanism follows: the offset shifts the raw margin and the band together, so a small-but-nonzero true margin with the same sign as the offset gets pushed further outside the band and reads as resolved — when a session with the opposite offset would have called it the other way. |m_raw| < |o| is blind to precisely that population, which is why it sits flat near 28% against a 37% target no matter how many passes you throw at it.
The reason the two standard-error bands get worse with data is the part I'd have gotten wrong: they answer "is the de-offsetted margin distinguishable from zero given sampling noise," and as c and k grow that SE collapses, the band narrows, and more cases get declared resolved — a sharper and sharper answer to the wrong question. The quantity that decides the case is the offset magnitude, not the sampling precision, so the band has to be anchored to |o|, not noise/√c. |m − o| < |o| is the only one of the four that tracks the target at every precision and tightens toward it from below, because it estimates the right comparison and just gets less noisy about it.
So the decision-rule correction is the real payload, and it inverts what I said: fix the anchor before spending passes. Under a magnitude band, escalation sharpens the classification; under an error band, escalation makes the report look cleaner while quietly settling more cases by the offset. "Escalate on a straddle" is only safe once the band is magnitude-anchored — with an SE band it's actively misleading. That's the one place the null control can't be trusted on its own: it hands you o, but plugged into the wrong band it makes things worse.
And the punchline is the honest ceiling — at an offset the size of the drift I measured, even a perfect precision budget leaves ~18% of verdicts wrong against a 4–5% zero-offset floor. The control classifies the regime; it doesn't get you out of it. That goes in the docs next to the definition, because it's the line that stops someone reading a clean re-run as a clean result.
On the issue — yes, and I'll reproduce the four-cell sim myself before it goes up rather than transcribe your numbers, so the table in the repo is one I've re-run (same discipline the tool preaches). Content as you laid it out: the four-rule table, the magnitude-anchored definition of unresolved, the note that neither escalation nor a precise control rescues a mis-anchored band, and the margin-distribution-next-to-flip-rate reporting line. I'll credit you by handle — say if you'd prefer a name or link — and I'll drop the draft here for you to check before it's published; hold me to "nothing in it overstates what the numbers support."
I would also segment stability by consequence. A judge can look reliable across the full mutation set while remaining inconsistent on the small class of payment, permission, or compliance failures that matter most.
Do you calculate mutation survival and confidence intervals per failure class or risk tier? That would separate harmless variance from instability that should block acceptance.
This is the sharper version of the stability question, and honestly — no, muteval computes the score and the Wilson interval globally, not per risk tier. It has the ingredients but not this: every mutant carries a severity, critical-pattern text (refund / PII / permission / compliance) escalates it, and there's a --fail-on-severity gate that blocks on any surviving high-severity mutant. So risk is tagged per mutant and gated — but the aggregate stability signal (the judge flip-rate, the CI) is pooled across everything.
And you've named exactly why that isn't enough: a global 8% flip-rate could be the judge being rock-solid on the easy 90% and genuinely unsure precisely on the payment/permission class — the one place you'd want it to block. Pooling hides the instability that matters most. Stratifying survival + CI + flip-rate by severity is the honest fix, and it's a clean extension of machinery that's already there. Going on the list.
What's nice is it converges with the cost point in the other thread: risk-tiering is also a budget lever — spend your (adaptive) N passes on the high-consequence mutants where you need the stability, single-pass the rest. "Judge harder where it's expensive to be wrong" turns out to be both the cheaper policy and the more trustworthy one.
Yes. The stratified view also gives you a cleaner stopping rule than a global target. For low-consequence mutants, one pass may be enough unless the result is borderline. For high-consequence mutants, keep sampling until the interval clears a predeclared accept-or-block boundary; if it never clears within the budget, abstention is the result. I’d also report how often severity escalation came from the text pattern versus the mutant’s declared class. Otherwise judge uncertainty can leak into both routing and evaluation. The useful artifact becomes a risk-tiered decision table: passes used, interval, final disposition, and whether a human override was required.
Circling back: your risk-tier point didn't just go "on the list" — it merged with two other threads on this post into one change. The shape is severity-stratified stability: survival + Wilson CI + flip-rate reported per tier, so a global 8% flip-rate can't hide a judge that's solid on the easy 90% and shaky exactly on the payment/permission class. And it converges with the cost thread — risk-tiering is also a budget lever: spend the (adaptive) passes where being wrong is expensive, single-pass the rest. Cheaper and more trustworthy at once. Thanks for the push; it was one of the sharper framings here.
Reporting a Wilson interval instead of a point percentage is the right discipline. If a judge has even a 10% flip rate, a single green check in CI is just an unpriced lottery ticket on prompt variance.
The bottleneck usually isn't statistical theory; it's the invoice. Running an N-pass majority vote across hundreds of mutants turns a linear eval budget into a painful fixed cost on every pull request. A squad shipping forty PRs a day with five model passes per mutant burns through their token allocation before lunch.
That cost convexity explains why so many teams retreat back to single-pass judges. Engineers understand the coin has an unknown bias, but the budget forces them to treat a lucky flip as deterministic verification. Short-circuiting with deterministic rule checks upstream is the only way to keep the expected cost from pricing the entire test suite out of CI.
Exactly — the invoice is the real constraint, and it's worse than linear because it's the product of mutants × cases × runs × judges. Five passes across a few hundred mutants per PR is how a linear eval budget becomes a fixed tax on every merge.
muteval leans on the levers you'd expect, and I'd call them load-bearing, not nice-to-haves: cheap rule-based checks run before the judge and short-circuit (a mutant a deterministic guardrail already kills never reaches a paid pass); an inert mutant whose output is byte-identical to the baseline reuses the baseline's outcomes for zero judge calls; an identical re-run is fully cached; and --max-calls fails closed before you overspend, not after. On a suite with real deterministic guardrails, most mutants die cheap and the judge only ever sees the survivors — which is exactly the upstream short-circuit you're describing.
But here's the honest gap, and I think it's where the cost curve actually bends: muteval runs a fixed N passes per mutant. You don't need five passes on a mutant whose first two agree 2–0 — you need them on the mutants sitting near the decision boundary. Adaptive N — stop early when the passes agree, escalate only when they don't (sequential testing) — is the right answer to the invoice, and muteval doesn't do it yet. Fixed-N is the naive version. That's going on the list; framing it as cost convexity is the clarifying lens.
We repeat 3x and majority-vote too, but the flaky list ended up being the more useful output. Every mutant that flipped traced back to an ambiguous line in the rubric, and tightening the wording cut the flip rate more than going from 3 to 5 runs did. One thing I'd add: keep the deterministic checks keyed on the mutation operator, not just the diff text. We had mutants that only changed an error message and a text-diff check killed them for the wrong reason, which inflated the kill rate. Do you re-run the flaky-flagged mutants with a larger N before deciding, or treat the flag itself as the verdict?
Same experience here, and it's why the flaky list is a first-class output now rather than a footnote: muteval attributes flips per eval (flaky_by_eval), so it points at the rubric dimension doing the flipping — "dimension X caused 8 of your 11 flips" — which turns "add runs" into "rewrite that line." Exactly your finding: tightening the ambiguous wording beats 3→5, because extra runs just average down noise you should've removed at the source.
The wrong-reason kill is a real one, and I like the operator-keying fix. muteval records which eval killed each mutant and under which operator (failing_eval/caught_by + operator + severity), so an error-message-only mutant killed by a text-diff check shows up in the report instead of silently inflating the rate — but you're right that surfacing it isn't preventing it. Keying a deterministic check to the operator (this check is only meaningful for these mutation classes) is the cleaner discipline, and it fits how operators already carry severity. Good gap to name.
On your question: today it's fixed N (runs_per_mutant) with a majority vote as the verdict, and flaky is a separate reliability flag — not the verdict itself. It does not re-run flagged mutants at higher N yet. That adaptive step — escalate N only on the mutants that flip, stop early on the ones that already agree — is the thing I keep circling back to; fixed N over-samples the easy mutants and under-samples the boundary ones. It's on the list as sequential testing. For now: majority is the call, flaky tells you not to trust that call, and a genuinely tied verdict is carried as unresolved rather than defaulted either way.
The flaky list being signal is underappreciated - verdicts do not flip uniformly. They cluster on the mutants where the rubric is ambiguous, so the flaky list is really a bug report about the eval question, not just about the judge. If half your flips come from one scoring dimension, the dimension needs rewriting before the judge needs more samples.
Second gotcha beyond re-runs: judge drift over time. A majority vote stabilizes a coin flip, but a model version bump quietly replaces the coin. Pin the judge version in every report, or last month's killed-rate and this month's are measuring different things.
Both of these are provenance gaps, and both are things muteval doesn't do yet.
On the flaky list being a bug report about the eval question — yes, and it's the better diagnosis. flaky is per-mutant today (which mutants flipped), not per-eval. But the data to do what you're describing is already there per run: every run records which eval fired, so aggregating flips by the eval that caused them would say "dimension X accounts for 8 of your 11 flips" — which turns "add samples" into "rewrite that rubric dimension," and rewriting is cheaper and more correct. That's the whole ethos of the thing (fix the eval, not the number), and it's a reporting change, not new measurement. Going on the list.
On judge drift — you're right, and it's the more dangerous one because it's silent. muteval pins nothing about the judge in the result today, so last month's killed-rate and this month's genuinely aren't comparable across a model bump. One honest wrinkle: muteval's evals are opaque callables (output, case) -> verdict, so it can't always introspect the judge — a user's llm_judge or a deepeval metric hides the model inside the closure. What it can do is record the model-under-test and its own judge's model, and carry a judge-version field in the report regardless, so a comparison across time either lines up or visibly doesn't.
Both come down to attribution: flakiness to the dimension, the score to the judge that produced it. Neither is measurement I'm missing — it's provenance I'm not recording. That's a cheap, honest gap to close.