I built an evaluation harness. I found nine bugs in it. Every single one would have made my results look better than they were.
Nine out of nine, all pointing the same way. That is not a coincidence, and I do not think it is specific to me or to my project. I think it is a structural problem with any evaluation you build for yourself.
Here is the mechanism, before any of the evidence.
Debugging is triggered by surprise
You do not audit numbers. You audit numbers that bother you.
When a result comes back disappointing, you go looking for the reason. You check the setup, you re-run it, you add logging, you find the bug. The bug gets fixed and the number moves.
When a result comes back good, none of that fires. Nothing feels wrong. There is no surprise to investigate. You write it up.
So the filter that removes measurement bugs from your work is applied unevenly. Hard against results you dislike. Softly against results you like. Every pass through that filter removes more unflattering bugs than flattering ones.
Run that loop for a few days and your instrument has drifted in one direction, and nothing in your process is designed to notice. The bugs that survive to publication are disproportionately the ones that helped you. Not because anyone was dishonest. Because they never triggered the thing that catches bugs.
I knew this argument in the abstract before I started. It did not stop me from writing nine of them.
What the harness did
Only enough context for the bugs to make sense.
Mutation testing changes your source code in small ways. Flip a comparison. Change a constant. Delete a raise. Then it runs your test suite and checks whether anything failed.
def withdraw(balance, amount):
- if amount <= 0:
+ if amount < 0:
raise ValueError("amount must be positive")
If the suite stays green, that is a fault your tests cannot detect.
Coverage tells you a line ran. This tells you whether anything would have complained if the line were wrong. Those are very different questions. A toy module with one happy-path test sits at 47% line coverage and a 9.5% mutation kill score.
The harness generated mutations, ran suites against them, and scored how many got caught. Roughly 40 hours of work.
The nine bugs, and which way each one pushed
Editable installs made mutations invisible.
pip install -eon src-layout packages resolved imports back to the original checkout, so mutations written to a temp copy never executed. Three targets scored 0.000. Direction: reads as "these test suites are terrible" rather than "my harness is broken." It made the problem I was solving look bigger.Parallel execution corrupted one target. Running mutants concurrently produced three different results across four runs on the one target doing real async I/O. A recorded improvement of 0.27 to 0.77 was noise. Spurious failures get counted as the mutation being detected, and detection was the number I was maximising. Direction: inflated the result.
A file picker chose an unrelated test file for the hardest target, feeding the model irrelevant context exactly where context mattered most. Direction: made a baseline look worse than it was.
A classifier categorised batches instead of individual tests. One good test in a batch of 69 would have marked all 69 as good. Direction: inflated quality.
A reconstruction step dropped shared imports and manufactured test failures that were not real. Direction: understated a baseline's capability.
An extractor only scanned top-level functions, so a valid
unittest.TestCaseresponse was discarded as "no test found." The retry loop then received a harness error instead of real pytest output, which disabled the exact mechanism I was measuring. Direction: understated the agent.self.assertEqual(...)was classified as "no assertion." I was testing a hypothesis about models writing assertion-free tests at the time. Direction: would have manufactured my own hypothesis and handed it back to me.A pre-registered metric was not computable on dunder-dispatched code. It read as a real near-zero rate instead of as undefined. Direction: false signal.
A documentation figure that had already survived two audits. One arm was recorded as having zero clean-pass failures when it had three. Direction: flattered a comparison.
Note that they do not all inflate the headline number. Three of them understate a baseline or an arm. That still counts as favourable, because a worse baseline makes the thing I built look better by comparison. "Favourable" means favourable to the story, not favourable to one metric.
Bug 9 is the one I find hardest to be relaxed about. It had been looked at twice. Two audits, both of which read past it, because the figure was consistent with what we expected to see and nothing about it invited a third look.
None of them was found by reading code
This is the part I would most want someone to take away.
Not one of those nine was caught by re-reading the function. I had already read the functions. Reading code that you wrote, looking for a bug you do not yet believe exists, is close to useless.
Every single one was caught the same way: by running a check whose outcome I had predicted in advance, and getting a different answer.
The clearest case was bug 5. I knew, independently and from earlier output, that one specific generated test was broken. So the prediction was simple. Remove that one test, and the suite goes green.
I removed it. The suite did not go green.
That contradiction is the only reason I found the dropped-imports bug before its numbers went into anything. There was no other signal. The scores it produced looked entirely plausible. They were plausible in a direction I liked, which is why nothing else would have prompted me to look.
The prediction is what does the work. A check you run without a prediction just produces another number, and you will interpret that number the same way you interpret all the others.
Two checks worth stealing
Both are cheap. Both caught things.
The canary. Overwrite the file under test with unparseable garbage and assert that the suite fails. This proves your mutation actually reaches the interpreter. If the suite passes while the file is syntactically invalid, you are not testing what you think you are testing. This is what would have caught bug 1 on day one.
The determinism gate. Run the same scoring three times, serially, and require byte-identical output. This proves execution is isolated. This is what catches bug 2.
The important part is that these prove different properties, and neither substitutes for the other.
The canary passes happily while concurrency silently corrupts your results. Your mutation reached the interpreter, so the canary is satisfied, and the numbers are still garbage.
The determinism gate passes happily while imports resolve to the wrong file. Three identical runs of the wrong thing are still perfectly deterministic. The gate is satisfied and the score is meaningless.
You need both, and you need to write down what each one actually proves, so you do not talk yourself into believing one covers the other.
The check that caught one of my own targets
Three hours before my deadline I did a clean-clone reproduction run to verify the reproducibility claim.
The determinism gate quarantined one of the project's own targets. Eleven of twelve reproduced exactly. The twelfth varied.
I reported it in the README instead of fixing it.
A check that has never caught anything is indistinguishable from a check that cannot catch anything. Mine had just caught something, and the something was mine. Removing that from the record would have made the project look better and the instrument look worse, which is exactly the trade this entire post is about.
What to actually do
Before you measure anything you built:
Write down what your instrument would look like if it were lying to you. Then build the check that catches specifically that, and run it before you have any results you are attached to.
Then write down which direction each possible lie would push your result.
That second list is the one that matters, because it is a list of the checks you will be least motivated to run. Every item on it is a place where a bug will feel like a finding. You will not notice those on your own. Nobody does. That is the whole reason the nine came out nine for nine.
Top comments (15)
Your canary and the determinism gate prove two properties, and there is a third one both of them pass while it is broken: that the mutation reached the interpreter as bytecode, not just the disk. CPython keys pyc invalidation on source mtime in whole seconds plus source size, so a mutant flipping
<to>is byte-size identical and, written in the same second as the pyc, gets skipped — on 3.14.6 the mutated file is on disk and the import still returns the original value, exit 0, nothing printed. The canary misses it because unparseable garbage is a different length, so the pyc really does get invalidated and the suite fails exactly as expected; I rewrote that same canary to be byte-size-preserving and it started firing. The gate misses it too, since staleness is deterministic, and three serial runs came back byte-identical. What makes this the default rather than a race isshutil.copytree: it preserves mtime and copies__pycache__along with the source, so the temp tree inherits a matching pyc before the first mutant is written.PYTHONDONTWRITEBYTECODE=1stops writing rather than reading and changed nothing, whilecompileall --invalidation-mode checked-hashfixed it. Direction fits your nine, since the mutant survives, the kill score drops, and the suites look worse than they are. I measured this on CPython 3.14.6 on macOS with plain imports rather than pytest, and I have not run your harness.You were right about the mechanism, and I've now confirmed it. Compile a pyc from a == b, overwrite it with a byte-identical a != b, force the mtime back - the stale bytecode executes. Reproduced on the first try.
It doesn't reach any scored result in my harness, and the reason is embarrassing: all four temp-copy call sites pass ignore_patterns("pycache", "*.pyc", ...). I wrote that line to avoid copying junk. It turns out to be load-bearing for correctness, and I didn't know that until you pointed out the mechanism. I've verified it by copying a target with 19 real pyc files and confirming 0 arrive. There's a second accidental layer too - ast.unparse reformats, so 0 of my 455 mutants are byte-length identical to the original anyway - but that's an artifact of the unparser, and I'm not treating it as protection.
The canary point is the part that stings, and it stands regardless. Unparseable garbage changes the length, so the pyc is genuinely invalidated, so the canary fires and tells me mutations reach the interpreter. It was answering a question I wasn't asking. I've added a byte-size-preserving canary alongside it - 8 of 12 targets are eligible, all pass, and the other 4 are reported N/A rather than silently skipped. Plus a regression test that I checked actually fails against a deliberately broken variant before trusting it.
On direction: agreed, and it's the same shape as bug 1. The mutant survives, the score drops, the suite looks worse than it is, and the problem the tool exists to solve looks bigger. That's the version of flattering that's hardest to notice, because the number that moves isn't the one you're proud of.
The thing I'd have missed entirely: a line added for tidiness was the only thing standing between me and a tenth bug. Someone cleaning that up later would have had no idea. It's now commented as load-bearing, with a test.
One boundary on that ignore_patterns line, now that it is marked load-bearing: it only holds while bytecode lands next to the source. Set
PYTHONPYCACHEPREFIXand the pyc goes to a central tree keyed on the absolute path of the copy, so your 19-file test still sees 0 arrive while the stale read comes straight back. I ran three branches on 3.14.6, same copytree with ignore_patterns each round, mutant flipping<to>at identical byte length with mtime forced back: no prefix plus a fixed work path prints False (fresh), prefix plus that same fixed work path prints True (stale, and still 0 pyc inside the copy), prefix plus a unique work dir per mutant prints False again. So it needs the work dir to be reused, which a fixed scratch dir does and tempfile.mkdtemp does not. Your byte-size-preserving canary does catch this one, provided it runs through the same copy path the real mutants take, and that is the property I would assert on rather than the pyc count. Plain import on macOS, not through your harness.Reproduced all three branches on 3.14.7 and got exactly what you described: no prefix fresh; prefix plus a reused fixed path stale with 0 pyc in the copy; prefix plus a unique path per call fresh again.
So the ignore_patterns line I'd just labelled load-bearing was only ever half the protection. The other half is that every copy site uses tempfile.TemporaryDirectory() - a unique path per mutant, which is your third branch. I checked all of them: _evaluate_one, verify_clean, the gate check, the batch rescore including its nested per-mutant copy, the baseline arm, and both verification scripts. None reuses a fixed path, and a full 12-target run with PYTHONPYCACHEPREFIX set diffs byte-identical against the committed verification file. So it doesn't bite here either.
But that's now two accidental protections in the same area, and neither was written for this reason. I wrote ignore_patterns to avoid copying junk and TemporaryDirectory because it cleans itself up. Somebody adding a --work-dir flag or a copy-caching optimisation would remove the second one and have no way to know. It's commented now, naming both halves and the env var explicitly.
The assertion point is the part I'd have missed, and it's the more general lesson. My regression test asserted pyc_count == 0, which is precisely the assertion that holds while the harness is stale under a prefix. A proxy that's true under my configuration and false under a supported environment variable. The test now runs a byte-size-preserving mutant through the real _evaluate_one under PYTHONPYCACHEPREFIX and asserts the observed outcome isn't survived - and there's a meta-test that reproduces your stale branch to prove the check actually has teeth, because a regression test nobody has seen fail is in the same category as the canary that couldn't detect this in the first place.
Assert the property, not the proxy. Same cost to write, doesn't have that failure mode. That generalises well past mutation testing, and I'm going to be applying it to a few other things.
Worth naming what's happened across both your comments: you've pointed at my checks twice, been right twice, and neither time did a result move. The scrutiny has landed entirely on the instrument. I think that's the actual argument for publishing the harness rather than the finding, and I hadn't seen it that clearly before you started poking at it. Thank you - this has been the most useful exchange I've had about this work.
Following up since you found two of these. Both fixes shipped, and the repo is public now: github.com/marvinoka4/killcheck
The byte-size-preserving canary is in, running on 8 of 12 targets, with the other 4 reporting N/A rather than silently skipping. The regression test asserts the property (observed behaviour changed) instead of the pyc count, with a meta-test reproducing your stale branch so the check is known to fail when it should. Both accidental protections, ignore_patterns and TemporaryDirectory, are commented as load-bearing and name PYTHONPYCACHEPREFIX explicitly.
Wrote the whole project up short, if useful: dev.to/marvinoka4/69-tests-all-pas...
"Debugging is triggered by surprise" is the most useful sentence I have read about evaluation this year, because it explains a one-directional bias without anyone being dishonest.
Two counter-measures that helped me: pre-register the number you expect before you run, so that a result matching your prediction still has to be audited (agreement with a guess is not evidence the instrument is sound). And keep a deliberate negative control in the suite, a query set where the score should be near zero. When the harness scores that too high you have found a flattering bug you would never have gone looking for, because nothing about it felt surprising.
The negative control is the better of your two suggestions, and I hadn't thought of it in those terms.
I was doing pre-registration, and it still let most of the nine through. Writing the expected number down helps when you're wrong, but it does nothing when the instrument is broken in a way that produces the number you predicted. Bug 7 is exactly that case: I predicted models would write assertion-free tests, and a classifier that misread self.assertEqual would have handed me my prediction with a confirmation attached.
The negative control is different because it inverts which direction is alarming. Everywhere else in the harness, high is good and low prompts investigation. On a set where the score should be near zero, high is the alarm. That gives you one place where the flattering failure is the surprising one, which is the only thing that reliably triggers debugging.
The design constraint I'd add is that the control has to be something your harness cannot accidentally pass for a boring reason. If it scores near zero because the target never loaded, you've built a check that's satisfied by the exact class of bug you're hunting. Which is what bug 1 did to me: three targets at 0.000 that read as a finding rather than a fault.
Your negative control is in, and it earned its place immediately. Repo is public: github.com/marvinoka4/killcheck
It came back at 7 of 51 rather than near zero, and the explanation turned out to be a real calibration fact rather than a bug. assert x is not None detects exactly one class of fault; all seven kills were that one operator, zero were anything else. A control that returned clean would have told me nothing. It also has the guard you'd expect, asserting separately that the suite passes on clean source and that the module actually imported, so it can't pass for the boring reason.
Short write-up of the whole thing: dev.to/marvinoka4/69-tests-all-pas...
Still thinking about your point on pre-registration. Writing the number down did nothing for the bugs that produced the number I predicted.
7 of 51 with all seven from one operator is a better outcome than a clean zero, for exactly the reason you give: a control that returns nothing tells you the control is insensitive, not that the suite is sound. What you actually measured is the suite's coverage profile, which is a more useful artifact than a pass/fail. It also suggests the complementary control: seed mutations the suite should catch and confirm it does, so you have the sensitivity curve from both ends rather than only the floor.
On pre-registration, that is a fair correction and I overstated it. Writing the number down only helps against the bias where a disappointing result triggers investigation and a pleasing one does not. It does nothing for a bug that produces the number you predicted, and yours were mostly that second kind. The honest version is that pre-registration narrows the surprise-triggered filter and leaves everything else untouched, which is a much smaller claim than the one I made.
Asserting separately that the suite passes on clean source and that the module actually imported is the more load-bearing of your two guards.
Some comments may only be visible to logged-in visitors. Sign in to view all comments.