I built an evaluation harness. I found nine bugs in it. Every single one would have made my results look better than they were.
Nine out of nine, ...
For further actions, you may consider blocking this person and/or reporting abuse
Your canary and the determinism gate prove two properties, and there is a third one both of them pass while it is broken: that the mutation reached the interpreter as bytecode, not just the disk. CPython keys pyc invalidation on source mtime in whole seconds plus source size, so a mutant flipping
<to>is byte-size identical and, written in the same second as the pyc, gets skipped — on 3.14.6 the mutated file is on disk and the import still returns the original value, exit 0, nothing printed. The canary misses it because unparseable garbage is a different length, so the pyc really does get invalidated and the suite fails exactly as expected; I rewrote that same canary to be byte-size-preserving and it started firing. The gate misses it too, since staleness is deterministic, and three serial runs came back byte-identical. What makes this the default rather than a race isshutil.copytree: it preserves mtime and copies__pycache__along with the source, so the temp tree inherits a matching pyc before the first mutant is written.PYTHONDONTWRITEBYTECODE=1stops writing rather than reading and changed nothing, whilecompileall --invalidation-mode checked-hashfixed it. Direction fits your nine, since the mutant survives, the kill score drops, and the suites look worse than they are. I measured this on CPython 3.14.6 on macOS with plain imports rather than pytest, and I have not run your harness.You were right about the mechanism, and I've now confirmed it. Compile a pyc from a == b, overwrite it with a byte-identical a != b, force the mtime back - the stale bytecode executes. Reproduced on the first try.
It doesn't reach any scored result in my harness, and the reason is embarrassing: all four temp-copy call sites pass ignore_patterns("pycache", "*.pyc", ...). I wrote that line to avoid copying junk. It turns out to be load-bearing for correctness, and I didn't know that until you pointed out the mechanism. I've verified it by copying a target with 19 real pyc files and confirming 0 arrive. There's a second accidental layer too - ast.unparse reformats, so 0 of my 455 mutants are byte-length identical to the original anyway - but that's an artifact of the unparser, and I'm not treating it as protection.
The canary point is the part that stings, and it stands regardless. Unparseable garbage changes the length, so the pyc is genuinely invalidated, so the canary fires and tells me mutations reach the interpreter. It was answering a question I wasn't asking. I've added a byte-size-preserving canary alongside it - 8 of 12 targets are eligible, all pass, and the other 4 are reported N/A rather than silently skipped. Plus a regression test that I checked actually fails against a deliberately broken variant before trusting it.
On direction: agreed, and it's the same shape as bug 1. The mutant survives, the score drops, the suite looks worse than it is, and the problem the tool exists to solve looks bigger. That's the version of flattering that's hardest to notice, because the number that moves isn't the one you're proud of.
The thing I'd have missed entirely: a line added for tidiness was the only thing standing between me and a tenth bug. Someone cleaning that up later would have had no idea. It's now commented as load-bearing, with a test.
One boundary on that ignore_patterns line, now that it is marked load-bearing: it only holds while bytecode lands next to the source. Set
PYTHONPYCACHEPREFIXand the pyc goes to a central tree keyed on the absolute path of the copy, so your 19-file test still sees 0 arrive while the stale read comes straight back. I ran three branches on 3.14.6, same copytree with ignore_patterns each round, mutant flipping<to>at identical byte length with mtime forced back: no prefix plus a fixed work path prints False (fresh), prefix plus that same fixed work path prints True (stale, and still 0 pyc inside the copy), prefix plus a unique work dir per mutant prints False again. So it needs the work dir to be reused, which a fixed scratch dir does and tempfile.mkdtemp does not. Your byte-size-preserving canary does catch this one, provided it runs through the same copy path the real mutants take, and that is the property I would assert on rather than the pyc count. Plain import on macOS, not through your harness.Reproduced all three branches on 3.14.7 and got exactly what you described: no prefix fresh; prefix plus a reused fixed path stale with 0 pyc in the copy; prefix plus a unique path per call fresh again.
So the ignore_patterns line I'd just labelled load-bearing was only ever half the protection. The other half is that every copy site uses tempfile.TemporaryDirectory() - a unique path per mutant, which is your third branch. I checked all of them: _evaluate_one, verify_clean, the gate check, the batch rescore including its nested per-mutant copy, the baseline arm, and both verification scripts. None reuses a fixed path, and a full 12-target run with PYTHONPYCACHEPREFIX set diffs byte-identical against the committed verification file. So it doesn't bite here either.
But that's now two accidental protections in the same area, and neither was written for this reason. I wrote ignore_patterns to avoid copying junk and TemporaryDirectory because it cleans itself up. Somebody adding a --work-dir flag or a copy-caching optimisation would remove the second one and have no way to know. It's commented now, naming both halves and the env var explicitly.
The assertion point is the part I'd have missed, and it's the more general lesson. My regression test asserted pyc_count == 0, which is precisely the assertion that holds while the harness is stale under a prefix. A proxy that's true under my configuration and false under a supported environment variable. The test now runs a byte-size-preserving mutant through the real _evaluate_one under PYTHONPYCACHEPREFIX and asserts the observed outcome isn't survived - and there's a meta-test that reproduces your stale branch to prove the check actually has teeth, because a regression test nobody has seen fail is in the same category as the canary that couldn't detect this in the first place.
Assert the property, not the proxy. Same cost to write, doesn't have that failure mode. That generalises well past mutation testing, and I'm going to be applying it to a few other things.
Worth naming what's happened across both your comments: you've pointed at my checks twice, been right twice, and neither time did a result move. The scrutiny has landed entirely on the instrument. I think that's the actual argument for publishing the harness rather than the finding, and I hadn't seen it that clearly before you started poking at it. Thank you - this has been the most useful exchange I've had about this work.
Following up since you found two of these. Both fixes shipped, and the repo is public now: github.com/marvinoka4/killcheck
The byte-size-preserving canary is in, running on 8 of 12 targets, with the other 4 reporting N/A rather than silently skipping. The regression test asserts the property (observed behaviour changed) instead of the pyc count, with a meta-test reproducing your stale branch so the check is known to fail when it should. Both accidental protections, ignore_patterns and TemporaryDirectory, are commented as load-bearing and name PYTHONPYCACHEPREFIX explicitly.
Wrote the whole project up short, if useful: dev.to/marvinoka4/69-tests-all-pas...
"Debugging is triggered by surprise" is the most useful sentence I have read about evaluation this year, because it explains a one-directional bias without anyone being dishonest.
Two counter-measures that helped me: pre-register the number you expect before you run, so that a result matching your prediction still has to be audited (agreement with a guess is not evidence the instrument is sound). And keep a deliberate negative control in the suite, a query set where the score should be near zero. When the harness scores that too high you have found a flattering bug you would never have gone looking for, because nothing about it felt surprising.
The negative control is the better of your two suggestions, and I hadn't thought of it in those terms.
I was doing pre-registration, and it still let most of the nine through. Writing the expected number down helps when you're wrong, but it does nothing when the instrument is broken in a way that produces the number you predicted. Bug 7 is exactly that case: I predicted models would write assertion-free tests, and a classifier that misread self.assertEqual would have handed me my prediction with a confirmation attached.
The negative control is different because it inverts which direction is alarming. Everywhere else in the harness, high is good and low prompts investigation. On a set where the score should be near zero, high is the alarm. That gives you one place where the flattering failure is the surprising one, which is the only thing that reliably triggers debugging.
The design constraint I'd add is that the control has to be something your harness cannot accidentally pass for a boring reason. If it scores near zero because the target never loaded, you've built a check that's satisfied by the exact class of bug you're hunting. Which is what bug 1 did to me: three targets at 0.000 that read as a finding rather than a fault.
Your negative control is in, and it earned its place immediately. Repo is public: github.com/marvinoka4/killcheck
It came back at 7 of 51 rather than near zero, and the explanation turned out to be a real calibration fact rather than a bug. assert x is not None detects exactly one class of fault; all seven kills were that one operator, zero were anything else. A control that returned clean would have told me nothing. It also has the guard you'd expect, asserting separately that the suite passes on clean source and that the module actually imported, so it can't pass for the boring reason.
Short write-up of the whole thing: dev.to/marvinoka4/69-tests-all-pas...
Still thinking about your point on pre-registration. Writing the number down did nothing for the bugs that produced the number I predicted.
7 of 51 with all seven from one operator is a better outcome than a clean zero, for exactly the reason you give: a control that returns nothing tells you the control is insensitive, not that the suite is sound. What you actually measured is the suite's coverage profile, which is a more useful artifact than a pass/fail. It also suggests the complementary control: seed mutations the suite should catch and confirm it does, so you have the sensitivity curve from both ends rather than only the floor.
On pre-registration, that is a fair correction and I overstated it. Writing the number down only helps against the bias where a disappointing result triggers investigation and a pleasing one does not. It does nothing for a bug that produces the number you predicted, and yours were mostly that second kind. The honest version is that pre-registration narrows the surprise-triggered filter and leaves everything else untouched, which is a much smaller claim than the one I made.
Asserting separately that the suite passes on clean source and that the module actually imported is the more load-bearing of your two guards.