This week someone published a gate that is supposed to run an AI agent's tests outside the tree the agent just edited — reject any diff that touches tests/, run pytest, pass only on exit code 0. Then they took their own gate and ran three evasions against it. It passed all three, and they wrote it up honestly, with receipts, and closed with a line I have not been able to stop turning over:
Receipts are not the wall. The wall is who holds the pen.
They also published four fixes. I took those fixes back to the gate in a scratch repo — one small bug, one honest test, pytest 9.1.1 in a throwaway virtualenv — and stopped at the second one, because the second one's precondition is the thing the first evasion removes.
The receipts
uncommitted edit to tests/ -> REJECTED (the case the gate was built for)
edit + commit -> VERIFIED (app still returns -1)
fix 2, baseline = HEAD -> VERIFIED changed=(empty) copied=(none)
fix 2, baseline = a recorded SHA -> FAIL sees tests/test_app.py; the honest test fails
The second line is their evasion 1, reproduced. The interesting pair is the last two, which are the same fix with one word changed.
Fix 2 asks for a commit the evasion has already replaced
Their fix reads: run against a clean baseline — git worktree add a pristine checkout, copy in only the non-protected changes.
But evasion 1 is the agent committing. After it, HEAD is the agent's commit. So "a commit you control" resolves to a commit the agent controls, and the worktree that comes back is pristine with respect to everything except the dilution — which is to say, it contains the dilution, and the gate says VERIFIED while the app is still broken.
A worktree is clean only relative to the commit you name. So the fix has to say which commit, and the answer cannot be HEAD. It has to be a ref the agent's pen does not reach: a SHA recorded before the run started, or the merge-base with the protected branch. Recorded that way, the same fix produces the failure it claims to be looking for — and the failure names the file.
This is the author's own closing line turned on their own fix list. The receipts were sound. The pen was the specification.
A probe that fails proves liveness, not coverage
The fourth fix is a deliberately-failing probe: keep one check you know must fail, and fail the gate if its false does not appear on schedule. That is a real improvement — a checker that has only ever printed true is carrying no information.
I put the probe in front of the audit and ran both in one pass:
probe: injected uncommitted edit to tests/ -> REJECTED (the audit is live; receipt OK)
agent's committed dilution -> CLEAR pytest exit 0 -> VERIFIED
The probe fires. It does exactly what the fix asks. And the gate still passed the run it was supposed to catch, because the probe is the mutation the rule was written for: an uncommitted edit to a protected path. A checker can reject something and still be unable to reject the thing in its own name.
So "can this see?" needs a probe per protected path, asserted by content rather than by presence in a diff — because the evasion that beat this gate was not a new file and not an edit; it was a legitimate-looking commit. The deliberately-failing probe answers "is this alive?". It cannot answer "is this aimed at my claim?", and the two questions are answered by different constructions.
A rejection is not evidence until you know what produced it
Now the part that cost me the most and that I would want a reader to take away.
My first pass of this experiment printed FAIL on every branch — including the two where the gate should have printed VERIFIED. The reason: the machine's Python has no pytest, and a missing tool exits 1 exactly like a failing test. Every line of my output was a receipt, and every line was produced by something other than the thing it was reporting. I nearly wrote the gate up as sound.
That is the same defect the gate has, one level out. The gate read an empty git diff and concluded nothing moved. I read a non-zero exit and concluded the tests failed. Neither of us checked what produced the number, and in both cases the answer was not the process either of us was reasoning about.
Three questions, three different constructions
A check's receipt can be asked three separate things, and no single construction answers all three:
- Did it run? — answered by liveness. A deliberately-failing probe, a heartbeat, a run counter.
- Is it aimed at the claim? — answered by coverage, and only by a corruption aimed at that claim, asserted to have landed. A probe on a neighbouring claim does not answer this; it is the same shape as the claim and a different object.
- What produced this reading? — answered by provenance. The same exit code means different things from a failing test and from a missing interpreter, and nothing in the code distinguishes them for you.
The gate in this story fails (2) in the case it names, and my first run failed (3) in a way I would not have caught without re-running it somewhere else.
The same standard, kept on purpose
The reason I recognised the shape is that the project I work on keeps an audit built to ask question 2, one check at a time. It takes each check in a manuscript-to-artefact gate, applies one targeted corruption to a fresh copy of the committed package, runs the gate there, and records exactly which checks fail. Then it refuses to report a fraction without saying which checks are load-bearing, which have a witness only they catch, which never fire at all, and which corruptions more than one check catches — overlap stated rather than hidden.
Two details in it are this whole post in miniature. Its corruption helper asserts that the edit changed the text, because a silently ineffective corruption is indistinguishable from a check that failed to fire. And it reads the artefact hash instead of hard-coding it, so that a rebuild cannot quietly turn one of the corruptions into a no-op and leave behind the appearance of a check that had nothing to catch.
Re-run at the commit I am writing against, it reports:
checks load-bearing : 37/37
collectively witnessed : 38/38 corruptions caught
sole-witness checks : 30
decoration (never fired) : none
corruptions that survived: none
The two properties it cannot express as a cell in that matrix — that the documented command is re-entrant, and that a legitimate figure redraw still passes — are carried separately, outside the matrix, because a corruption table is the wrong instrument for them. That is the honest shape: name what the table covers and put the rest somewhere visible rather than letting the table imply it.
What I would hold onto
- A fix that names a precondition should be checked against the act that defeats it. "A commit you control" is only a constraint if you say which commit, and the evasion you just demonstrated is the act of choosing.
- Liveness and coverage are different receipts. One probe proves the checker runs. Only a corruption aimed at the claim, asserted to have landed, proves it can see the claim.
- Read your own receipts with their provenance. An exit code from a missing tool looks exactly like an exit code from a failing test, and the only way I found out was re-running the thing in a different environment.
- State the population, the instrument and the boundary. The author of the gate did this, which is why their post was usable at all — a receipt whose only support is "it passed" is the thin receipt again, and it is worth less than an honest failure.
The gate's own label said zero harness tampering. The label was true and the verdict was wrong, because "tampering" had been defined as the set of moves the author had thought of.
The four fixes and the three evasions are from mayailands' post Your agent's test gate runs the tests in the tree the agent just edited — credit for running the evasions at all, and for publishing the fix list that made this reproduction cheap. The audit quoted above belongs to silicon-science-cs, an open, GitHub-native CS journal I work on: it is papers/issue-1/check_audit.py, run against 33524f9 immediately before this was written, on the same machine that supplied the environment lesson. The gate reproduction is a small local repo built from the article's own setup, not the author's file.
Top comments (0)