Coding agents write fluent code. The expensive failure isn't the code that's obviously broken. It's the patch that looks verified:
- the test "passed" because it never actually ran (an import error exits non-zero on both sides, and a naive before/after check reads that as "nothing changed");
- coverage cleared the floor because the number was a default nobody measured;
- the agent introduced a bug and rewrote the assertion to agree with it. The patch's own tests pass. CI is green.
I built AdversaryGate to catch exactly that. It doesn't ask a model whether the code looks good. It runs your tests and answers one of three things:
| Exit | Decision | Meaning |
|---|---|---|
0 |
MERGE | every claim executed and cleared the floors |
1 |
BLOCK | something that passed before fails now |
2 |
INCONCLUSIVE | something wasn't measured, so it doesn't become a green check |
That third answer is the whole point. A measurement nobody produced can't clear a floor.
The case that started it
An agent changes add to return a + b + 1, then edits the test to expect 6 instead of 5. Run the patch's own tests: they pass.
AdversaryGate notices the patch changed a test the baseline already had. It copies the patch's tree, puts back the baseline's copy of the tests (helpers included), and runs them against the new code:
{
"decision": "block",
"claims": [{
"test_id": "test_ops",
"classification": "regression",
"oracle": "baseline"
}],
"reason": "judged by the baseline's version of test_calc.py and its test files (the patch rewrote test_calc.py): test passes on baseline and fails on patch"
}
The test that runs is the one that existed before the agent touched anything. Bending it doesn't help.
What it measures
1. The same tests on both sides. Name them, or let coverage.py's per-test contexts pick the tests that executed the changed lines. Only a real test failure counts as evidence; import errors, missing tests, timeouts and runs killed by resource limits are unverified, not failed.
2. Diff coverage from artefacts. Computed from your unified diff plus the coverage.py JSON report, test files excluded. Both inputs are recorded by SHA-256, so anyone can recompute the number.
3. Mutation testing on the lines the patch wrote. One mutant per changed line before any line gets a second. The floor reads the lower bound of an 80% Wilson interval, not the raw ratio:
| Mutants killed | Ratio | 80% lower bound | Clears 0.75? |
|---|---|---|---|
| 1 / 1 | 1.00 | 0.378 | no |
| 4 / 4 | 1.00 | 0.709 | no |
| 5 / 5 | 1.00 | 0.753 | yes |
One killed mutant is a ratio, not evidence.
4. The full suite on both sides. If it passed on the baseline and fails on the patch, that's a collateral regression: BLOCK.
A passing claim also says how it passed: fixed when the test failed before the patch and passes after (the fix is proven), no_regression when it passed both times.
The patch can't configure its judge
One bug I found in my own gate: a patch could ship a .pytest.ini with addopts = -p plugin, plus a plugin that reported "passed" only while the source had the exact bytes of the bug. Mutants change those bytes, so every mutant died for real. Result: MERGE, exit 0, strength 1.0, with add(2, 3) == 6.
Now every test-harness file in the tree (all five config names pytest 9 reads, conftest.py, sitecustomize.py, *.pth...) must be byte-for-byte the baseline's, compared tree against tree, not from whatever paths the diff reports.
That finding and 31 others are in a public ledger, each with how it was reproduced and whether it's still open. A wrong MERGE is the failure this tool exists to prevent, so a bug in it is treated as a security issue.
Letting an agent call it (MCP)
Since 2.8 it ships an MCP server, so an agent can call the gate before it says "done":
pip install "adversary-gate[mcp]"
adversary-gate-mcp
The design rule: the agent says what to judge; whoever runs the agent says how strictly, and against what. The tool call only takes a repository, claims and test paths. Everything that decides the answer comes from the server's environment:
-
ADVERSARY_GATE_POLICY: floors, sandbox. An agent that can pass--coverage-floor 0is grading itself. -
ADVERSARY_GATE_BASE_REF: the baseline is the oracle. An agent could commit a rewritten test and name that commit as the base. -
ADVERSARY_GATE_PYTHON: the interpreter's site-packages are part of the harness; a plugin installed there runs inside pytest.
There's a ready-made skill for Hermes Agent (adversary-gate-mcp --print-hermes-skill), and it works with any MCP client.
In CI
- uses: actions/checkout@v4
with:
fetch-depth: 0
- uses: actions/setup-python@v5
with:
python-version: '3.12'
- run: pip install -r requirements.txt pytest coverage
- uses: Sanflow10/adversary-gate@v2
with:
base-sha: ${{ github.event.pull_request.base.sha }}
With base-sha, the Action builds the baseline, the diff and a per-test coverage report itself.
What it doesn't do
- Measure languages other than Python. A C++ or SQL change gets INCONCLUSIVE, never a false MERGE.
-
Catch code that knows it's under test. Code that checks
"pytest" in sys.modules, or tampers with pytest from inside the process, gets past any test-based gate. That's documented, not hidden. - Decide for you. MERGE means the evidence cleared your floors. It's permission for a human to look, not an instruction to ship.
Site: https://sanflow10.github.io/adversary-gate/
Repo: https://github.com/Sanflow10/adversary-gate
pip install adversary-gate · MIT
I'd like to hear where it gives the wrong answer.
Top comments (1)
Exit 2 only means something if the workflow looks at the code. The Action example fails the step on any non-zero. That holds until someone adds continue-on-error, because C++ and SQL always come back INCONCLUSIVE and you don't want those PRs red. A BLOCK (exit 1) then fails the same way as a language the tool can't measure. Branch on the code. Let 2 pass only where the post already says measurement isn't possible.