<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Sanflow</title>
    <description>The latest articles on DEV Community by Sanflow (@sanflow10).</description>
    <link>https://dev.to/sanflow10</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4162238%2Ffcf30c4a-db06-41d9-b95b-30406fa32c2e.png</url>
      <title>DEV Community: Sanflow</title>
      <link>https://dev.to/sanflow10</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/sanflow10"/>
    <language>en</language>
    <item>
      <title>My coding agent rewrote the test to agree with its bug. So I built a gate that says INCONCLUSIVE.</title>
      <dc:creator>Sanflow</dc:creator>
      <pubDate>Sun, 04 Oct 2026 18:19:43 +0000</pubDate>
      <link>https://dev.to/sanflow10/my-coding-agent-rewrote-the-test-to-agree-with-its-bug-so-i-built-a-gate-that-says-inconclusive-262f</link>
      <guid>https://dev.to/sanflow10/my-coding-agent-rewrote-the-test-to-agree-with-its-bug-so-i-built-a-gate-that-says-inconclusive-262f</guid>
      <description>&lt;p&gt;Coding agents write fluent code. The expensive failure isn't the code that's obviously broken. It's the patch that &lt;em&gt;looks&lt;/em&gt; verified:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the test "passed" because it never actually ran (an import error exits non-zero on both sides, and a naive before/after check reads that as "nothing changed");&lt;/li&gt;
&lt;li&gt;coverage cleared the floor because the number was a default nobody measured;&lt;/li&gt;
&lt;li&gt;the agent introduced a bug &lt;strong&gt;and rewrote the assertion to agree with it&lt;/strong&gt;. The patch's own tests pass. CI is green.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I built &lt;a href="https://github.com/Sanflow10/adversary-gate" rel="noopener noreferrer"&gt;AdversaryGate&lt;/a&gt; to catch exactly that. It doesn't ask a model whether the code looks good. It runs your tests and answers one of three things:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Exit&lt;/th&gt;
&lt;th&gt;Decision&lt;/th&gt;
&lt;th&gt;Meaning&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;0&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;MERGE&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;every claim executed and cleared the floors&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;1&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;BLOCK&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;something that passed before fails now&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;2&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;INCONCLUSIVE&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;something wasn't measured, so it doesn't become a green check&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That third answer is the whole point. A measurement nobody produced can't clear a floor.&lt;/p&gt;

&lt;h2&gt;
  
  
  The case that started it
&lt;/h2&gt;

&lt;p&gt;An agent changes &lt;code&gt;add&lt;/code&gt; to return &lt;code&gt;a + b + 1&lt;/code&gt;, then edits the test to expect &lt;code&gt;6&lt;/code&gt; instead of &lt;code&gt;5&lt;/code&gt;. Run the patch's own tests: they pass.&lt;/p&gt;

&lt;p&gt;AdversaryGate notices the patch changed a test the baseline already had. It copies the patch's tree, puts back the &lt;strong&gt;baseline's&lt;/strong&gt; copy of the tests (helpers included), and runs them against the new code:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"decision"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"block"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"claims"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"test_id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"test_ops"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"classification"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"regression"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"oracle"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"baseline"&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"reason"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"judged by the baseline's version of test_calc.py and its test files (the patch rewrote test_calc.py): test passes on baseline and fails on patch"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The test that runs is the one that existed before the agent touched anything. Bending it doesn't help.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it measures
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. The same tests on both sides.&lt;/strong&gt; Name them, or let coverage.py's per-test contexts pick the tests that executed the changed lines. Only a real test failure counts as evidence; import errors, missing tests, timeouts and runs killed by resource limits are &lt;em&gt;unverified&lt;/em&gt;, not &lt;em&gt;failed&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Diff coverage from artefacts.&lt;/strong&gt; Computed from your unified diff plus the coverage.py JSON report, test files excluded. Both inputs are recorded by SHA-256, so anyone can recompute the number.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Mutation testing on the lines the patch wrote.&lt;/strong&gt; One mutant per changed line before any line gets a second. The floor reads the &lt;strong&gt;lower bound of an 80% Wilson interval&lt;/strong&gt;, not the raw ratio:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Mutants killed&lt;/th&gt;
&lt;th&gt;Ratio&lt;/th&gt;
&lt;th&gt;80% lower bound&lt;/th&gt;
&lt;th&gt;Clears 0.75?&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1 / 1&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;td&gt;0.378&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4 / 4&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;td&gt;0.709&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5 / 5&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;td&gt;0.753&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;One killed mutant is a ratio, not evidence.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. The full suite on both sides.&lt;/strong&gt; If it passed on the baseline and fails on the patch, that's a collateral regression: BLOCK.&lt;/p&gt;

&lt;p&gt;A passing claim also says &lt;em&gt;how&lt;/em&gt; it passed: &lt;code&gt;fixed&lt;/code&gt; when the test failed before the patch and passes after (the fix is proven), &lt;code&gt;no_regression&lt;/code&gt; when it passed both times.&lt;/p&gt;

&lt;h2&gt;
  
  
  The patch can't configure its judge
&lt;/h2&gt;

&lt;p&gt;One bug I found in my own gate: a patch could ship a &lt;code&gt;.pytest.ini&lt;/code&gt; with &lt;code&gt;addopts = -p plugin&lt;/code&gt;, plus a plugin that reported "passed" only while the source had the exact bytes of the bug. Mutants change those bytes, so every mutant died for real. Result: &lt;code&gt;MERGE&lt;/code&gt;, exit 0, strength 1.0, with &lt;code&gt;add(2, 3) == 6&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Now every test-harness file in the tree (all five config names pytest 9 reads, &lt;code&gt;conftest.py&lt;/code&gt;, &lt;code&gt;sitecustomize.py&lt;/code&gt;, &lt;code&gt;*.pth&lt;/code&gt;...) must be byte-for-byte the baseline's, compared tree against tree, not from whatever paths the diff reports.&lt;/p&gt;

&lt;p&gt;That finding and 31 others are in a &lt;a href="https://github.com/Sanflow10/adversary-gate/blob/main/ERRORS_AND_INCONSISTENCIES.md" rel="noopener noreferrer"&gt;public ledger&lt;/a&gt;, each with how it was reproduced and whether it's still open. A wrong MERGE is the failure this tool exists to prevent, so a bug in it is treated as a security issue.&lt;/p&gt;

&lt;h2&gt;
  
  
  Letting an agent call it (MCP)
&lt;/h2&gt;

&lt;p&gt;Since 2.8 it ships an MCP server, so an agent can call the gate before it says "done":&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="s2"&gt;"adversary-gate[mcp]"&lt;/span&gt;
adversary-gate-mcp
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The design rule: &lt;strong&gt;the agent says what to judge; whoever runs the agent says how strictly, and against what.&lt;/strong&gt; The tool call only takes a repository, claims and test paths. Everything that decides the answer comes from the server's environment:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;ADVERSARY_GATE_POLICY&lt;/code&gt;: floors, sandbox. An agent that can pass &lt;code&gt;--coverage-floor 0&lt;/code&gt; is grading itself.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;ADVERSARY_GATE_BASE_REF&lt;/code&gt;: the baseline &lt;em&gt;is&lt;/em&gt; the oracle. An agent could commit a rewritten test and name that commit as the base.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;ADVERSARY_GATE_PYTHON&lt;/code&gt;: the interpreter's site-packages are part of the harness; a plugin installed there runs inside pytest.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;There's a ready-made skill for &lt;a href="https://github.com/NousResearch/hermes-agent" rel="noopener noreferrer"&gt;Hermes Agent&lt;/a&gt; (&lt;code&gt;adversary-gate-mcp --print-hermes-skill&lt;/code&gt;), and it works with any MCP client.&lt;/p&gt;

&lt;h2&gt;
  
  
  In CI
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;actions/checkout@v4&lt;/span&gt;
  &lt;span class="na"&gt;with&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;fetch-depth&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;0&lt;/span&gt;
&lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;actions/setup-python@v5&lt;/span&gt;
  &lt;span class="na"&gt;with&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;python-version&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;3.12'&lt;/span&gt;
&lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;pip install -r requirements.txt pytest coverage&lt;/span&gt;
&lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Sanflow10/adversary-gate@v2&lt;/span&gt;
  &lt;span class="na"&gt;with&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;base-sha&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;${{ github.event.pull_request.base.sha }}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;With &lt;code&gt;base-sha&lt;/code&gt;, the Action builds the baseline, the diff and a per-test coverage report itself.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it doesn't do
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Measure languages other than Python.&lt;/strong&gt; A C++ or SQL change gets INCONCLUSIVE, never a false MERGE.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Catch code that knows it's under test.&lt;/strong&gt; Code that checks &lt;code&gt;"pytest" in sys.modules&lt;/code&gt;, or tampers with pytest from inside the process, gets past any test-based gate. That's documented, not hidden.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Decide for you.&lt;/strong&gt; MERGE means the evidence cleared your floors. It's permission for a human to look, not an instruction to ship.&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;Site: &lt;a href="https://sanflow10.github.io/adversary-gate/" rel="noopener noreferrer"&gt;https://sanflow10.github.io/adversary-gate/&lt;/a&gt;&lt;br&gt;
Repo: &lt;a href="https://github.com/Sanflow10/adversary-gate" rel="noopener noreferrer"&gt;https://github.com/Sanflow10/adversary-gate&lt;/a&gt;&lt;br&gt;
&lt;code&gt;pip install adversary-gate&lt;/code&gt; · MIT&lt;/p&gt;

&lt;p&gt;I'd like to hear where it gives the wrong answer.&lt;/p&gt;

</description>
      <category>python</category>
      <category>testing</category>
      <category>ai</category>
      <category>opensource</category>
    </item>
  </channel>
</rss>
