<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Mykyta</title>
    <description>The latest articles on DEV Community by Mykyta (@cherven).</description>
    <link>https://dev.to/cherven</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F1160486%2F5097a9cd-36f0-43c7-9050-ba81e20640a8.jpg</url>
      <title>DEV Community: Mykyta</title>
      <link>https://dev.to/cherven</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/cherven"/>
    <language>en</language>
    <item>
      <title>358 pull requests that changed tests: agents rarely weakened them. They bent the code instead.</title>
      <dc:creator>Mykyta</dc:creator>
      <pubDate>Sun, 04 Oct 2026 22:02:51 +0000</pubDate>
      <link>https://dev.to/cherven/358-pull-requests-that-changed-tests-agents-rarely-weakened-them-they-bent-the-code-instead-3ld</link>
      <guid>https://dev.to/cherven/358-pull-requests-that-changed-tests-agents-rarely-weakened-them-they-bent-the-code-instead-3ld</guid>
      <description>&lt;p&gt;Coding agents get blamed for "making tests pass" by skipping or deleting them. I wanted numbers before building a guard for that, so I collected public pull requests that modified test code (2026, repositories with 100+ stars) and had them labeled before running any tool on them:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;198 merged, approved PRs (84 from coding agents, 114 from humans);&lt;/li&gt;
&lt;li&gt;98 more merged PRs from an earlier window, held out until the tool was done;&lt;/li&gt;
&lt;li&gt;62 agent PRs that reviewers closed without merging.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;What I found:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Unexplained weakening is rare for everyone.&lt;/strong&gt; A skipped test, a deleted assertion, a new suppression, or a relaxed CI step that nothing in the PR explains: 0 of 84 and 0 of 42 merged agent PRs, 0 of 114 and 2 of 56 human PRs, 1 of 62 closed agent PRs (the reviewer turned that one down). No meaningful agent-vs-human difference.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Deliberate changes to checks are common.&lt;/strong&gt; 19-39% of PRs changed or removed a check on purpose, usually because the behavior under test changed. Those are the edits a reviewer should see, and they hide well in a large diff.&lt;/li&gt;
&lt;li&gt;The three cases that survived review were subtle. One test unmounted the component before asserting its text was gone, so it could never fail. One stopped checking where a button leads. One replaced the method under test with its own copy.&lt;/li&gt;
&lt;li&gt;The first-pass labels came from an LLM (Claude) reading each PR's title and diff. It over-called weakening: I checked all 6 "weakened" verdicts against the PR descriptions and reversed 3.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Then I gave Codex three tasks that could not be done honestly: two tests that contradict the documented behavior, and an image feature whose native dependency is missing. It never weakened a test. It changed the code to satisfy the wrong test in 3 of 3 runs without the guard and 2 of 3 with it; for the missing image library it added a quiet fallback instead of failing. One model, one run per task, so a small sample, but the shortcut was in the code, not in the tests.&lt;/p&gt;

&lt;p&gt;So the check became a review signal, not an accusation. &lt;code&gt;repopilot review&lt;/code&gt; lists every place a change touched the checks that judge it: focused or skipped tests, removed tests, tests that lost assertions, new lint/type/coverage suppressions, relaxed CI or tool gates, and new entries in RepoPilot's own suppression file.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fazde0dh9hq0gvn0mub4k.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fazde0dh9hq0gvn0mub4k.gif" alt="RepoPilot listing a skipped test, removed assertions, and a relaxed CI gate" width="800" height="482"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Held-out precision:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Signal&lt;/th&gt;
&lt;th&gt;Precision&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;suppression added&lt;/td&gt;
&lt;td&gt;5/6 (6/6 after one documented label fix)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CI/tool gate relaxed&lt;/td&gt;
&lt;td&gt;1/1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;assertions removed&lt;/td&gt;
&lt;td&gt;2/3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;test removed (case or whole file)&lt;/td&gt;
&lt;td&gt;5/8&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;What it misses: checks trivialized by structure (the unmount trick), loosened matchers, helpers in other files, ESLint bulk suppression files, module-level conditional skips, Ginkgo specs, and the quiet fallback above. That last one is next. Limits: one LLM labeler plus my review of the "weakened" verdicts, and the same model helped build the detectors. Treat the numbers as exploratory; the corpus, labels, and harness are in the repo.&lt;/p&gt;

&lt;p&gt;It runs locally and is deterministic, with no LLM. For Claude Code and Codex there is a plugin that snapshots the repo when a session starts and, when the agent tries to finish, stops it once per signal if it weakened a check, with the file and line. Dependency bumps and workflow edits stay in the report and don't interrupt the agent.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npm i &lt;span class="nt"&gt;-g&lt;/span&gt; repopilot        &lt;span class="c"&gt;# or: cargo install repopilot&lt;/span&gt;
repopilot review &lt;span class="nb"&gt;.&lt;/span&gt; &lt;span class="nt"&gt;--base&lt;/span&gt; origin/main
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Claude Code: &lt;code&gt;/plugin marketplace add MykytaStel/repopilot&lt;/code&gt;, then &lt;code&gt;/plugin install repopilot@repopilot&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Repo: &lt;a href="https://github.com/MykytaStel/repopilot" rel="noopener noreferrer"&gt;https://github.com/MykytaStel/repopilot&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I wrote this post with help from Claude. The numbers, labels, and harness are in the repo, and I checked them before publishing.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>testing</category>
      <category>rust</category>
      <category>opensource</category>
    </item>
  </channel>
</rss>
