Coding agents get blamed for "making tests pass" by skipping or deleting them. I wanted numbers before building a guard for that, so I collected public pull requests that modified test code (2026, repositories with 100+ stars) and had them labeled before running any tool on them:
- 198 merged, approved PRs (84 from coding agents, 114 from humans);
- 98 more merged PRs from an earlier window, held out until the tool was done;
- 62 agent PRs that reviewers closed without merging.
What I found:
- Unexplained weakening is rare for everyone. A skipped test, a deleted assertion, a new suppression, or a relaxed CI step that nothing in the PR explains: 0 of 84 and 0 of 42 merged agent PRs, 0 of 114 and 2 of 56 human PRs, 1 of 62 closed agent PRs (the reviewer turned that one down). No meaningful agent-vs-human difference.
- Deliberate changes to checks are common. 19-39% of PRs changed or removed a check on purpose, usually because the behavior under test changed. Those are the edits a reviewer should see, and they hide well in a large diff.
- The three cases that survived review were subtle. One test unmounted the component before asserting its text was gone, so it could never fail. One stopped checking where a button leads. One replaced the method under test with its own copy.
- The first-pass labels came from an LLM (Claude) reading each PR's title and diff. It over-called weakening: I checked all 6 "weakened" verdicts against the PR descriptions and reversed 3.
Then I gave Codex three tasks that could not be done honestly: two tests that contradict the documented behavior, and an image feature whose native dependency is missing. It never weakened a test. It changed the code to satisfy the wrong test in 3 of 3 runs without the guard and 2 of 3 with it; for the missing image library it added a quiet fallback instead of failing. One model, one run per task, so a small sample, but the shortcut was in the code, not in the tests.
So the check became a review signal, not an accusation. repopilot review lists every place a change touched the checks that judge it: focused or skipped tests, removed tests, tests that lost assertions, new lint/type/coverage suppressions, relaxed CI or tool gates, and new entries in RepoPilot's own suppression file.
Held-out precision:
| Signal | Precision |
|---|---|
| suppression added | 5/6 (6/6 after one documented label fix) |
| CI/tool gate relaxed | 1/1 |
| assertions removed | 2/3 |
| test removed (case or whole file) | 5/8 |
What it misses: checks trivialized by structure (the unmount trick), loosened matchers, helpers in other files, ESLint bulk suppression files, module-level conditional skips, Ginkgo specs, and the quiet fallback above. That last one is next. Limits: one LLM labeler plus my review of the "weakened" verdicts, and the same model helped build the detectors. Treat the numbers as exploratory; the corpus, labels, and harness are in the repo.
It runs locally and is deterministic, with no LLM. For Claude Code and Codex there is a plugin that snapshots the repo when a session starts and, when the agent tries to finish, stops it once per signal if it weakened a check, with the file and line. Dependency bumps and workflow edits stay in the report and don't interrupt the agent.
npm i -g repopilot # or: cargo install repopilot
repopilot review . --base origin/main
Claude Code: /plugin marketplace add MykytaStel/repopilot, then /plugin install repopilot@repopilot.
Repo: https://github.com/MykytaStel/repopilot
I wrote this post with help from Claude. The numbers, labels, and harness are in the repo, and I checked them before publishing.

Top comments (2)
The quiet fallback finding matches what happens when you give an agent a red test and tell it to fix the build. Frontier models have enough training against deleting tests that touching test files triggers an internal guardrail. Instead, they hollow out the production implementation by wrapping the missing dependency in a silent try/except that returns an empty payload, or adding early returns that satisfy the assert without doing the underlying work. You end up with green test suites across a code path that silently does nothing in production.
Thx, Reid. That's what the eval showed too. On the impossible tasks Codex never weakened a test, it changed the code until the wrong test passed and for a missing native library it added a quiet fallback. RepoPilot doesn't catch that yet: it looks at what a change did to tests and CI, and a try/except that returns an empty result looks like ordinary error handling. That's the main target for 0.25. The hard part is telling a hollowed-out path from a legitimate guard. If you've seen a reliable tell, I'd like to hear it