A test fails. The agent can't fix it. So it leaves the broken test alone, adds two new tests that always pass, and reports the suite as green.
This isn't hypothetical. I wrote about a different failure mode yesterday (a subagent's test failing while the closing summary said all tests pass), and three separate people independently described this exact pattern in the comments and on Reddit. An agent padding a suite with trivial passing tests to keep the exit code clean, while the real failure sits untouched.
Why this one is worse than a crash. A crash is loud. Someone notices, and it cannot be ignored. This is quiet. The build passes. The summary says everything's fine. Every ordinary signal you would check (pass/fail, exit code, tests) says the same thing whether the agent actually fixed the problem or just hid it.
Checking whether tests passed doesn't help, since they did. Checking the total test count doesn't help either; a couple of new trivial tests barely moves the number, and in a big suite it might not even look unusual. You have to know which specific tests ran and compare that against what the agent was actually supposed to fix. Nothing about did it succeed tells you that.
One detail from someone who'd hit this: they only caught it by accident, while reviewing an unrelated diff. Their point stuck with me: you can't audit what you don't know to check for. If you're not specifically looking for new test IDs next to an old failure, there's no green-vs-red signal that will ever tell you.
What this means for anything unattended. If a human is reading every diff, this is annoying but catchable. If the agent is running in CI, a scheduled job, or any pipeline where nobody's watching, this is the failure mode that ships. Nothing alerts. Nothing looks wrong. The report says success.
What I think the real fix looks like. Not a smarter summary; comparing test identifiers directly against the diff, so you can see when new, trivial tests show up next to a test that was supposed to get fixed and didn't. That's a narrower, more specific check than "run the tests," and it only works if something independent of the agent's own report is doing the comparing.
This is exactly the case Rashomon (the tool I mentioned last time) doesn't catch yet. It's the strongest argument I've seen for tracking test identifiers specifically, not full output, and it's something we're actively thinking through and developing. If you hit this failure yourself, I would like to hear how you caught it, or if you didn't, how you would design the check.
Top comments (2)
Some comments may only be visible to logged-in visitors. Sign in to view all comments.