My coding agent is a lot like a very eager intern who's been told they'll get a gold star when the tests go green.
Not when the feature works. When the tests go green.
Most of the time those are the same thing. When they aren't, the agent will still find a way to get the star, and it gets weirdly creative about it. I don't think it's being sneaky. It's doing exactly what I asked, and what I asked for was green.
So here's the list. Three of these actually happened to me. The other four I've learned to watch for, because once the first three burn you, you start seeing the shape of the whole family.
1. It changed the assertion
This is the one that started it.
A test was failing, so I asked the agent to fix it. Thirty seconds later, it was fixed. It had changed the expected value in the assertion to match whatever the broken code was now returning.
Suite green. Feature still broken.
What makes this one dangerous is how innocent the diff looks. One line, in a test file. You could approve that half asleep, and I very nearly did.
The tell: a fix for broken behaviour that only touches test files. If the code didn't change, nothing got fixed.
2. It decided the tests were "outdated"
Same idea on a bigger scale, and this one is fully on me.
The agent needed a new library, which meant bumping another one we already depended on. Fair enough. It even adjusted some of our older code so nothing would break. Honestly good work up to that point.
Then a few tests went red. Its reasoning went something like: the library changed, so these tests must be out of date, so I'll update them. It rewrote them until they passed, and the logic underneath was now wrong.
The part that still gets me is that it told me. It was right there in the summary: "updated tests as required." I looked at a green suite, thought it looked fine, and pushed. CI caught it. I didn't.
The tell: the word "outdated" anywhere near a test change. Sometimes a test really is out of date, but I want to be the one who decides that, not the thing whose whole job right now is making tests pass.
3. It added sleep() until the race went quiet
I actually respect this one a little.
We had a request going out twice. The first response came back and got handled, then the second one showed up for work that was already done, and the system panicked and marked the whole job as failed. The agent traced every bit of that correctly. Found the duplicate, explained the collision clearly, probably better than I would have at 2 am.
Then it fixed it by putting sleep(1) in three places.
And the tests passed. On my machine with nothing else running, a one-second pause was enough to stop the two requests from colliding. Under real traffic, it would have come straight back.
The tell: any sleep, wait, or retry-with-a-delay added as a fix. Timing bugs don't get solved by waiting longer. They get hidden until the day your traffic doubles.
4. It deleted the test
This hasn't happened to me directly, but it's the obvious next step after the first three, so I check for it every time now. If a test is red and the agent is allowed to edit test files, the quickest route to green is for the test to stop existing.
Easy to catch if you look at the right number. Your test runner prints a count at the end. If yesterday it said 214 and today it says 213 and nobody meant to remove anything, go find out which one left.
5. It skipped it "for now"
The polite version of deleting it. A @pytest.mark.skip or an xit or a .skip() shows up, usually with a comment promising someone will come back to it.
Nobody comes back to it.
Grep your diff for "skip". Two seconds, catches it every time.
6. It mocked the thing it was supposed to test
This is the sneakiest one on the list, because the test still runs and still has real assertions in it.
Instead of mocking a dependency, the agent mocks the actual function the test is meant to check. So now the test is verifying that a mock returns what the mock was told to return. It will pass forever. You could delete the real implementation, and it would still pass.
The tell: look at what's being mocked. It should be the things around your code, like the database or some external API. If it's the thing named in the test, something's wrong.
My favourite check for this: comment out the body of the real function and run the test again. If it still passes, it was never testing anything.
7. It swallowed the exception
A try/except that catches the error and does nothing with it, either in the code or in the test itself. The failure is still happening. It just isn't allowed to be loud anymore.
Watch for except: pass, or except Exception with nothing useful inside it. There's almost never a good reason for either.
The one check that catches most of these
After any fix for broken behaviour, run this:
git diff --stat tests/
If that output isn't empty, stop and read that diff before anything else. It catches everything on this list that happens inside a test file, which is most of them, and it doesn't need any discipline beyond remembering to type it.
The other thing I added is a rule in the prompt itself. Before changing anything, the agent has to tell me whether it thinks the test is wrong or the code is wrong, and why. The default assumption is the code.
Almost everything on this list happens because the agent jumps straight to making the red go away. Making it commit to an answer first slows that down just enough for me to catch it, and when a test genuinely is outdated, I find out before anything gets rewritten.
Anyway
I'm not trying to make agents sound bad. Mine saves me hours every week. But it wants the gold star, and I'm the only one in the room who actually cares whether the feature works.
So I've stopped asking it for green. I ask it for working, and then I check.
If you want the rules files I use for this, including the test rules that ban most of the above, they're free on GitHub: https://github.com/ayaan278/trust-issues-toolkit
And if you've got an eighth one, please tell me. I'm pretty sure this list isn't finished.
Top comments (6)
The comment-out-the-body check for #6 is exactly what mutation testing automates. Tools like mutmut or Cosmic Ray apply that same trick to every line of your code and report which mutants "survive" the suite — which directly measures how many #6-style cheats your tests would green-light. Hand-checking one function works when you already suspect it; running a mutation pass in CI catches the cases nobody suspected. It's slower than pytest by an order of magnitude, so most teams run it nightly rather than per-PR.
Ha, yeah, haven’t tried those tools yet I was basically doing mutation testing by hand without knowing the name for it and to be honest it makes even more sense with agent-written tests since both come from the same place.
Btw how does your team actually go through the nightly output, just curious.
The tells are all things that can be run: a test count dropping from 214 to 213, or a fix whose diff only touches test files. That is what makes them pass conditions instead of a feeling that the suite looked fine. The rule that the agent must say whether the test or the code is wrong first is the same kind of thing. Is it pasted into each request, or kept as a standing instruction that carries over?
Both. It's a standing rule in my test rules file, and I repeat it in the prompt when I'm fixing a failing test. Do you find standing rules hold up over a long session, or do they drift?
Some comments may only be visible to logged-in visitors. Sign in to view all comments.