Testers have been badgering developers for years to write unit tests. Now even those devs who don’t like tests can generate tests.
So now testers say, well, we don’t trust these tests. And they have good reason, too. Unit tests may pass, while the code is broken.
First, let’s explain the badgering. We don’t have a lot of time. Didn’t before AI, and definitely not now.
So when we get our hands on the app, or the new features you created, dear developer, we want it as stable as possible. We don’t want it to break the moment we touch it.
Now, one way is to check every build from scratch. But like I said, we don’t have time. If we have proof that you checked the code in some way, that means we don’t have to waste time on it.
So you’ve got tests? That’s proof. Thank you!
Except…
Well, where did the tests come from? Ah. Generated. I see. Hmm.
And I assume there are no bugs in the code, right? Because, if there are, and you’ve got some tests covering that bug, then you and I are pretty screwed.
What do I mean? Here’s an example.
Here’s a function that checks the expiration of a credit card.
def is_expired(exp_month, exp_year, today):
return date(exp_year, exp_month, 1) < today
Looks ok. But not really. Credit cards have MM/YY expiration, and that means that they do not expire until the END of the month.
This code has a bug in it. It causes a card to expire on the 2nd of the month.
No worries, a test will catch that bug, right? Here’s what my genie created:
def test_card_expired_last_year():
assert is_expired(9, 2025, date(2026, 9, 16)) is True
def test_card_expiring_this_month():
assert is_expired(9, 2026, date(2026, 9, 16)) is True
The first one’s great. It checks that a card with 09/25 expiration is indeed expired.
The second one is weird. We know the card should not expire. It’s not the end of the month yet, so we expect the result to be False, not True.
Why did it generate the wrong test?
Well, if you generate tests for the code, you get tests that pass for the code. Not for the requirement. This is called a Tautological test – it confirms that the code works as written. But not as it should really work.
What the genie didn’t have is context. So it guesses when it generates the code, filling holes in what it has been told, and what “good code looks like”.
But without context (and let’s face it, sometimes with it) the code may be wrong.
The bigger problem is that now the tests make it look like it works as intended.
And what happens next? The code breaks when we touch it.
That’s why testers have trust issues.
One more thing – if a human read the code, chances are the bug could have been caught in the review.
So what can we do?
- Ask which tests are generated and which we can trust.
- Ask what code is reviewed by a human.
- Know where the risky logic sits and test it extensively.
- Spot the passing tests (that shouldn’t pass) and fix them.
- Teach the devs to do that.
The irony is that while we finally got more tests, our confidence in them is lower.
Originally published at testingil.com.
I'm Gil Zilberfeld. I teach API testing and test automation, and I write about what AI-generated code does to quality.
Top comments (0)