DEV Community

Gil Zilberfeld
Gil Zilberfeld

Posted on Originally published at testingil.com

The Test Was Wrong. Rewriting.

Did you ever see this line when you let your code agent run wild?

“The test was wrong. Rewriting.”

Can the genie really make mistakes?

Ok, enough jokes.

If you’re surprised, this may be the first time you’re generating tests. It happens a lot – seven times in a single build in one of my checks.

Let’s walk through a couple of scenarios, and see how we got to this place.

For our purpose, let’s assume that we generated both code and tests. But the analysis is true for when the code agent generates just the test.

How did we get here?

The genie generates test code. This code is just lines of text in a file – unproven until it runs. It can be proven wrong syntactically, and by some heuristics, logically. Meaning, knowing it’s wrong without running it.

So if the genie runs this check only (we don’t really know unless it decides to tell us), it’s a guess. Not 50-50, but still not a real proof.

Our genie is nothing but a truth chaser. Rightly so, it runs the test. If the generated test fails, we have a fork in the road.

  1. The code is right, and the test is wrong.
  2. The code is wrong, and the test is right.

Nah, it’s not that simple. Both can be wrong. It can also be:

  1. The genie understands the test doesn’t even run the code. Or,
  2. The genie realizes the test isn’t designed correctly to check what it needs to.

Happened to me, where a test that should have recreated a race condition turned out to never run the race at all.

I want you to understand that at least in one of the stations, the genie made a mistake.

  1. Creating the code
  2. Creating the test
  3. Validating the test vs the code
  4. Running the test
  5. Evaluating the test for its purpose

#4 is the easiest to spot. Why?

Because it’s not reasoning. It’s a deterministic run – pass or fail. So a failure means something is definitely wrong.

In all other stages, whatever the genie does with the code is

a) Hidden from us

b) Non-deterministic, meaning we got the result today, but maybe if we ran this yesterday we’d get a different answer, and

c) It actually caught something. Or thinks it did.

Everything except the actual run is opaque and not repeatable. Not what we’d call proof.

A crime may have happened here. Or not.


This is not an anti-genie rant.

Let me be honest. I don’t – I can’t – read everything it logs. I run the agent, and look at the end results. Tests passing? Cool. I usually don’t look back at the logs.

Heck, the race condition issue? I asked Claude to go through the build logs to find where tests were found wrong.

I use coding agents all the time. And looking at me, you’d say – I trust them. Because we were taught that trust looks exactly like this.

But, this is not trust. It’s a bet. Lots of them. The app does work mostly, and when I find something I ask for a fix. And if the agent finds something, it fixes it.

But remember the options? The genie can make mistakes. Also in fixes. And in new code. And in replacement tests.

Who says the fix is the correct one?

My app, even in production, does not carry the risks of fully scaled apps. Imagine hundreds of developers, each with their own agent making those bets on your finance apps. Or law practices. Or online election management.

Are you scared? I know I am.

The way out is chunking the tasks to be smaller and manageable. Smaller pieces of code produce smaller logs. Ones we can read better.

Reviewability – if it’s not a word, it should be – is now a delivery capability.


Originally published at testingil.com.

I'm Gil Zilberfeld. I teach API testing and test automation, and I write about what AI-generated code does to quality.

Top comments (0)