You've got a 200-day Duolingo streak. The owl is proud of you.
Then you land in Rome, sit down in a trattoria, and order. Every word is owl-approved. The grammar too. You think.
The waiter nods, smiles, and brings you something you don't really want. Is beef cheeks a thing?
Knowing the words isn't knowing the language. Even with a 200-day streak.
Code is made of words. Words that computers understand. If we could talk to computers with natural language, everything would be great. The problem is not with the words we speak, it's with what the computer understands.
And they understand ones and zeros. That's not going to change.
How do we get from natural language to zeros and ones? We use "code". Our code is an abstraction above ones and zeros.
As any tourist can tell you, knowing a few keywords is not enough to convey the meaning of what you want, and reach the outcome you want. The better you speak the language and know the nuances, the better results you'll get.
Let's look at the process of "getting results". I'll keep it simple, because in the real world this is a lot more complex.
Requirements -> Design -> Code -> Build -> Verification
Each stage includes a translation.
- Requirements – The translation of what the customer wants to what to build
- Design – Translation of what to build to how to build it
- Code – Translation of how to build in our words to computer commands
- Build – Translation of computer commands to zeros and ones
- Verification – Translation of what we see back into what we meant to build
With each translation there are translation errors. We know that, so we put processes in place to plug the leaks. Requirements review, design reviews, code reviews, automation tests, canary releases.
We know there are going to be problems, so we try to minimize the risks.
Final point – except for the translation from computer commands to ones and zeros – we control (or think we do) everything.
So, is AI making code worse?
Not necessarily. I've seen good AI code that is better than some developers' code.
But that's not the point. The code is a translation, this time by a genie, whose way of thinking we can't question or understand. Because it doesn't really think.
We get a dish. But not exactly as we wanted it.
And when we start delegating requirement specification, design and verification to AI, we're losing more and more controlling points.
The outcome is mostly ok, because coding agents get better. But we all know the big problems fall through the cracks.
We all know about security problems and performance issues introduced by AI code. But even the functional stuff suffers because of unintended side effects, weird behavior, and other decisions made by the agent.
If the code worked perfectly, we wouldn't ask the question. But we know we'll need to get in there and fix stuff, and the code will resist. And it will be painful, especially if it's the first time we see it.
And improve code you've never read? That's risky business.
So what can we do?
- Review, review and review. Not just code. Docs, plans, anything the genie produces. You can't fix what you haven't read.
- Work in small chunks. Our review capabilities are a lot more effective in small chunks. That's also about finding errors in code and architecture. For us and the genie.
- Refactor. Once you've read it, if you don't like the code, change it. Tell the genie how to do it better the next time.
And yes, code gets worse. Just because we don't improve it. No matter who wrote it.
Originally published at testingil.com.
I'm Gil Zilberfeld. I teach API testing and test automation, and I write about what AI-generated code does to quality.
Top comments (0)