Scroll through any developer platform right now, and you will immediately notice a massive trend: hating on AI.
It has become so popular that I honestly cannot tell who is just chasing the current trend, and who is genuinely struggling to tame these tools. But the frustration usually stems from a fundamental misunderstanding of what AI actually is.
If you hand a surgical scalpel to someone without surgical training, it does not magically make them a surgeon.
The secret is not the scalpel itself. It is the hard-earned knowledge of how to wield it. AI in software development is exactly the same.
1. Letting the Scalpel Run Wild
I will be controversial here: I don't think "one-shotting" huge features with AI is wrong. I also don't think trusting AI-written code is the root issue.
The real mistake happens before a single prompt is ever written. It is a lack of proper setup.
Most developers let AI agents run wild in massive, existing repositories heavily burdened with accumulated tech debt, anti-patterns, and smelly code. The AI agent scans the repo, assumes this bad code is the acceptable standard, and perfectly replicates your worst practices.
Worse, some teams assume that because AI covers so much ground, they no longer need traditional code quality tools. This is how a tool becomes dangerous.
2. Building the Surgeon's Operating Room
You cannot just write a .cursorrules or agents.md file, ask the AI to "be a good developer," and expect 100% compliance. You have to remove the implicitness. You must frame the AI into conventions that are strictly enforced by tools.
To actually harness AI-driven development, you need an environment built on hard guardrails:
- Write Real Tests, Not Echoes: Stop letting AI generate tests just to satisfy existing code. Build real business logic and conformance tests that actively fail when an agent strays outside a domain or breaks a class.
- Build Regression Fixtures: Record fixtures that allow you to run the new feature code against the previous version. You need an automated way to compare and catch unexpected regressions immediately.
- Gate Everything in CI: Pull requests generated by AI need strict automated gates. If it doesn't pass the pipeline, it doesn't get merged.
- Demand Visual Evidence: Let agents run E2E testing using browser tools. Have them capture and provide actual visual evidence on the PR showing that the feature works.
The goal is not to hold the AI's hand. The goal is to build an infrastructure that maximizes your absolute confidence in the code sitting in that PR.
3. From Token Furnace to Reliable Worker
I currently run a fully autonomous AI factory. My experience has been extremely positive, but getting here wasn't magic.
Initially, I had built a "token furnace", which was basically a system that just burned through API calls without delivering reliable value. It only became a reliable worker when I stopped focusing on the AI and started focusing entirely on the infrastructure. I worked on the environment until it was so fine-tuned that I actually trusted leaving it running alone.
My biggest takeaway? Make everything bounded.
Stop the Snake from Eating Its Own Tail
In my case, the issue usually isn't that reviewing agents hallucinate or fight over trivial syntax. More often than not, I find their actual code review findings to be completely valid. For my factory, the real killer has been documentation rot and context bloat.
To keep agents aligned, you have to provide enough context and compound findings from previous sessions. But this documentation quickly bloats and rots. References get scattered across the repository, and soon, a perfectly capable agent that would have otherwise executed flawlessly gets misguided because an outdated file gave it conflicting instructions.
When my agents start spinning in circles or blocking PRs, it is rarely because the AI failed at reasoning. It is almost always a context problem.
This is why you must remain in control. Human-in-the-loop (HITL) is an absolutely normal and necessary part of autonomous development. If an agent gets stuck in a loop because it is battling stale docs, you need infrastructure that automatically parks the PR. The agents do the heavy lifting, but a human must step in to clean up the context and act as the circuit breaker.
4. The Architect and the Executor
Using AI correctly means the AI works for you, not the other way around. The same rule applies to the development process. The human is supposed to make the decisions, the AI is supposed to execute them, and the infrastructure is supposed to validate that execution and ensure quality.
If an AI makes a rogue decision that negatively impacts your codebase, only the human is to blame. You cannot blame the tool just because you are not using it correctly.
Conclusion
The problem with AI in software development isn’t the AI. The problem is expecting a scalpel to perform the surgery on its own in a dirty operating room.
Stop blaming the tool. Start building the infrastructure, write tests that actually challenge the code, and bound your agents so they work for you, not against each other.
Over to you: Have you tried automating any part of your PR review process? What guardrails have actually worked for your team? Let me know in the comments!
Top comments (0)