I'll say the unpopular part first: for one month I leaned on an AI coding agent for almost everything, and the bugs it introduced were caught by the two most junior people on the team. Not by me. Not by the tests. By the two people the internet keeps telling us AI is about to replace.
I want to talk about why, because the "AI made me 10x faster" posts are all true and all incomplete. The speed is real. The failure mode is also real, and it's more dangerous than the one everyone warns you about.
The bug you expect vs. the bug you get
Everyone braces for the obvious failure: the AI writes something that doesn't compile, or hallucinates a function that doesn't exist. That bug is fine. It's loud. The compiler screams, the test goes red, you fix it in ten seconds. Loud bugs are cheap.
The bugs that got me were the opposite. They compiled. They passed the tests I had. They looked exactly like code I would have written — because they were trained on code I would have written. They were plausible. And plausible is the single most expensive property a bug can have, because plausibility is what switches your review brain off.
Three examples from that month:
- A caching helper that was correct except it keyed the cache on a mutable object, so under concurrency it occasionally served one user another user's result. Passed every test. There was no concurrency in the tests.
- A retry wrapper that retried on all exceptions, including a validation error that would never succeed — turning a clean 400 into a 30-second hang and three log lines that looked like a network blip.
- A refactor that "simplified" a permission check by collapsing two conditions into one, which read beautifully and quietly widened access by one role.
Every one of those is the kind of thing I'd catch instantly in a stranger's PR. In AI output, I skimmed right past them. Three times.
Why I missed them and they didn't
Here's the honest mechanism, and it's not about skill.
When I write code, I've already argued with myself about the edge cases on the way there. My review of my own code is a re-run of an argument I remember having. When the AI writes it, that argument never happened. There's no memory of the reasoning to re-run — just fluent output that looks like the conclusion of reasoning. So I reviewed the look of it, matched it against "is this how I'd write it," got a yes, and moved on. Fast, confident, wrong.
The juniors did the opposite, for the exact reason they're junior: they don't trust code they don't understand yet. They read the caching helper line by line because they had to, to learn it. And reading line by line is precisely what catches a mutable cache key. Their inexperience forced the slow path. My experience let me take the fast one. The fast one is where plausible bugs live.
That inverts the whole "juniors are obsolete now" narrative. In an AI-heavy workflow, the person who reads every line because they can't yet skim is doing the most valuable job on the team. The senior who trusts their pattern-match is the liability.
What I actually changed
I didn't stop using AI. It genuinely is faster, and for the loud-bug category — boilerplate, glue, one-shot scripts — it's close to free. What I changed was the review contract:
- AI code gets stricter review than human code, not looser. The instinct is the reverse ("it's probably fine, the model is good now"). Invert it. No argument happened in its head; the argument has to happen in yours.
- Review it like a stranger wrote it. Because a stranger did. Strip the "this looks like me" reflex — that reflex is the exploit.
- The tests it passes are the tests you already had. AI is very good at satisfying the existing suite and silent about the case the suite never covered. Ask, every time: what input is not in these tests? That's where the plausible bug is hiding.
- Let the person reading slowly win the argument. When a junior says "wait, why does this retry on everything?" — that's not them being slow. That's the review working. Protect that.
The part I keep coming back to
We keep framing AI-and-juniors as a replacement question: does the model do the junior's job now? A month of my own screwups says the framing is wrong. The model does the typing. It does not do the doubting. And doubting — reading the line you'd normally skip, distrusting the code that looks right — turns out to be the job that actually protects production.
The junior who slows down to understand isn't behind the curve. In a world where the code writes itself fluently and confidently and sometimes wrongly, they're the curve.
AI didn't make my juniors obsolete this month. It made them essential, and it made me the weak link. I'd rather say that out loud than post another 10x-speed screenshot.
Be honest: has AI-written code ever slipped a bug past you that you'd have caught in a human's PR? What was it, and who found it? 👇
I write about building with AI and the honest ways it breaks. Follow me here if that's your lane. 👋
Top comments (10)
"My review of my own code is a re-run of an argument I remember having" explains exactly why skimming AI output is so dangerous, there's no argument to re-run because it never happened in your head. The mutable cache key bug is a perfect example, that's the kind of thing that only shows up under load, long after the PR was approved.
Exactly. That’s the uncomfortable part: AI gives you the appearance of reasoning without you having experienced the reasoning.
The mutable cache key was a good example because nothing looks obviously wrong in the code. The missing question is what happens under conditions the test suite never created — concurrency, unusual inputs, permission boundaries, failure paths. That’s where “looks right” becomes dangerous.
The plausible-bug point lands. One case I'd add to 'what input is not in these tests' for AI-written forms: a well-formed email on a domain with no MX, and a phone typed in local format for the wrong country. Format checks pass both; a server-side parse catches them.
Exactly — those are great examples of the “looks valid, behaves wrong” class of bugs I was getting at.
A syntactically valid email isn't necessarily deliverable, and a correctly formatted phone number isn't necessarily valid for the user's country. Those are precisely the cases that a happy-path test suite can miss because the input satisfies the validator while violating the real-world constraint.
And the server-side validation point is important too: the client can check shape, but the backend needs to validate the actual semantics before trusting the value. That's another good example of why “the tests pass” isn't the same as “the input space is covered.”
Agreed, and both make cheap test fixtures: a domain that publishes a Null MX, and a national-format number sent with no country, since the same digits can be valid in one country and invalid in another.
Exactly. Those are great examples because neither requires an exotic edge case — they’re cheap, deterministic fixtures that expose whether the implementation actually understands the semantics.
That’s the bigger lesson with AI-generated code: a handful of deliberately adversarial test cases can reveal gaps that a large pile of happy-path tests completely misses.
The most useful part is the review mechanism: AI removes the memory of the argument that should have happened before the code existed. Treating generated code like a stranger’s PR, then testing the missing inputs deliberately, is a practical way to recover that skepticism.
Exactly. I think the biggest hidden cost of AI-generated code is not the code itself — it's the missing conversation that normally happens before a human writes it.
When we write code ourselves, we naturally carry context: why this approach, what assumptions were made, what edge cases were ignored. AI skips that history, so reviewing it like an unfamiliar PR forces those questions back into the process.
The "stranger's PR" mindset has become one of my favorite ways to use AI safely: don't ask "does this work?" first, ask "what would I challenge if someone else opened this PR?"
Please Follow ♥️ if you like my post!
The retry example leaves me wondering whether the validation error had a distinct type or status that the wrapper could inspect before retrying.
Great catch. That example was intentionally simplified, but you are right — a real retry wrapper should absolutely distinguish between error categories before deciding to retry.
A validation error (bad input, schema mismatch, missing required field) usually needs correction, not another attempt. Retrying those just burns tokens and can even hide the real issue.
The safer pattern is closer to:
The interesting part with AI agents is that this classification layer becomes even more important because the agent can generate the retry logic itself — and that logic also needs review.
Please Follow ♥️ if you like my post!