TL;DR — Coding assistants have made writing code nearly free, but they haven't made reviewing, verifying, and trusting that code any cheaper. The result is a new bottleneck — the diff tax — that most tooling still ignores because it's optimized for generation speed, not reviewability. The teams getting real value are the ones who redesigned their workflow around verification, not around prompt quality.
Every pitch for an AI coding assistant leads with the same number: lines of code generated, time saved, tickets closed per sprint. Those numbers are real. They're also measuring the wrong half of the job.
Writing code was never the expensive part of software engineering. It felt expensive because it was slow and tedious, so when a tool made it fast, the improvement was visible and dramatic. But the actual cost of software has always lived downstream of writing: in review, in verification, in understanding why a change does what it does, in catching the one wrong assumption buried in two hundred otherwise-correct lines. Coding assistants didn't eliminate that cost. They didn't even reduce it. They just moved it, and in a lot of cases they made it worse by increasing the volume of code that needs to pass through that same narrow verification channel.
Call it the diff tax. Every AI-generated change still has to be read, reasoned about, and trusted by a human before it ships — and reading code for correctness is slower than writing it, especially code you didn't write and whose author can't explain its intent beyond "the model suggested it." A ten-minute prompt that produces a two-hundred-line diff can cost an engineer an hour of review. That math only works in your favor if the diff is trivial, well-scoped, or extensively self-tested. Most assistant output is none of those things by default.
Where the help is real
This isn't an argument that coding assistants are overrated across the board. In specific, bounded contexts, they are a genuine productivity unlock, and it's worth being precise about which contexts those are.
Mechanical, pattern-matched work. Boilerplate, CRUD scaffolding, converting a function from one language idiom to another, writing the fortieth similar test case — tasks where correctness is easy to check by inspection because the pattern is already established in the codebase.
Unfamiliar API surface. Figuring out the right call signature for a library you've never used is a search problem, not a reasoning problem, and assistants are excellent search engines with context.
First-draft generation under a tight spec. If you can describe the exact shape of the output — input/output types, edge cases, error behavior — the assistant's job collapses to pattern completion, and verification collapses to checking against the spec you already wrote.
Local, mechanical refactors. Renaming, extracting, restructuring within a single file or function where the diff is large but the semantic change is small and easy to confirm.
Notice the common thread: in every one of these cases, the verification cost stays low relative to the generation benefit. The assistant isn't making a judgment call. It's filling in a shape that's already fully specified by context, convention, or an explicit spec. The human's review job is a bounded check, not an open-ended investigation.
Where it quietly breaks down
The failure mode isn't that assistants write bad code. Most of the time the code runs, passes a cursory read, and even passes tests you didn't think to write. The failure mode is that assistants are confident at exactly the moments they should be uncertain, and that confidence is what breaks the review process.
Ambiguous requirements are the clearest case. When a spec has a gap, a human engineer either flags it or picks an interpretation and flags that choice loudly in a comment or a PR description. An assistant picks an interpretation and presents it with the same tone of certainty it uses for syntax. The reviewer now has to reverse-engineer which parts of the diff reflect a real decision versus which parts are just plausible-looking filler — and the assistant gives no signal about where that line is.
Cross-system reasoning is the second case. Assistants are strong within the context window they can see. They are much weaker at reasoning about behavior that emerges from the interaction of services, caches, retries, and timing that live outside that window — the kind of bug that shows up in production under load, not in a unit test. This is exactly the work senior engineers spend the most time on, and it's the work assistants are least equipped to do reliably, because it requires a model of the system that no single file or prompt captures.
Security and trust boundaries are the third. An assistant will happily generate code that works but quietly widens a trust boundary — logging a secret, skipping input sanitization, assuming a caller is authenticated because the surrounding code implied it. These are exactly the mistakes that pass a casual review, because the code "looks right." That's the dangerous category: not code that obviously fails, but code that obviously succeeds and is subtly wrong.
Tooling is optimizing for the wrong metric
Most coding-assistant tooling today is still built around a single optimization target: get from prompt to plausible code as fast as possible. Autocomplete latency, context window size, multi-file edit capability — all generation-side metrics. Almost none of the mainstream tooling optimizes for the thing that actually determines whether the output creates value: how fast and how confidently a human can verify it.
That's backwards. If writing code is now nearly free, the scarce resource is reviewer attention, and tooling should be designed around conserving it. A few concrete implications follow from taking that seriously:
Smaller diffs by default, even if it takes more turns to get to the final result. A ten-line change you can verify in thirty seconds beats a two-hundred-line change you have to trust blind.
Self-generated tests as a verification artifact, not an afterthought. An assistant that writes a failing test before the fix, then shows the fix making it pass, hands the reviewer evidence instead of just an assertion.
Explicit flagging of assumptions. If the model filled a spec gap, that choice should be visible in the diff — a comment, a note, anything — rather than silently absorbed into code that reads as settled fact.
Provenance over the diff. Which lines came from the assistant verbatim, which were edited by a human, which were accepted without changes — this is the kind of metadata that makes review targeted instead of exhaustive.
None of this is exotic. It's the same discipline good engineers already apply to their own pull requests: keep diffs reviewable, make assumptions visible, give the reviewer evidence instead of asking for trust. Coding assistants didn't remove the need for that discipline. They removed the friction that used to force it, because a human writing two hundred lines by hand naturally paces themselves into smaller, more legible units of work. An assistant has no such constraint unless you build it in.
The actual skill shift
The engineers getting durable value out of these tools aren't the ones with the best prompts. They're the ones who've rebuilt their personal workflow around fast, cheap verification — tight specs, aggressive test scaffolding, deliberately small units of generated change, and a habit of treating every AI-authored diff as a claim that needs evidence, not a gift that needs gratitude. The tools will keep getting better at generation. The organizations that win won't be the ones that generate the most code. They'll be the ones that figured out, early, that generation was never the bottleneck.
Top comments (0)