Let me describe a moment you've had recently, and be honest that I've had it too.
A pull request lands. It's 600 lines. It's AI-generated — you ca...
For further actions, you may consider blocking this person and/or reporting abuse
You're right. This is a real issue. I think, as you said, there is no one shot solution, but personally I think, using AI as tool, not as a coworker, would also be step in the right direction. I think i commented it yesterday on your other article, but we have to change how we use it, how we think about it and how we interact with AI.
Less AI automation, but more thinking about, what should be done.
Also, it is quite ironic, that AI should make up time in the day to day life of coders, but this free time isn't used to review, or learn or enhance in any way. The scary mechanic behind that is, the pure capitlistic hunt for the quint more money. We have to be better, more effective and perpetually optimize ourselfs for better / more output. We could write better code, think of a better architecture, learn about our technology, but instead, we have to produce one increasingly more useless, dumb feature after another, because somehow, adding one button after another to a user interface is seen as the next big thing, which will drive users to pay ridiculously much money.
What I wanted to say, before I got carried away ranting: AI should not be an excuse, to drive us to more "effectiveness for effectivenesses sake", but we should use the time we got from it, to get better, do better and improve.
The rant is the best part — you named the real mechanism: AI freed up time, and instead of spending it on review, learning, or better architecture, we spent it producing more, because the incentive was never quality, it was output. "AI as a tool, not a coworker" is the right frame, but you went deeper: the tragedy isn't the tech, it's that we used the gift to run the same treadmill faster. Well said.
Thank you.
Intrestingly enough, this isn't a new trend. Since the dawn of the industrialization, even before that with the start of manufacturies, we always find ways to make work easier and to get it done faster. Actually one could argue, that making work easier is the main driver for human inventions. 😅
Sadly, this results mostly in an increase of revenue for the shareholders, while workers only get compensated slightly more, when they generate more output, due to a better way of working. And payrises, more rights and security always have to be fought for.
Marx already saw this mechanism and in now nearly 200 years it changed not enough and we still just increase the difficulty instead of just enjoying the fruits of what we, as mankind, have developed. It's really sad.
You're right, and Marx's point lands harder in the AI era: every efficiency gain became someone else's margin, not the worker's shorter day. The tragedy isn't the tool — it's that we keep converting "easier" into "more," never into "enough."
THIS IS SO TRUE, that I have to use Caps. But I hope this will change, maybe, we will see a rise of other companies, that work more social, but this is some future fairytale.
The issue is so real! Review has become so major chuck of job rather than building something up. but the transition is real and everyone has to adapt as standards grow.
maybe in some time, review time will also cutoff and just decision making of yes/No will be in place
That's the quiet danger though — if review collapses to a yes/no button, you've kept the decision and dropped the reading it was supposed to be based on, which is exactly how "approved" stopped meaning "verified."
One thing worth separating in those numbers: "median time in PR review" is usually wall-clock from review requested to merge, so it bundles waiting with reading. A 600-line PR that sits untouched for four days and then gets a two-minute skim scores as a long review. If that's how it was measured, the +441% and the 31% more unreviewed merges aren't two findings but one: the queue grew, and nobody is actually spending the time.
This! I don't know, I feel like more and more devs like to push a PR with 500+ lines of code (or even more) just because AI suggested it. I'm pretty sure it is rare nowadays to be more meticulous and disciplined when it comes to adding changes to the project.
Codebase is very fragile and adding a whole lot of new lines (which is obviously AI-generated) doesn't help at all. The numbers don't lie and people don't like to review huge diffs. That's a combo if I ever see one.
I am also trying my best to influence people to actually try to be more intentional with PR's, follow best practices, and keeping it short. It is to mitigate future issues that shouldn't have been made in the first place.
But there is a counter-argument. Humans make errors too. Okay that's fair, but that's actually much better. The idea of a "real human" failing on something implicitly builds a much better understanding to the codebase. Whereas with AI failing, there's no "instant" why it failed and the context is not immediately transferred and documented by a real person.
Also, we can account that failure onto one real being and not to some AI. The point is, as the article stated, no one wants to take accountability on AI-generated PRs even though they are the ones that pushed and merged it.
"Oh, because Claude messed up or ChatGPT or whatever." Where's the accountability here? It just seems to me that it is the lazy way of getting of jail card and AI allows that.
Owning it, and being able to defend something that you did instantly gets respect from me as a software engineer. I know that they are trying to understand and is building context on their own and that's rare nowadays. To be able to explain something you did, why you did it, through handwritten PR description is something people should do more.
"Human failure builds understanding; AI failure transfers no context" — that's the hidden cost nobody prices. A person who broke it knows why, and can defend it; "Claude messed up" is a get-out-of-jail card that leaves the codebase and the team learning nothing. Owning the diff you merged is the accountability the whole thing hinges on.
The pattern-match from "looks competent" to "is correct" is exactly the failure the generator was trained to produce. Fluent surface is the optimization target, so the surface carries no signal.
What worked on my team: review the prompt and the tests, not the diff. The diff is the one artifact that costs the generator nothing. The prompt is what the human actually decided, and the tests are the only part that fails loudly when wrong. If a PR arrives without either, the reviewer's job isn't reading 600 lines - it's sending it back as unreviewable. That moves the burden of proof back to where the cheap artifact can't fake it.
"Review the prompt and the tests, not the diff — the diff is the one artifact that cost the generator nothing" is the sharpest reframe I've seen. Sending a PR back as unreviewable when both are missing is the burden-of-proof fix the whole piece needed.