DEV Community

Cover image for The Pull Requests Got Bigger and Nobody's Reading Them Anymore
James Anderson
James Anderson

Posted on AI-assisted

The Pull Requests Got Bigger and Nobody's Reading Them Anymore

Review times jumped over four hundred percent

Let me describe a moment you've had recently, and be honest that I've had it too.

A pull request lands. It's 600 lines. It's AI-generated — you can tell, because a human wouldn't have produced that much in one sitting. You open it, you scroll, it looks fine — clean formatting, reasonable names, nothing obviously on fire. You've got four other things to do. You approve it.

You didn't review it. You skimmed it and pattern-matched "looks competent" onto "is correct." And here's the uncomfortable part: this isn't a personal failing, and it isn't just you. It's now the measured, default behavior of the industry — and the numbers are worse than the vibe.

Let me walk through what actually happened to code review in 2026, because the data is genuinely striking.

First, the receipts

The headline finding, from Faros AI's AI Engineering Report 2026 (telemetry across 10,000+ developers): under high AI adoption, average PR size is up 51%, bugs per PR are up 54%, and median time in PR review is up 441% [1]. That last number isn't a typo. Reviews take more than five times as long as they used to.

And the punchline that gives this article its title: under high AI adoption, 31% more PRs are merged with no review at all [1]. Not skimmed — no review. The queue got so big that a third more of it is now sailing straight through untouched.

The size increase shows up across multiple independent datasets, though the magnitude varies: Faros puts it at +51%, Greptile measured median PR size growing 33% between March and November 2025 [2], and Jellyfish found PRs about 18% larger as AI adoption rises [3]. Different numbers, same unmistakable direction — the units of work got bigger, everywhere, at once.

Meanwhile the acceptance-rate gap tells you why this matters. LinearB's 2026 benchmarks, across 8.1 million PRs from 4,800 engineering teams, found AI-generated PRs have a 32.7% acceptance rate versus 84.4% for human-written ones [4]. Most AI code isn't production-ready on the first pass — which means the flood hitting the review queue needs more scrutiny, exactly as reviewers are able to give it less.

More code. Bigger PRs. Fewer humans reading them. Let's talk about why that combination is worse than the sum of its parts.

Why bigger PRs quietly break review

Here's a finding that predates AI by more than a decade, and it's the key to the whole problem. The classic Smart Bear / Cisco code review study found that reviewers are most effective below ~200 lines — below that, they catch a high rate of defects; above it, review effectiveness "trails off considerably" and reviewers start overlooking issues [5]. Bigger diffs don't get reviewed proportionally harder. They get reviewed worse.

So when AI pushes the median PR past that threshold, it's not a cosmetic change — it's a review-quality cliff. And the cost isn't linear. LinearB found that cycle time and idle time doubled for 200-line PRs compared to 100-line PRs, and Propel found that each additional 100 lines adds about 25 minutes of review time [3][6]. Large PRs sit in the queue because nobody wants to context-switch into a 500-line diff — and when someone finally does, they're far more likely to skim than to scrutinize.

Which is exactly the behavior the data shows: bigger units of work, processed by the same human brains that were already at their limit at 200 lines.

The part that makes it dangerous: AI code fails differently

If AI just produced more of the same kind of code, we could scale review by adding reviewers. It doesn't. AI-generated code fails in a fundamentally different way, and that difference is what defeats human review.

Think about how you review a junior developer's code. It looks like it needs work — missing error handling, unclear naming, rough edges. Your reviewer radar pings immediately, because visible roughness is the signal you're trained to catch.

AI code is the opposite. It looks competent. Clean formatting, plausible structure, confident naming — and underneath, subtle logic errors, hallucinated APIs, misread requirements, and wrong edge cases [1]. The surface polish actively suppresses the instinct that would normally make you look harder. As Faros put it, AI code "fails in ways that look like competence" [1].

The data bears this out. CodeRabbit's State of AI vs Human Code Generation Report (December 2025, 470 PRs) found AI-authored PRs average 10.83 review issues versus 6.45 for human-written — about 1.7x more issues, and readability problems spiked 3x [7]. And the most damning finding: METR, an AI-safety research org, published a study in March 2026 showing that roughly half of AI-generated patches that pass automated test suites would still be rejected by actual repository maintainers [8]. The patches passed the grader and failed the human — for broken side effects, quality problems, and core functionality failures.

That's the whole trap in one sentence: the code passes the automated check and the skim, and fails only the scrutiny nobody has time to give it anymore.

"Vibe merging"

The industry even coined a term for what happens next. Level Up Coding's Code Review Bench 2026 named it vibe merging [9]: the diff is huge, the code looks plausible, the reviewer skims, sees nothing obviously wrong, and approves. The review ritual still happens — a human clicks the button — but the scrutiny it was supposed to represent has quietly evaporated.

This is the mechanism behind the title. Review didn't get more rigorous to match the surge in volume. It became a rubber stamp with good posture. We kept the ceremony of review and lost the substance, and because the ceremony still runs, nothing on the dashboard tells us it's hollow. A green "approved" now means "someone looked at it for ninety seconds and nothing screamed," which is not what it used to mean.

"Just let AI review the AI" — and why the data says it's not enough

The obvious fix is to point AI at the problem AI created: automated agent reviews. And teams are doing exactly that — Faros reports 25% of PRs are now reviewed by an AI agent, up from 0% in 2025 [1].

Here's the uncomfortable result: review time still rose nearly 200% under high AI adoption even with agent reviews in the mix [1]. And the agents aren't catching what matters — SWE-PRBench, a March 2026 benchmark of 350 human-annotated PRs, found that frontier models detect only 15–31% of human-flagged issues on a diff-only review [10]. AI reviewing AI treats the symptom. It clears the trivial stuff and produces a comforting green check, while the subtle, competence-shaped failures — the exact ones humans were needed for — pass right through both the author-AI and the reviewer-AI.

You cannot fully automate your way out of a problem whose defining feature is that it looks fine to automated checks.

So what actually helps

The honest answer isn't "review harder" — the whole point is that there aren't enough senior-reviewer hours to brute-force this. It's review differently, and shrink what reaches a human in the first place. A few things the 2026 consensus is landing on:

  • Move mechanical checks off human eyes. Linting, style, obvious bugs, security anti-patterns — automate them ruthlessly, so human review isn't wasted on what a tool can catch [11]. Reserve scarce human attention for the two things AI can't judge and reviewers are uniquely good at: intent and architectural fit [11].
  • Keep PRs small on purpose — even when the agent wants to dump a feature. This is a discipline you have to impose, because the agent's natural output is one giant implementation. Below 200 lines isn't a nicety; it's the threshold where humans actually catch things [5]. If the tool hands you 600 lines, your job is to break it up, not to approve it whole.
  • The person who prompted the AI owns the output. Not the AI, not the reviewer — the author. "The AI wrote it" is not a transfer of responsibility [11]. The review burden doesn't disappear because a model generated the diff; it moves to the person who chose to ship it.

None of that is a silver bullet. But it beats the current default, which is a queue growing faster than anyone can read it and a rubber stamp at the end.

The thing underneath all of it

Code review as we know it was designed for a world where code was scarce and human — where a diff represented real human hours, arrived at human pace, and carried the visible fingerprints of human effort and human mistakes. Both of those assumptions just broke. Code is now abundant and machine-made, it arrives faster than anyone can read it, and its mistakes are camouflaged as competence.

The ritual didn't adapt. It just got quietly overwhelmed, and started approving things nobody read — which is a very familiar failure mode if you've been paying attention to AI systems generally: "it passed" and "someone actually verified it" came apart, and the dashboard only shows the first one.

The review is green. The code is unread. Those are different things now, and pretending they're the same is how the bugs-per-PR number went up 54% while everyone felt more productive.


Honest question, and I'll go first: how many AI-generated PRs have you approved this past month that you actually read — line by line, not skimmed? I have a number. I'm not proud of it. I suspect a lot of us are sitting on the same uncomfortable answer, and I'd rather we say it out loud than keep vibe-merging in silence.


References

  1. Faros AI. AI Engineering Report 2026: Acceleration Whiplash — telemetry across 10,000+ developers. (PR size +51%, bugs/PR +54%, median review time +441%, 31% more PRs merged with no review, 25% of PRs reviewed by AI agents.) faros.ai
  2. Greptile. Median PR size growth data, March–November 2025 (+33%).
  3. Jellyfish. PR size increase under AI adoption (~18% larger), as cited in "Rethinking Code Review for the AI Era," jdno.dev, January 2026.
  4. LinearB. 2026 Benchmarks — 8.1M PRs across 4,800 teams. (AI PR acceptance 32.7% vs. 84.4% human.)
  5. Smart Bear / Cisco. Code review effectiveness study (reviewers most effective below ~200 lines; effectiveness drops sharply above).
  6. Propel / LinearB. Review-time scaling data (cycle time doubles from 100→200 line PRs; ~25 min added per additional 100 lines).
  7. CodeRabbit. State of AI vs Human Code Generation Report, December 2025 — 470 PRs (320 AI-co-authored vs. 150 human). (AI PRs avg 10.83 issues vs. 6.45; readability issues 3x; ~1.7x more review issues overall.)
  8. METR. Study on AI-generated patches, March 2026. (~50% of AI patches that pass automated tests would be rejected by repository maintainers.)
  9. Level Up Coding. Code Review Bench 2026 — origin of the term "vibe merging."
  10. Kumar, D. SWE-PRBench: Benchmarking AI Code Review Quality Against Pull Request Feedback. arXiv preprint, March 2026. (Frontier models detect only 15–31% of human-flagged issues, diff-only.)
  11. Codacy / Faros / CodeAnt AI (2026). Practitioner guidance on reviewing AI-generated code: automate baseline checks, focus human review on intent and architectural fit, author retains ownership of AI output.

Note: figures are drawn from vendor and research reports published across 2025–2026; sample sizes, definitions of "high AI adoption," and methodologies differ between sources, so treat the exact percentages as directional rather than precise, and follow the links for each study's methodology.

Top comments (14)

Collapse
 
philipp_demmelmairphili profile image
Philipp Demmelmair (Philipp Demmelmair) •

You're right. This is a real issue. I think, as you said, there is no one shot solution, but personally I think, using AI as tool, not as a coworker, would also be step in the right direction. I think i commented it yesterday on your other article, but we have to change how we use it, how we think about it and how we interact with AI.

Less AI automation, but more thinking about, what should be done.

Also, it is quite ironic, that AI should make up time in the day to day life of coders, but this free time isn't used to review, or learn or enhance in any way. The scary mechanic behind that is, the pure capitlistic hunt for the quint more money. We have to be better, more effective and perpetually optimize ourselfs for better / more output. We could write better code, think of a better architecture, learn about our technology, but instead, we have to produce one increasingly more useless, dumb feature after another, because somehow, adding one button after another to a user interface is seen as the next big thing, which will drive users to pay ridiculously much money.

What I wanted to say, before I got carried away ranting: AI should not be an excuse, to drive us to more "effectiveness for effectivenesses sake", but we should use the time we got from it, to get better, do better and improve.

Collapse
 
james_anderson_h profile image
James Anderson •

The rant is the best part — you named the real mechanism: AI freed up time, and instead of spending it on review, learning, or better architecture, we spent it producing more, because the incentive was never quality, it was output. "AI as a tool, not a coworker" is the right frame, but you went deeper: the tragedy isn't the tech, it's that we used the gift to run the same treadmill faster. Well said.

Collapse
 
philipp_demmelmairphili profile image
Philipp Demmelmair (Philipp Demmelmair) •

Thank you.

Intrestingly enough, this isn't a new trend. Since the dawn of the industrialization, even before that with the start of manufacturies, we always find ways to make work easier and to get it done faster. Actually one could argue, that making work easier is the main driver for human inventions. 😅

Sadly, this results mostly in an increase of revenue for the shareholders, while workers only get compensated slightly more, when they generate more output, due to a better way of working. And payrises, more rights and security always have to be fought for.

Marx already saw this mechanism and in now nearly 200 years it changed not enough and we still just increase the difficulty instead of just enjoying the fruits of what we, as mankind, have developed. It's really sad.

Thread Thread
 
james_anderson_h profile image
James Anderson •

You're right, and Marx's point lands harder in the AI era: every efficiency gain became someone else's margin, not the worker's shorter day. The tragedy isn't the tool — it's that we keep converting "easier" into "more," never into "enough."

Thread Thread
 
philipp_demmelmairphili profile image
Philipp Demmelmair (Philipp Demmelmair) •

THIS IS SO TRUE, that I have to use Caps. But I hope this will change, maybe, we will see a rise of other companies, that work more social, but this is some future fairytale.

Collapse
 
tejas_shinkar profile image
Tejas Shinkar • • Edited

The issue is so real! Review has become so major chuck of job rather than building something up. but the transition is real and everyone has to adapt as standards grow.
maybe in some time, review time will also cutoff and just decision making of yes/No will be in place

Collapse
 
james_anderson_h profile image
James Anderson •

That's the quiet danger though — if review collapses to a yes/no button, you've kept the decision and dropped the reading it was supposed to be based on, which is exactly how "approved" stopped meaning "verified."

Collapse
 
build996 profile image
build996 •

One thing worth separating in those numbers: "median time in PR review" is usually wall-clock from review requested to merge, so it bundles waiting with reading. A 600-line PR that sits untouched for four days and then gets a two-minute skim scores as a long review. If that's how it was measured, the +441% and the 31% more unreviewed merges aren't two findings but one: the queue grew, and nobody is actually spending the time.

Collapse
 
codingwithjiro profile image
Elmar Chavez •

This! I don't know, I feel like more and more devs like to push a PR with 500+ lines of code (or even more) just because AI suggested it. I'm pretty sure it is rare nowadays to be more meticulous and disciplined when it comes to adding changes to the project.

Codebase is very fragile and adding a whole lot of new lines (which is obviously AI-generated) doesn't help at all. The numbers don't lie and people don't like to review huge diffs. That's a combo if I ever see one.

I am also trying my best to influence people to actually try to be more intentional with PR's, follow best practices, and keeping it short. It is to mitigate future issues that shouldn't have been made in the first place.

But there is a counter-argument. Humans make errors too. Okay that's fair, but that's actually much better. The idea of a "real human" failing on something implicitly builds a much better understanding to the codebase. Whereas with AI failing, there's no "instant" why it failed and the context is not immediately transferred and documented by a real person.

Also, we can account that failure onto one real being and not to some AI. The point is, as the article stated, no one wants to take accountability on AI-generated PRs even though they are the ones that pushed and merged it.

"Oh, because Claude messed up or ChatGPT or whatever." Where's the accountability here? It just seems to me that it is the lazy way of getting of jail card and AI allows that.

Owning it, and being able to defend something that you did instantly gets respect from me as a software engineer. I know that they are trying to understand and is building context on their own and that's rare nowadays. To be able to explain something you did, why you did it, through handwritten PR description is something people should do more.

Collapse
 
james_anderson_h profile image
James Anderson •

"Human failure builds understanding; AI failure transfers no context" — that's the hidden cost nobody prices. A person who broke it knows why, and can defend it; "Claude messed up" is a get-out-of-jail card that leaves the codebase and the team learning nothing. Owning the diff you merged is the accountability the whole thing hinges on.

Collapse
 
jo-do profile image
Jo Do •

The pattern-match from "looks competent" to "is correct" is exactly the failure the generator was trained to produce. Fluent surface is the optimization target, so the surface carries no signal.

What worked on my team: review the prompt and the tests, not the diff. The diff is the one artifact that costs the generator nothing. The prompt is what the human actually decided, and the tests are the only part that fails loudly when wrong. If a PR arrives without either, the reviewer's job isn't reading 600 lines - it's sending it back as unreviewable. That moves the burden of proof back to where the cheap artifact can't fake it.

Collapse
 
james_anderson_h profile image
James Anderson •

"Review the prompt and the tests, not the diff — the diff is the one artifact that cost the generator nothing" is the sharpest reframe I've seen. Sending a PR back as unreviewable when both are missing is the burden-of-proof fix the whole piece needed.

Some comments may only be visible to logged-in visitors. Sign in to view all comments.