A team can absorb more AI-generated code for a while. Volume climbs, pull requests get bigger, reviewers work a little harder. Then something shifts: on the largest changes, review time stops climbing and starts flattening out. That plateau is not a sign the pipeline is healthy. It is the moment reviewers quietly stopped reading.
Two measured primary sources from real engineering organizations converge on this exact warning, and both say the fix is not a faster reviewer. It is changing what a reviewer is shown.
The plateau Salesforce recorded
Salesforce described the shift in detail in Scaling Code Reviews: Adapting to a Surge in AI-Generated Code on their engineering blog in January 2026, by Ravi Boyapati. The internal numbers are specific: code volume up roughly 30%, pull requests regularly spilling past 20 files and 1,000 lines of change, review latency rising quarter over quarter.
The worrying figure is not the latency. It is what happened to the biggest pull requests. Review time on the largest changes began to plateau, and in places decline. More code was being reviewed in less time. That combination is not efficiency. It is the tell that reviewers had switched from evaluating the change to validating it, defaulting to surface approval because exhaustive review on a 1,000-line, 20-file diff no longer felt worth the effort.
This reframes the whole problem: when review effort exceeds the perceived value of the change, reviewers stop engaging deeply, and it is a systems failure, not an individual one.
Why file-by-file diffs break under load
Salesforce's diagnosis is that the traditional review model stops working for a structural reason, not a volume one. File-by-file diff scanning assumes a reviewer can reconstruct intent by reading changes in sequence. That holds when changes are small and incremental. AI-generated pull requests violate the assumption in a predictable way: they span backend logic, configuration, tests, and UI in one submission, with the narrative structure stripped out.
The failure modes are concrete. Conceptual coherence is lost, so reviewers infer purpose from disconnected fragments. Cognitive load grows faster than PR size, so reviewers spend more time navigating files than reasoning about behavior. Review incentives invert, pushing toward mechanical approval. Without AI-assisted signals surfacing risk, the reviewer is left as the sole filter for both trivial syntax and architectural flaws at once.
Intent reconstruction and context as the levers
Salesforce's response was to re-architect review around intent reconstruction rather than diff scanning. Instead of treating the diff as flat text, the system preserves structural boundaries (files, functions, modules) through token-aware chunking so semantic units stay intact. It builds higher-order abstractions progressively, then merges related changes across the codebase so backend logic and its UI counterpart are reviewed together rather than in isolation.
The second lever is context. The Salesforce design treats context as a first-class dependency: unified context drawn from work items, prior pull requests, historical defects, and codebase patterns, surfaced incrementally through progressive disclosure. The reviewer still decides, but attention is directed at the high-risk areas instead of being spread evenly across the diff.
Adoption is up, and the time effect is still a coin flip
The LeadDev 2026 State of AI-Driven Software Releases report, surfaced through LeadDev's April 2026 writeup "Use of AI has us creating more code than we can review," found 68% of teams say AI already influences their approach to code review, and 28% now use AI-powered code review tools, up from 17% the prior year. Among those using AI, 86% use it to identify issues before a human ever looks at the code.
That last number matters. Review is moving left: the first pass at catching defects is increasingly automated, before a human reads anything. But adoption climbing does not mean delivered review capacity is climbing with it. The report's time effect is close to a coin flip across teams, which is exactly what you would expect if most tools save a little time but the structural problem of scrutiny under load is untouched.
What the measured sources converge on
Put the Salesforce and LeadDev evidence next to the peer-reviewed ContextCRBench work (ACM FSE 2026 companion, DOI 10.1145/3803437.3805235), which showed that LLM code review accuracy improves when the model is given issue and PR context rather than the code alone, and a consistent picture emerges.
The bottleneck is not throughput, and it never really was. It is scrutiny: whether anyone is actually reasoning about what the change does, at the scale a single large queue demands. On the human side, the answer Salesforce converged on is to preserve intent and feed context so the reviewer can reason. On the model side, the answer ContextCRBench converges on is the same: context, not bare code, is what makes the evaluation accurate.
A tool that can reconstruct what a change is trying to do, tie related changes together, and bring in the surrounding context is doing the part of review that file-by-file scanning cannot. Whether that tool is built internally like Prizm or bought from a vendor, the test is the same. Does it make the reviewer reason about intent, or does it just make the diff easier to skim?
That distinction is the difference between a review workflow that absorbs AI volume and one that just looks like it absorbed it. A plateau in review time on your largest PRs is the earliest sign you are on the wrong side of it. I walked through the same plateau signal in an earlier post on what the best coding agent's error rate means for your review load, and it keeps showing up in the sources.
What to check on your own repositories
The concrete check, before buying or building anything, is to pull your own numbers. Group your closed PRs by size and look at review time per line for the largest band. If it is flat or declining while smaller PRs are still climbing, you are in the plateau Salesforce described. That is your signal that reviewers are approving on surface rather than reasoning.
Then look at your largest PRs and ask whether anyone can still reconstruct, from the diff alone, what the change was trying to accomplish. If the answer is no, that is not a reviewer failure. It is a workflow that stopped presenting changes in a way a human can evaluate.
Top comments (0)