Hello, I'm Shrijith Venkatramana, and I'm building LiveReview — a blast-radius aware AI code review built for your business-critical systems. Star us to help devs discover the project, give it a try, and share your feedback to help improve the product.
AI coding agents have created a peculiar problem for software engineering: writing code is getting cheaper, but establishing that the code is safe to ship still requires scarce human attention.
At Meta, this problem became measurable. Over one year, significant lines of code per human-landed diff increased by 105.9%, while diffs per developer per month increased by 51%. More than 80% of the increase in diff volume was attributed to agentic AI.
Meanwhile, the share of diffs reviewed within 24 hours was declining. Some large engineering groups had thousands of pending reviews.
Meta's response was to automate a specific class of code reviews rather than attempt to automate all of them.
The system, called RADAR (Risk Aware Diff Auto Review), combines machine-learned risk scores, LLM-based code review, deterministic checks, and conservative eligibility policies. Its central idea is simple: let low-risk changes proceed through automated verification, and reserve human attention for changes that need human judgment.
In their 2026 paper, Automating Low-Risk Code Review at Meta: RADAR, Risk Calibration, and Review Efficiency, Chris Adams, Nachiappan Nagappan, Peter Rigby, and their colleagues report results from more than 535,000 RADAR-reviewed diffs.
The reported outcomes are worth understanding in detail: the system reduced review latency while exhibiting substantially lower revert and production-incident rates than non-RADAR changes.
Here is how it works, why the design matters, and what engineering teams can learn from it.
Code generation scaled faster than code review
Code review has always served several purposes: finding defects, maintaining consistency, sharing knowledge, and making engineers accountable for changes.
But these benefits come with an operational cost. Every change entering a review queue competes for attention from someone who must understand its context, inspect its implementation, and decide whether it is safe to merge.
AI agents change the economics of this process.
Consider a developer who previously produced five diffs per day and now produces ten with coding agents. If each diff still requires the same human review effort, review demand doubles even if the developer's working hours remain constant.
At Meta, the pressure was larger than a simple increase in the number of changes. Significant lines of code per human-landed diff had increased by 105.9%, too. Larger changes can demand more review effort, although lines of code alone are an imperfect measure of complexity.
The result was a capacity mismatch:
- Code production became cheaper and faster.
- Human review capacity remained constrained by attention and calendar time.
- Review queues grew, and timely reviews became less common.
This is a queueing problem as much as a tooling problem. If changes arrive faster than reviewers can process them, the queue grows. Adding more code-generation capacity can make the bottleneck worse.
Meta's approach was to change which changes consume human review capacity.
Routine formatting changes, dead-code removal, mechanical refactoring, and similar low-risk work can often be verified through a combination of deterministic checks and automated semantic inspection. Changes involving security, business logic, or substantial architectural decisions deserve more scrutiny.
The objective is to allocate review effort according to risk rather than treat every diff as equally deserving of human attention.
RADAR is a funnel, not an LLM with merge permissions
RADAR does not simply ask an LLM whether a patch looks safe and then merge whatever it approves.
It uses multiple stages, with different rules depending on who or what produced the change.
Meta already had RACER, its GenAI-powered code-editing and refactoring system. Developers can delegate tasks such as dead-code removal, framework migrations, complexity reduction, lint fixes, and test generation to RACER. The system generates diffs in a sandbox, runs verification, and submits changes for review.
RADAR adds an automated path for qualifying changes to progress through review and landing.
The pipeline combines three important kinds of evidence:
Static eligibility and safety checks. These enforce hard constraints involving the source of a change, its scope, review requirements, and restricted areas of the codebase.
Diff Risk Score (DRS). This machine-learning model estimates the risk associated with a change, primarily its likelihood of contributing to a production incident. Its purpose is to identify changes that should be excluded from automated approval.
Automated Code Review (ACR). An LLM-based review agent examines the actual code changes, looking for safe patterns and signals of defects or risk. It can distinguish a large mechanical refactor from a change that modifies important behavior in ways that require human judgment.
All applicable gates must pass.
A failure at any stage sends the change to the appropriate human-review path. Certain deterministic codemods have a separate, more permissive route after the transformation itself has been approved.
The distinction matters. Static rules can enforce known constraints but cannot understand every semantic consequence of a change. An LLM can interpret code behavior but should not be the only safety mechanism. A risk model can prioritize changes but cannot prove that a particular patch is correct.
RADAR combines these mechanisms instead of asking one component to solve the entire problem.
Risk calibration determines how much work gets automated
The most consequential design decision is the risk threshold.
DRS expresses a change's risk as a percentile. A threshold of P5 admits only the lowest-risk 5% of changes under the relevant scoring distribution. P20 admits the lowest-risk 20%, and P50 admits the lowest-risk 50%.
These are eligibility thresholds, not guarantees that a given percentage of all submitted diffs will be approved. Changes must still satisfy the remaining requirements.
A stricter threshold reduces the number of changes eligible for automation. A more permissive threshold increases potential coverage, but also exposes the system to more risk.
Meta's paper reports an operational policy change from P25 to P50. After calibration, the RADAR approval rate reached 60.31%, while the RADAR Verification pass rate was 26.31%.
Why can the approval rate exceed the risk-score percentile?
Because the two numbers measure different things. P50 describes the risk-score eligibility boundary. The 60.31% figure measures the approval rate among the relevant eligible diffs, according to the paper's metric definition. They are not interchangeable measures of overall codebase coverage.
A simple hypothetical illustrates the underlying economics.
Suppose a team receives 10,000 diffs each month, and each human review consumes 20 minutes of reviewer time.
If 30% of those diffs can safely bypass human review, the system eliminates 3,000 manual reviews:
3,000 diffs x 20 minutes = 60,000 minutes
60,000 / 60 = 1,000 reviewer-hours
That is approximately 125 eight-hour workdays of reviewer capacity per month.
This is an illustrative calculation, not Meta's measured labor savings. It assumes that every qualifying diff would otherwise receive a 20-minute review and that automated review does not create equivalent additional human work elsewhere.
The broader point is that even partial automation can have a meaningful operational effect when applied to a high-volume workflow.
Risk calibration makes this a controllable trade-off. An organization can begin with a conservative threshold, observe production outcomes, and expand automation as evidence accumulates.
Reliability comes from layered controls and operational discipline
The system's safety design goes beyond scoring individual diffs.
For AI-generated changes, RADAR uses the ACE (AI Commit Eligibility) pipeline. Eligible changes must pass safety checks, satisfy the applicable DRS threshold, and pass automated code review.
RACER runbooks receive additional controls because different automated tasks have different operational histories.
Each runbook must demonstrate an acceptable record over a 60-day lookback window, including zero production incidents and low revert and human-rejection rates. The system also imposes per-runbook daily limits, configurable risk thresholds, and denylisting for runbooks or sensitive use cases that should not be allowed to land automatically.
Established runbooks can receive more permissive thresholds and higher volume limits. The paper describes limits reaching 2,000 diffs per day for high-volume use cases.
This is a useful distinction between model confidence and operational confidence.
A model may judge an individual change to be low-risk. The organization must also establish that the process generating that change is trustworthy, that its historical performance supports the decision, and that its volume cannot overwhelm downstream systems.
The human-authored path has its own safeguards. RADAR Verification can allow eligible changes to ship with deferred review. RADAR Approval applies stricter criteria to determine whether that review can be waived entirely. Authors retain control over whether to use the automated path or request human review.
For automated bot changes, the paper also describes a configurable landing delay, during which a human can reject a pending change.
These controls make the system operationally governable. Teams can expand or pause automation, restrict particular sources, and adjust risk policies without treating the LLM as an unquestionable authority.
The resulting architecture is closer to a risk-managed production pipeline than a conventional code-review chatbot.
The production numbers: more throughput, fewer incidents
The paper reports the following outcomes for RADAR.
| Metric | Reported result |
|---|---|
| Diffs reviewed by RADAR | 535,290 |
| Diffs landed by RADAR | 331,720 |
| Peak daily review throughput | 25,000+ diffs |
| RADAR approval rate after calibration | 60.31% |
| RADAR Verification pass rate | 26.31% |
| Revert rate relative to non-RADAR diffs | 1/3 |
| Production-incident rate relative to non-RADAR diffs | 1/50 |
The safety figures deserve careful interpretation.
RADAR-reviewed changes had approximately one-third the revert rate of non-RADAR changes and one-fiftieth the production-incident rate. The paper reports statistically significant differences for these comparisons.
These results are consistent with the system selecting changes that are safer to automate. They do not establish that automated review itself caused the lower incident rate.
Selection is central to the design: RADAR deliberately accepts only changes meeting eligibility requirements and risk criteria. The RADAR population therefore differs from the non-RADAR population before the review process even begins.
That is a feature of the system, but it complicates causal interpretation.
The authors also note that production incidents were manually reviewed by domain experts, and none of the incidents in the reported analysis were judged to have been detected by a human reviewer. That observation is useful context, although it does not establish that human review would never have caught any of those defects.
The right conclusion is that RADAR demonstrated substantial production-scale automation with strong observed safety outcomes. It would be premature to conclude that LLM-based review universally outperforms human review.
Why review latency fell: a queueing and economic perspective
The paper reports two distinct efficiency outcomes for RADAR-handled changes compared with human-reviewed diffs:
- Median time to close was reduced by over 330%, as reported by the authors.
- Median diff review wall time was reduced by 35%.
The first metric measures elapsed time from publication to closure. The second measures time spent waiting for and undergoing review.
These measures capture different bottlenecks. A diff may wait for a reviewer, undergo several rounds of feedback, and then wait again. Automating eligible changes removes much of their dependence on reviewer availability.
There is a terminology caveat worth noting: a conventional percentage reduction cannot exceed 100%. The paper's reported “over 330% reduction” in time to close should therefore be interpreted using the authors' reported comparison rather than as a literal reduction in elapsed time. The 35% reduction in review wall time is more straightforward to interpret.
Consider the queueing dynamics. Let:
-
lambda= the rate at which diffs arrive. -
mu= the rate at which the review system can process them.
When arrival demand approaches or exceeds effective review capacity, waiting times can rise sharply. Increasing reviewer capacity helps, but so does removing eligible changes from the human queue.
RADAR effectively reduces the arrival rate seen by human reviewers:
lambda_human = lambda_total x (1 - p)
Here, p is the fraction of incoming diffs that are safely handled without human review. This is a simplified model: it assumes the automated fraction is stable and ignores differences in review effort between changes.
If 40% of changes qualify, the human queue receives only 60% of the original volume, assuming all else remains constant.
That does not mean total engineering cost falls by exactly 40%. Automation introduces compute, model-inference, validation, and operational costs. Some rejected diffs still need human review. Teams must also monitor incidents and maintain the eligibility policies.
The economic value comes from reallocating expensive human attention. Reviewers can spend more time on architecture, security, business logic, and difficult debugging, rather than routine changes that are amenable to automated verification.
In that sense, RADAR is a form of capacity engineering: it changes the demand placed on a constrained resource rather than relying exclusively on increasing that resource's supply.
What engineering teams can learn from Meta's approach
The paper's findings suggest several practical principles for organizations adopting coding agents.
Start with a narrow automation envelope. Identify recurring changes with clear correctness criteria: mechanical migrations, dead-code removal, formatting, and similarly constrained tasks. Do not begin by allowing unrestricted automated merges.
Separate risk prediction from semantic review. A risk score, static checks, and an LLM-based review agent answer different questions. Combining them provides more useful control than treating any one signal as sufficient.
Evaluate automation sources independently. A proven deterministic transformation and an experimental coding agent should not inherit the same permissions merely because both produce diffs. Track each workflow's incident history, reverts, rejection rates, and volume.
Make risk thresholds operational controls. Start conservatively. Increase coverage only when the measured outcomes justify it, and preserve the ability to disable a source or return changes to human review.
Measure reliability and throughput together. Approval rates and time savings alone can reward an unsafe system. Monitor production incidents, reverts, validation failures, and the share of changes that actually reach production.
Treat the results as organization-specific. Meta operates a large monorepo with extensive testing, established deployment controls, and mature review infrastructure. A smaller organization with weaker tests or less reliable rollback mechanisms should expect different outcomes.
There is also an important limit to the evidence. The paper uses production telemetry, observational comparisons, and a difference-in-differences-style analysis of efficiency outcomes rather than randomized assignment. The authors acknowledge that the findings may not generalize directly to other repositories and engineering environments.
That makes the study valuable as an account of a production system and its measured outcomes, rather than a universal benchmark for the quality of AI code review.
Conclusion: automate the routine, preserve attention for the consequential
Meta's RADAR illustrates a different way to think about AI-assisted software engineering.
The challenge is no longer just generating correct code. It is maintaining a review process that can keep pace with the volume of changes that coding agents produce.
RADAR addresses this by making automated review conditional on risk, source history, eligibility rules, semantic analysis, and deterministic validation. Its production deployment processed more than half a million diffs, landed more than 331,000, and exhibited substantially lower revert and production-incident rates than non-RADAR changes.
The deeper engineering lesson is that reliable automation depends on controlling where it operates, how its permissions expand, and what happens when a change fails a safety gate.
As AI agents become more productive, human review may become more valuable precisely because it can be concentrated on the changes where judgment matters most.
Discussion question: If you were introducing risk-aware code review in your organization, which changes would you allow to land without human review first—and what production evidence would you require before expanding that scope?
Your team's attention is limited, and the deluge of AI-generated code is making it harder to keep production reliable and secure without slowing you down.
I'm building LiveReview, a blast-radius aware AI code review built for your business-critical systems.
Instead of presenting every diff with equal emphasis, LiveReview scores each change by blast radius — how far its impact reaches through your call graph — so you can focus attention where it actually matters.
Spend code review effort where business risk is highest — not spread evenly across every diff.
⭐ Star it on GitHub:
HexmosTech
/
LiveReview
Blast-Radius Aware AI Code Review for Business-Critical Systems
LiveReview: Risk-Aware AI Code Review for Business-Critical Systems
LiveReview is an AI code reviewer that scores every hunk of a diff by blast radius: how far a change reaches through your call graph, how much persistent state it touches, and how well-tested it is. A 3-line change to a shared auth check can outrank a 300-line UI tweak. Your team's attention goes to the highest-risk code first, not spread evenly across every diff.
blast-radius-demo.mp4
LiveReview's Blast Radius & Review Priority scoring, live in the diff viewer.
How does Blast Radius scoring work? (a more technical explanation)
Here's the goal:
- A 3-line fix in a function used by 40 other files, that also writes to a database, should score high.
- A 300-line UI change in one file, fully covered by tests…
Click below to try LiveReview with your codebase:





Top comments (0)