DEV Community

Tess Ainsley
Tess Ainsley

Posted on

Your pull request review time is mostly waiting and rework

Time to merge is the number teams quote when they say AI review sped things up. It bundles at least three separate costs, and reading the diff is only one of them. When the number drops after adding a tool, the drop usually comes from the other two.

Two sets of measurements make the split concrete. CodeRabbit's guide to reviewing AI-generated diffs cites LinearB's 2026 benchmark report: AI-assisted pull requests are 2.6 times larger than unassisted ones, take 4.6 times longer to receive a first review, and have a 30-day acceptance rate of 32.7 percent against 84.4 percent for manual PRs. CodeRabbit is careful to label those as observed associations rather than causal effects, and the association framing is the honest way to read them. The acceptance gap is the part worth sitting with. If roughly two of every three AI-assisted changes are not accepted within 30 days, the cycle spans several passes rather than one review.

The second set is narrower and better instrumented. Obada Kraishan's study of five autonomous coding agents covers 37,623 provenance-labeled PRs across 2,807 repositories, with the pipeline code released for replication. It is a preprint, not peer reviewed. Review effort concentrates unevenly: Copilot PRs drew the most human reviews and change requests, and Claude Code PRs waited a median of 12.6 hours for a first human review. Post-merge outcomes differ by vendor too, with Codex PRs reverted about half as often as human PRs, 6.1 percent against 11.5 percent, while Devin PRs were reverted more often at 14.5 percent.

That 12.6-hour median is the number to keep. Reviewers do not read diffs for twelve hours. The wait is queue time: an agent opens a PR, and nothing happens until a human picks it up. A tool that reads faster than a human shortens the few minutes at the end of that window and leaves the twelve hours at the front of it untouched. Teams that buy review speed by buying reading speed often watch the merge metric barely move, and this is why. The 441 percent increase in code-review time that shows up in the vibe coding literature sits on top of a queue that was already the slow part.

Wait, read, and rework are different numbers

Split the cycle at four timestamps: when the PR opens, when the first review lands, when the first change request lands, and when it is approved and merged. Wait time is the gap between the first pair. Reading time is between the second and third. Rework is everything after a change request, including the author's fix, the re-review, and any further round trip.

Instrument those four per PR, split by provenance and by repository, and the tool decision falls out of the data instead of the demo. If wait dominates, the lever is the trigger. If rework dominates, the lever is whether the tool produces one precise change request or a stream of comments the author answers in four passes. If reading time dominates and your diffs are large, reading speed is finally the right thing to buy.

The rework share is where vendor differences show up most clearly. Copilot PRs drew the most reviews and change requests in the Kraishan dataset, which means more round trips per change, not slower readers. Round trips are also the cost that a review tool can multiply. Ten findings in one pass is one loop. The same ten findings spread across four review cycles is four loops, and each loop carries a context reload, a re-review, and a chance that the author argues with a comment that already stopped being true.

Trigger timing is a review feature

When a review runs matters as much as what it says. A review that fires on PR open puts its feedback into the same window as the wait, so the author and the assigned reviewer see the same first pass. A review that fires after CI, or only when a human asks for it, adds its latency on top of the queue instead of overlapping it. Ask any vendor for the trigger list: open, ready-for-review, CI green, manual comment command. Where the tool sits in that list determines whether it can touch the biggest number in the cycle.

Rules cut round trips only when they fire on the right files

Stack Overflow's post on coding guidelines for AI and people contains the line that ties standards to cycle time. Code review will be most engineers' first look at code they did not write. Heroku chief architect Vish Abrams makes the related point there that principles seasoned engineers assume, like DRY, are not common knowledge to an agent. Rules are supposed to move the standards check earlier so a human reviewer does not spend the first pass on naming and layout, which is exactly the work that generates change requests when it is missed.

A rule only reduces that work if it fires on the right files. A payments rule that fires on the CLI tool produces comments the author has to triage, and triage is a round trip. Kodus imports the rule files teams already keep, including AGENTS.md, CLAUDE.md, .cursorrules, Copilot instruction files, Windsurf rules, and docs/coding-standards, scopes each rule to a path glob, and discovers nested files so a services/billing/CLAUDE.md applies to services/billing/** without extra setup. On self-hosted deployments it writes a per-file trace to the API log under the marker [kody-rules-eval], listing the rule ids selected into the prompt for each reviewed file. The same page documents a limit worth knowing before you plan around it: an unlicensed Community Edition instance evaluates at most 10 rules per review, oldest first, so a rule past the tenth may not fire at all. A rule that silently does not run is one fewer thing checked before a human, and one more thing that turns up later as a comment.

The trace matters beyond debugging. It is the difference between a tool that states your standards load correctly and one where you can point at the rule id and the file it evaluated. Everything else in this section assumes that mechanism works.

What each tool documents about the rework half

The criteria below are the ones that affect round trips: which rule files a tool ingests, whether rules are scoped, and whether you can see what ran on a given file. Unknown means the source I read did not say.

Tool Rule files ingested Path scoping Per-file trace
Kodus AGENTS.md, CLAUDE.md, .cursorrules, Copilot instructions, Windsurf rules, docs/coding-standards Glob per rule, nested files discovered and scoped [kody-rules-eval] log marker per reviewed file on self-hosted
CodeRabbit Repo configuration and custom pre-merge checks, per its review guide Unknown Unknown
GitHub Copilot .github/copilot-instructions.md and nested instruction files, per the patterns Kodus documents as shared conventions Unknown Unknown

For large multi-repo organizations there is a second question that sits under the same heading, which is whether the reviewer can see the dependents of a changed file at all. A change that breaks a caller in another repository is a rework loop nobody can avoid by reading the diff harder, and cross-repo context is the capability that decides it. The same logic that pushes effort toward coupled code applies here, since the codebase matters more than the volume when you decide where review attention goes.

Instrument before you buy

Run the four timestamps for a month before you evaluate anything. Report wait, reading, and rework separately, split by provenance and by repository, and you will know which of the three is eating the cycle. Then ask each vendor which of the three its product moves: the trigger window, the number of change requests per merged PR, or the time a human spends reading. A tool that can answer for the first two is reducing cycle time. A tool that can only answer for the third is reducing the smallest number on most teams, and its benchmark will still look impressive, because reading speed is measurable while a queue nobody instrumented stays invisible.

Claims checked 2026-09-29.

Top comments (0)