DEV Community

Cole Halton
Cole Halton

Posted on

Agent PRs revert at 6.1% or 14.5%, depending on the vendor

If you have ever tried to answer "is agent-written code worse?" with one number, the new preprint from Obada Kraishan at Texas Tech is the most useful thing I have read on it this month, mostly because it refuses to hand you that number.

Not All Agents Are Equal: Code Quality and Post-Merge Maintenance Across Five Autonomous Coding Agents in the Wild follows 37,623 provenance-labeled pull requests across 2,807 GitHub repositories from December 2024 to July 2025. Of those, 33,596 were opened by five commercial agents (OpenAI Codex, Devin, GitHub Copilot, Cursor, Claude Code) and 4,027 by a matched human baseline in the same repositories over the same window. The pipeline enriches them with 58,792 cached GitHub API responses and then follows each merged change for 90 days.

Every PR keeps its vendor label, so the analysis can separate "agents" from individual products. That distinction carries most of the result.

What the pipeline actually measures

The comparison runs on three fronts. Patch-level quality covers 8,933 PRs and 1,348,822 added lines, scoring security smells across eight CWE classes for Python, JavaScript and TypeScript, plus maintainability signals like TODO/FIXME density, over-long lines, comment ratio, branching and nesting depth. Post-merge maintenance covers 26,283 merged PRs over the full 90-day window, measuring size-normalized churn and revert detection. On top of that it counts human and bot reviews and change requests per PR.

The design is observational. Nothing was randomized, and the repositories are public repos with more than 100 stars, so this is a slice of open-source work rather than your internal monorepo. Keep both of those in mind before you paste a number into a slide.

Reverts split by vendor, and the split is the story

Within 90 days of merge, the human baseline reverted 11.5% of its PRs. Here is how the agents did:

  • OpenAI Codex: 6.1% (n = 17,756, odds ratio 0.50, 95% CI 0.44 to 0.57)
  • Devin: 14.5% (n = 2,185, OR 1.31, CI 1.11 to 1.54)
  • GitHub Copilot: 12.5% (n = 2,094, OR 1.10, CI 0.93 to 1.31)
  • Cursor: 11.4% (n = 946, OR 1.00, CI 0.79 to 1.25)
  • Claude Code: 10.5% (n = 267, OR 0.90, CI 0.60 to 1.36)

Pool those and you get a number that describes almost nothing. The spread inside "agent" is wider than the gap between the best agent and the humans. Any comparison that treats "AI code" as a single category is averaging over a 2.4x difference in revert rate, and the average lands somewhere no individual tool actually sits.

Read the sample sizes before the percentages

Codex is 17,756 of the 33,596 agent PRs, so the pooled agent baseline is mostly Codex behavior wearing a trench coat. Claude Code's 10.5% rests on 267 PRs, and its interval runs from 0.60 to 1.36. The honest reading of that row is "no detectable difference from humans," and "Claude Code is safer" is not supported by 267 observations. Same story for Copilot (interval crosses 1, p = .457) and Cursor (OR exactly 1.00). Devin and Codex are the two rows whose intervals clear the baseline.

This is the failure mode I keep running into in reviewer benchmarks: a headline percentage pulled from a slice too thin to carry it. Per-turn instruction compliance also decays the longer an agent session runs, and session length is not controlled here, so a vendor whose users run longer sessions gets penalized or flattered by a variable nobody measured. Same class of problem as reviewing a patch without ever looking at execution.

The security-smell result is real but narrow

Pooled agent code was less likely than human code to contain a security smell (OR 0.63), driven by fewer hardcoded credentials and eval-style constructs. A size-stratified check puts the whole effect in the largest PRs (XL bucket, delta -0.08, p = .025) with no difference in the four smaller buckets. The honest version is: on very large diffs, the agents hardcode fewer secrets than humans do. Worth knowing, and considerably smaller than "agent code is more secure."

Review effort is where your volume problem shows up

Copilot PRs drew the most human reviews and change requests. Claude Code PRs waited the longest for a first human review, median 12.6 hours, and they are also the biggest by far: median 495 changed lines, median nesting four levels deep, and the highest branch density at .063 per line. Bigger PRs wait longer. That matches the queue data I keep seeing, and it is the part of this paper a team can act on today, because it is a routing problem you can measure with the tools you already have.

What to instrument in your own repo

You do not need 2,807 repositories. You need provenance and a fixed window.

  1. Label PRs by author type at creation time (agent, human, mixed). Retrofitting this later is guesswork.
  2. Pick a fixed slice, say 90 days of merged PRs, and keep the pool constant so the mix does not drift under you.
  3. Track four numbers: revert rate, size-normalized churn, first-review latency, and change requests per PR.
  4. Split every one of them by agent, not by "AI." The paper's own table is the argument for that.
  5. Record session length or PR size as a covariate, or you cannot tell a quality difference from a routing difference.

The pipeline code, statistical reports and figures are released for replication, so you can diff your methodology against theirs before you trust your own numbers.

Picking a reviewer for agent-authored PRs

No reviewer tool fixes the measurement problem, but a few properties make it possible. The question is whether the tool can see provenance and repo history, and whether you can run it where your code lives. CodeRabbit, Greptile and Qodo are hosted-first with self-hosting gated behind enterprise contracts, which is fine if you are already on those contracts. Kodus is open source and self-hostable on a standard plan with your own model keys and org-wide rules, which is the property that matters when the agent PRs live in a private repo. The tool directory scores 27 of these against the same nine standards, so compare on documented capability first and then verify with your own numbers. If you want a sense of how a review bot reads a diff versus how the code actually runs, that gap is its own post.

Where this changes my own measurement

Vendor identity explains more of the post-merge outcome than agent-versus-human does. That sounds like a boring sentence until you try to answer "how do we review the growing volume of AI-generated code" and realize the answer depends on which agent wrote the diff, how large it is, and how long it sat in the queue. Measure those in your own history before you buy anything, and treat any benchmark that reports one number for "AI" as unfinished.

Top comments (0)