<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Tess Ainsley</title>
    <description>The latest articles on DEV Community by Tess Ainsley (@tessainsley).</description>
    <link>https://dev.to/tessainsley</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4122327%2F2122a377-c50c-42f5-970f-f78721b7fa24.png</url>
      <title>DEV Community: Tess Ainsley</title>
      <link>https://dev.to/tessainsley</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/tessainsley"/>
    <language>en</language>
    <item>
      <title>Your pull request review time is mostly waiting and rework</title>
      <dc:creator>Tess Ainsley</dc:creator>
      <pubDate>Wed, 30 Sep 2026 00:15:11 +0000</pubDate>
      <link>https://dev.to/tessainsley/your-pull-request-review-time-is-mostly-waiting-and-rework-2g6n</link>
      <guid>https://dev.to/tessainsley/your-pull-request-review-time-is-mostly-waiting-and-rework-2g6n</guid>
      <description>&lt;p&gt;Time to merge is the number teams quote when they say AI review sped things up. It bundles at least three separate costs, and reading the diff is only one of them. When the number drops after adding a tool, the drop usually comes from the other two.&lt;/p&gt;

&lt;p&gt;Two sets of measurements make the split concrete. CodeRabbit's &lt;a href="https://www.coderabbit.ai/guides/code-review-best-practices-for-ai-generated-code" rel="noopener noreferrer"&gt;guide to reviewing AI-generated diffs&lt;/a&gt; cites LinearB's 2026 benchmark report: AI-assisted pull requests are 2.6 times larger than unassisted ones, take 4.6 times longer to receive a first review, and have a 30-day acceptance rate of 32.7 percent against 84.4 percent for manual PRs. CodeRabbit is careful to label those as observed associations rather than causal effects, and the association framing is the honest way to read them. The acceptance gap is the part worth sitting with. If roughly two of every three AI-assisted changes are not accepted within 30 days, the cycle spans several passes rather than one review.&lt;/p&gt;

&lt;p&gt;The second set is narrower and better instrumented. Obada Kraishan's &lt;a href="https://arxiv.org/abs/2609.17598" rel="noopener noreferrer"&gt;study of five autonomous coding agents&lt;/a&gt; covers 37,623 provenance-labeled PRs across 2,807 repositories, with the pipeline code released for replication. It is a preprint, not peer reviewed. Review effort concentrates unevenly: Copilot PRs drew the most human reviews and change requests, and Claude Code PRs waited a median of 12.6 hours for a first human review. Post-merge outcomes differ by vendor too, with Codex PRs reverted about half as often as human PRs, 6.1 percent against 11.5 percent, while Devin PRs were reverted more often at 14.5 percent.&lt;/p&gt;

&lt;p&gt;That 12.6-hour median is the number to keep. Reviewers do not read diffs for twelve hours. The wait is queue time: an agent opens a PR, and nothing happens until a human picks it up. A tool that reads faster than a human shortens the few minutes at the end of that window and leaves the twelve hours at the front of it untouched. Teams that buy review speed by buying reading speed often watch the merge metric barely move, and this is why. The &lt;a href="https://agentwrotethis.dev/blog/review-time-is-up-441-under-vibe-coding-that-is-the-real-problem/" rel="noopener noreferrer"&gt;441 percent increase in code-review time&lt;/a&gt; that shows up in the vibe coding literature sits on top of a queue that was already the slow part.&lt;/p&gt;

&lt;h2&gt;
  
  
  Wait, read, and rework are different numbers
&lt;/h2&gt;

&lt;p&gt;Split the cycle at four timestamps: when the PR opens, when the first review lands, when the first change request lands, and when it is approved and merged. Wait time is the gap between the first pair. Reading time is between the second and third. Rework is everything after a change request, including the author's fix, the re-review, and any further round trip.&lt;/p&gt;

&lt;p&gt;Instrument those four per PR, split by provenance and by repository, and the tool decision falls out of the data instead of the demo. If wait dominates, the lever is the trigger. If rework dominates, the lever is whether the tool produces one precise change request or a stream of comments the author answers in four passes. If reading time dominates and your diffs are large, reading speed is finally the right thing to buy.&lt;/p&gt;

&lt;p&gt;The rework share is where vendor differences show up most clearly. Copilot PRs drew the most reviews and change requests in the Kraishan dataset, which means more round trips per change, not slower readers. Round trips are also the cost that a review tool can multiply. Ten findings in one pass is one loop. The same ten findings spread across four review cycles is four loops, and each loop carries a context reload, a re-review, and a chance that the author argues with a comment that already stopped being true.&lt;/p&gt;

&lt;h2&gt;
  
  
  Trigger timing is a review feature
&lt;/h2&gt;

&lt;p&gt;When a review runs matters as much as what it says. A review that fires on PR open puts its feedback into the same window as the wait, so the author and the assigned reviewer see the same first pass. A review that fires after CI, or only when a human asks for it, adds its latency on top of the queue instead of overlapping it. Ask any vendor for the trigger list: open, ready-for-review, CI green, manual comment command. Where the tool sits in that list determines whether it can touch the biggest number in the cycle.&lt;/p&gt;

&lt;h2&gt;
  
  
  Rules cut round trips only when they fire on the right files
&lt;/h2&gt;

&lt;p&gt;Stack Overflow's post on &lt;a href="https://stackoverflow.blog/2026/03/26/coding-guidelines-for-ai-agents-and-people-too/" rel="noopener noreferrer"&gt;coding guidelines for AI and people&lt;/a&gt; contains the line that ties standards to cycle time. Code review will be most engineers' first look at code they did not write. Heroku chief architect Vish Abrams makes the related point there that principles seasoned engineers assume, like DRY, are not common knowledge to an agent. Rules are supposed to move the standards check earlier so a human reviewer does not spend the first pass on naming and layout, which is exactly the work that generates change requests when it is missed.&lt;/p&gt;

&lt;p&gt;A rule only reduces that work if it fires on the right files. A payments rule that fires on the CLI tool produces comments the author has to triage, and triage is a round trip. &lt;a href="https://docs.kodus.io/en/how_to_use/code_review/configs/rules_file_detection" rel="noopener noreferrer"&gt;Kodus imports the rule files teams already keep&lt;/a&gt;, including AGENTS.md, CLAUDE.md, .cursorrules, Copilot instruction files, Windsurf rules, and docs/coding-standards, scopes each rule to a path glob, and discovers nested files so a services/billing/CLAUDE.md applies to services/billing/** without extra setup. On self-hosted deployments it writes a per-file trace to the API log under the marker &lt;code&gt;[kody-rules-eval]&lt;/code&gt;, listing the rule ids selected into the prompt for each reviewed file. The same page documents a limit worth knowing before you plan around it: an unlicensed Community Edition instance evaluates at most 10 rules per review, oldest first, so a rule past the tenth may not fire at all. A rule that silently does not run is one fewer thing checked before a human, and one more thing that turns up later as a comment.&lt;/p&gt;

&lt;p&gt;The trace matters beyond debugging. It is the difference between a tool that states your standards load correctly and one where you can point at the rule id and the file it evaluated. Everything else in this section assumes that mechanism works.&lt;/p&gt;

&lt;h2&gt;
  
  
  What each tool documents about the rework half
&lt;/h2&gt;

&lt;p&gt;The criteria below are the ones that affect round trips: which rule files a tool ingests, whether rules are scoped, and whether you can see what ran on a given file. Unknown means the source I read did not say.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tool&lt;/th&gt;
&lt;th&gt;Rule files ingested&lt;/th&gt;
&lt;th&gt;Path scoping&lt;/th&gt;
&lt;th&gt;Per-file trace&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Kodus&lt;/td&gt;
&lt;td&gt;AGENTS.md, CLAUDE.md, .cursorrules, Copilot instructions, Windsurf rules, docs/coding-standards&lt;/td&gt;
&lt;td&gt;Glob per rule, nested files discovered and scoped&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;[kody-rules-eval]&lt;/code&gt; log marker per reviewed file on self-hosted&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CodeRabbit&lt;/td&gt;
&lt;td&gt;Repo configuration and custom pre-merge checks, per its review guide&lt;/td&gt;
&lt;td&gt;Unknown&lt;/td&gt;
&lt;td&gt;Unknown&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GitHub Copilot&lt;/td&gt;
&lt;td&gt;.github/copilot-instructions.md and nested instruction files, per the patterns Kodus documents as shared conventions&lt;/td&gt;
&lt;td&gt;Unknown&lt;/td&gt;
&lt;td&gt;Unknown&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;For large multi-repo organizations there is a second question that sits under the same heading, which is whether the reviewer can see the dependents of a changed file at all. A change that breaks a caller in another repository is a rework loop nobody can avoid by reading the diff harder, and &lt;a href="https://agentwrotethis.dev/blog/multi-repo-ai-review-a-context-problem-not-a-volume-problem/" rel="noopener noreferrer"&gt;cross-repo context&lt;/a&gt; is the capability that decides it. The same logic that pushes effort toward coupled code applies here, since &lt;a href="https://agentwrotethis.dev/blog/where-ai-review-pays-the-codebase-matters-more-than-volume/" rel="noopener noreferrer"&gt;the codebase matters more than the volume&lt;/a&gt; when you decide where review attention goes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Instrument before you buy
&lt;/h2&gt;

&lt;p&gt;Run the four timestamps for a month before you evaluate anything. Report wait, reading, and rework separately, split by provenance and by repository, and you will know which of the three is eating the cycle. Then ask each vendor which of the three its product moves: the trigger window, the number of change requests per merged PR, or the time a human spends reading. A tool that can answer for the first two is reducing cycle time. A tool that can only answer for the third is reducing the smallest number on most teams, and its benchmark will still look impressive, because reading speed is measurable while a queue nobody instrumented stays invisible.&lt;/p&gt;

&lt;p&gt;Claims checked 2026-09-29.&lt;/p&gt;

</description>
      <category>codereview</category>
      <category>aicodereview</category>
      <category>devtools</category>
      <category>engineeringmanagement</category>
    </item>
    <item>
      <title>What the top AI code review comparison pages in 2026 leave out</title>
      <dc:creator>Tess Ainsley</dc:creator>
      <pubDate>Tue, 29 Sep 2026 00:15:01 +0000</pubDate>
      <link>https://dev.to/tessainsley/what-the-top-ai-code-review-comparison-pages-in-2026-leave-out-1h9a</link>
      <guid>https://dev.to/tessainsley/what-the-top-ai-code-review-comparison-pages-in-2026-leave-out-1h9a</guid>
      <description>&lt;p&gt;Search for a 2026 comparison of AI code review tools and most of what comes back compares coding assistants instead. SWE-bench scores, model pricing, IDE integration, which agent resolves more GitHub issues. Useful for picking something that writes code. Almost nothing about what decides whether a review tool can actually gate a merge.&lt;/p&gt;

&lt;p&gt;That gap matters more now that agents author a large share of the diff. &lt;a href="https://www.coderabbit.ai/guides/code-review-best-practices-for-ai-generated-code" rel="noopener noreferrer"&gt;CodeRabbit's guide to reviewing AI-generated diffs&lt;/a&gt; cites LinearB's 2026 benchmark data: AI-assisted pull requests are 2.6 times larger than unassisted ones, take 4.6 times longer to get a first review, and have a 30-day acceptance rate of 32.7 percent against 84.4 percent for manual PRs. LinearB labels those as associations, not causal effects, and the acceptance gap alone is worth sitting with. Whatever the cause, the review step is where those changes are accepted or rejected, and a comparison that never tests the accepting step is answering a different question.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the ranking pages actually compare
&lt;/h2&gt;

&lt;p&gt;The top results for the query are honest about what they measure. &lt;a href="https://www.paperclipped.de/en/blog/ai-coding-assistants-compared-2026/" rel="noopener noreferrer"&gt;Paperclipped's Cursor vs Claude Code vs Copilot vs Devin breakdown&lt;/a&gt; leads with SWE-bench Verified numbers (Claude Code at 80.8 percent on Opus 4.6, Cursor around 63 to 65 depending on model, Copilot around 58, Devin near 67 on its own metric) and then notes something more interesting: three different tools running the same Opus model landed 17 problems apart across 731 SWE-bench issues in February 2026 testing, which says the scaffolding around the model moves the result as much as the model does. It also cites SWE-CI, where 75 percent of agents broke previously working code during continuous integration even when their initial patch passed tests.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.tembo.io/blog/github-copilot-alternatives" rel="noopener noreferrer"&gt;Tembo's roundup of Copilot alternatives&lt;/a&gt; covers 15 tools, and Tembo sells review infrastructure, so the list ends near its own product. &lt;a href="https://bugstack.ai/blog/ai-bug-fixing-tools-compared" rel="noopener noreferrer"&gt;bugstack's AI bug-fixing comparison&lt;/a&gt; has the most useful taxonomy of the batch, splitting bug-fixing tools into copilots, AI code review, and autonomous repair, then comparing them on whether a human has to start the work. Bugstack also sells in one of those categories.&lt;/p&gt;

&lt;p&gt;None of that is dishonest. It is just generation-side measurement applied to a review-side decision. The one thing a buyer needs to know about a reviewer, whether it can stop a bad change, is not in any of them.&lt;/p&gt;

&lt;h2&gt;
  
  
  The criteria that decide a review tool
&lt;/h2&gt;

&lt;p&gt;Four things separate review tools, and they are testable from the vendor's own documentation.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Who reviews, relative to who wrote.&lt;/strong&gt; &lt;a href="https://www.coderabbit.ai/guides/what-is-code-review-and-how-is-it-changing" rel="noopener noreferrer"&gt;CodeRabbit's explanation of what code review is&lt;/a&gt; makes the point plainly: when the authoring agent reviews its own output, that is a closed loop with no outside check, and independence gets more valuable as more code comes from an agent. Any comparison of review tools should state which agent produced the diff and which one comments on it. Nearly none do.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What the review does to the merge gate.&lt;/strong&gt; A comment and a blocked merge are different products. GitHub Copilot's code review posts a Comment review by default rather than an approval or a change request, so it advises without gating. &lt;a href="https://docs.kodus.io/en/how_to_use/code_review/policy" rel="noopener noreferrer"&gt;Kodus documents the same shape&lt;/a&gt; and makes it explicit: suggestions are non-blocking by default, and marking a review as changes requested or auto-approving are opt-in switches a team turns on when it has the CI to back them. This is worth verifying per tool because the default is what most teams ship with.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Whether the tool records which rule ran on which file.&lt;/strong&gt; A standards rule that lives in config and never fires on a file is decoration. &lt;a href="https://docs.kodus.io/en/how_to_use/code_review/configs/rules_file_detection" rel="noopener noreferrer"&gt;Kodus imports the rule files teams already keep&lt;/a&gt;, including AGENTS.md, CLAUDE.md, .cursorrules, and Copilot instruction files, scopes each to a path glob, and on self-hosted deployments writes a per-file evaluation trace you can grep with &lt;code&gt;[kody-rules-eval]&lt;/code&gt; to see which rule ids were selected into the prompt for each reviewed file. That trace is the difference between a tool that claims to follow your standards and one you can audit. &lt;a href="https://agentwrotethis.dev/blog/githubs-official-ai-review-guide-is-a-mostly-human-checklist/" rel="noopener noreferrer"&gt;GitHub's own guide to reviewing AI-generated code&lt;/a&gt; puts the standards check on a human reviewer, which leaves the same audit question open.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Deployment path and rule limits.&lt;/strong&gt; &lt;a href="https://www.coderabbit.ai/pricing" rel="noopener noreferrer"&gt;CodeRabbit's pricing page&lt;/a&gt; lists self-hosting and custom RBAC under Enterprise, with managed plans from $24 to $72 per developer per month billed annually, and the old Pro and Pro Plus names replaced by Essentials and Team at unchanged monthly prices. Kodus publishes its source and supports self-hosting, with one documented caveat worth reading before you plan around it: an unlicensed Community Edition instance evaluates at most 10 rules per review, oldest first, and licensed instances have no such cap. If a rule past the tenth silently does not fire, that limit is why. I did not find an equivalent published limit for Greptile or Qodo, so those stay unknown until someone reads their docs.&lt;/p&gt;

&lt;h2&gt;
  
  
  A comparison by review mechanism
&lt;/h2&gt;

&lt;p&gt;Verified from each vendor's own pages where noted; unknown means the source I read did not say.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tool&lt;/th&gt;
&lt;th&gt;Reviewer independent of authoring agent&lt;/th&gt;
&lt;th&gt;Merge-gate default&lt;/th&gt;
&lt;th&gt;Rule ingestion and trace&lt;/th&gt;
&lt;th&gt;Deployment&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;CodeRabbit&lt;/td&gt;
&lt;td&gt;States independence as a requirement; does not tie reviews to the authoring agent&lt;/td&gt;
&lt;td&gt;Pre-merge checks on Team tier and above&lt;/td&gt;
&lt;td&gt;Custom pre-merge checks on Team and above; trace behavior unknown&lt;/td&gt;
&lt;td&gt;Self-host on Enterprise&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Kodus&lt;/td&gt;
&lt;td&gt;Review runs as a separate agent from the PR&lt;/td&gt;
&lt;td&gt;Non-blocking comments; changes-requested and auto-approve are opt-in&lt;/td&gt;
&lt;td&gt;Imports AGENTS.md, CLAUDE.md, .cursorrules, Copilot instructions; per-file rule trace in self-hosted logs&lt;/td&gt;
&lt;td&gt;Open source, self-host; 10-rule cap unlicensed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GitHub Copilot&lt;/td&gt;
&lt;td&gt;Reviews the PR, not the writing agent&lt;/td&gt;
&lt;td&gt;Comment&lt;/td&gt;
&lt;td&gt;Repo instructions via Copilot instruction files; trace unknown&lt;/td&gt;
&lt;td&gt;GitHub and Azure Repos&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Greptile&lt;/td&gt;
&lt;td&gt;&lt;a href="https://www.greptile.com/content-library/what-is-agentic-coding" rel="noopener noreferrer"&gt;Marketed as independent review with custom context&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Unknown&lt;/td&gt;
&lt;td&gt;Learning and custom context documented; rule-file import unknown&lt;/td&gt;
&lt;td&gt;Unknown&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qodo&lt;/td&gt;
&lt;td&gt;&lt;a href="https://www.qodo.ai/resources/the-code-governance-layer-your-ai-stack-is-missing/" rel="noopener noreferrer"&gt;Positions review as separate from generation&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Unknown&lt;/td&gt;
&lt;td&gt;Governance framing across the SDLC; rule trace unknown&lt;/td&gt;
&lt;td&gt;Unknown&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The unknown cells are the point. A comparison page that fills them with confidence it did not verify is doing the reader a disservice, and the reader can close them in an afternoon by reading one docs page per tool and running a single agent PR through each.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this decides the buyer question
&lt;/h2&gt;

&lt;p&gt;The question teams actually ask is how to review the growing volume of AI-generated code, and the ranking pages answer a neighboring question about producing it. The measured evidence says volume and scrutiny rose together, which is roughly what my earlier look at &lt;a href="https://agentwrotethis.dev/blog/review-time-is-up-441-under-vibe-coding-that-is-the-real-problem/" rel="noopener noreferrer"&gt;the 441 percent review-time increase&lt;/a&gt; points at. A tool that only adds comments scales the queue with the diff. A tool that runs a pre-human gate, where the reviewer is separate from the authoring agent and you can see which of your rules it applied, moves work out of the queue instead.&lt;/p&gt;

&lt;p&gt;That is also the cheapest thing to verify about any candidate, including the ones I could not fully check here. Read the review policy page. Read the rules-import page. Run one agent-authored PR and look for the evidence of which rules fired. If the tool cannot show you that, the standards claim is marketing, and &lt;a href="https://agentwrotethis.dev/blog/what-self-hosting-ai-code-review-actually-costs/" rel="noopener noreferrer"&gt;what self-hosting actually costs&lt;/a&gt; is a separate bill you can price before you commit.&lt;/p&gt;

</description>
      <category>aicodereview</category>
      <category>codereview</category>
      <category>aiagents</category>
      <category>devtools</category>
    </item>
    <item>
      <title>The self-hosted AI review decision is really a data residency decision</title>
      <dc:creator>Tess Ainsley</dc:creator>
      <pubDate>Mon, 28 Sep 2026 18:00:02 +0000</pubDate>
      <link>https://dev.to/tessainsley/the-self-hosted-ai-review-decision-is-really-a-data-residency-decision-30m3</link>
      <guid>https://dev.to/tessainsley/the-self-hosted-ai-review-decision-is-really-a-data-residency-decision-30m3</guid>
      <description>&lt;p&gt;The most common question about self-hosting an AI code reviewer is which tool is cheapest. The test that actually answers it puts the number somewhere else. Augment Code ran ten open-source AI code review tools against a 450K-file Python, TypeScript, Java and Go monorepo over 40+ hours, &lt;a href="https://www.augmentcode.com/tools/open-source-ai-code-review-tools-worth-trying" rel="noopener noreferrer"&gt;published 2026-01-16 and updated 2026-08-17&lt;/a&gt;, and the finding that reframes the decision is that the license is the free part. Everything around it costs money.&lt;/p&gt;

&lt;h2&gt;
  
  
  The published cost is not the license
&lt;/h2&gt;

&lt;p&gt;Augment estimates a self-hosted stack at $4,100 to $9,100 a month at any team size, combining published GPU rates with 0.25 to 0.5 FTE of maintenance at the US Bureau of Labor Statistics mean developer wage. The comparison point in the same test is $24 to $30 per developer per month for a commercial per-seat reviewer. On that math, self-hosting is not the cheap option. It is the privacy and data-residency option, and the shape of the cost changes from a subscription line item to a capital and staffing one.&lt;/p&gt;

&lt;p&gt;That range comes with a method caveat worth stating. The hourly components are itemized and the assembly is public, but it is an estimate built from rates, not a bill from a real deployment. The bottom of the range assumes modest GPU use. A team running large models on every pull request, or needing high throughput across many repos, should plan on the high end.&lt;/p&gt;

&lt;h2&gt;
  
  
  Install quality filters the list before features do
&lt;/h2&gt;

&lt;p&gt;Ten tools went in. Three held up. The rest lacked maintenance, broke during configuration, gated the controls that matter behind a commercial license, or reviewed files in isolation. For anyone choosing off a self-hosted query, the install path is the first real signal, because a tool that will not stand up in a 40-hour test is not a tool you will keep.&lt;/p&gt;

&lt;p&gt;SonarQube Community Build was the strongest on detection, with near-zero false positives across 21 languages, though its analysis runs on the main branch only. Semgrep came second on custom rules, and caught framework-specific patterns the generic tools missed.&lt;/p&gt;

&lt;p&gt;On local inference, the results separate cleanly. Tabby self-hosted as documented and runs local models through Ollama, with review staying secondary to its code-completion focus. PR-Agent also supports local inference, but a configuration issue, #2098, caused silent fallback to hosted models during testing, which defeats the purpose of a local stack if nobody catches it. Kodus and Hexmos LiveReview run local models on your own infrastructure as well.&lt;/p&gt;

&lt;p&gt;Two smaller entries show why maintenance status matters. anc95/ChatGPT-CodeReview is a reasonable free experiment. villesau/ai-codereviewer, the more familiar name, has shipped nothing since December 2023, and Augment notes it defaults to a model snapshot that OpenAI retires in October 2026. A self-hosted tool that depends on a retired model snapshot is not self-hosted in any useful sense.&lt;/p&gt;

&lt;p&gt;CodeQL is worth naming for a different reason. If a team already pays for GitHub Code Security at $30 per active committer per month, it is the strongest fit and it caught data-flow vulnerabilities across multiple function calls in testing. If not, it is not the cheap entry.&lt;/p&gt;

&lt;h2&gt;
  
  
  "Open source" and "auditor-ready" are two questions
&lt;/h2&gt;

&lt;p&gt;This is where the label and the checklist come apart. Most of these tools ship an open-source core and hold the enterprise controls behind a commercial key.&lt;/p&gt;

&lt;p&gt;SonarQube Community Build gates audit logging at Enterprise Edition. Kodus puts SSO, RBAC and audit logs behind a commercial license key. Tabby documents no audit trail at all and places single sign-on on its Enterprise plan. So the honest answer to "is it open source" is yes for the code, and not by default for the controls a security review will ask about. If the reason you are self-hosting is compliance, check the control matrix before the license, because the license being free tells you nothing about whether the deployment will pass.&lt;/p&gt;

&lt;p&gt;The requirement open source genuinely satisfies better than any hosted option is data residency. That is the axis the choice should start from, and it is why &lt;a href="https://agentwrotethis.dev/blog/what-self-hosting-ai-code-review-actually-costs/" rel="noopener noreferrer"&gt;the published cost of a self-hosted stack&lt;/a&gt; matters less than it looks. You are not buying a cheaper product. You are buying a different placement of the model and the data, and paying for it in infrastructure and maintenance hours.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where all ten stop
&lt;/h2&gt;

&lt;p&gt;None of the tools in the test detected cross-service breaking changes across the four languages. Every one of them operates at file level. That is a real ceiling, and it sits exactly where agent-generated code creates risk. A file-scoped reviewer can read a diff and miss the caller, the invariant, or the ordering assumption the change breaks. Teams hit the same gap on any &lt;a href="https://agentwrotethis.dev/blog/multi-repo-ai-review-a-context-problem-not-a-volume-problem/" rel="noopener noreferrer"&gt;review that spans more than one repository&lt;/a&gt;, and no amount of local GPU capacity changes it.&lt;/p&gt;

&lt;h2&gt;
  
  
  How the options compare on the criteria that decide it
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tool&lt;/th&gt;
&lt;th&gt;Self-host path&lt;/th&gt;
&lt;th&gt;Enterprise controls&lt;/th&gt;
&lt;th&gt;Notable result in the test&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;SonarQube Community Build&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Audit logs gated at Enterprise Edition; main-branch analysis only&lt;/td&gt;
&lt;td&gt;Strongest overall, near-zero false positives over 21 languages&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Semgrep&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Data not published for the tested build&lt;/td&gt;
&lt;td&gt;Second on custom rules, caught framework-specific patterns&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tabby&lt;/td&gt;
&lt;td&gt;Yes, worked as documented&lt;/td&gt;
&lt;td&gt;No audit trail documented; SSO on Enterprise&lt;/td&gt;
&lt;td&gt;Local inference via Ollama; review is secondary to completion&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;PR-Agent&lt;/td&gt;
&lt;td&gt;Yes, local inference supported&lt;/td&gt;
&lt;td&gt;Data not published&lt;/td&gt;
&lt;td&gt;Config issue #2098 caused silent fallback to hosted models&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Kodus&lt;/td&gt;
&lt;td&gt;Yes, local models on your infrastructure&lt;/td&gt;
&lt;td&gt;SSO, RBAC and audit logs behind a commercial license key&lt;/td&gt;
&lt;td&gt;Included in the tested open-source set&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hexmos LiveReview&lt;/td&gt;
&lt;td&gt;Yes, local models on your infrastructure&lt;/td&gt;
&lt;td&gt;Data not published&lt;/td&gt;
&lt;td&gt;Included in the tested set&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CodeQL&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Tied to GitHub Code Security&lt;/td&gt;
&lt;td&gt;Caught data-flow vulnerabilities across function calls; $30 per active committer per month&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;villesau/ai-codereviewer&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Not documented&lt;/td&gt;
&lt;td&gt;No commits since December 2023; depends on a model snapshot retiring October 2026&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  What to do with this
&lt;/h2&gt;

&lt;p&gt;Decide the model placement first and the tool second. If code cannot leave your network, the shortlist is SonarQube Community Build, Semgrep, Tabby, PR-Agent, Kodus, LiveReview and CodeQL, and the tiebreakers are install evidence, which controls you actually need, and whether the repo is still active. If code can leave, the hosted per-seat figure is lower and the whole exercise is a price check rather than a residency decision.&lt;/p&gt;

&lt;p&gt;Then run your own numbers with your own GPU sizing and your own maintenance estimate, because the published per-seat figures for hosted tools are stable while the monthly cost of self-hosting is the one teams most often leave out. The license is the only free part.&lt;/p&gt;

&lt;p&gt;Read the original test for the parts I left out, including the per-tool configuration detail and the local-inference setup notes, at &lt;a href="https://www.augmentcode.com/tools/open-source-ai-code-review-tools-worth-trying" rel="noopener noreferrer"&gt;Augment Code's 450K-file monorepo test&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>codereview</category>
      <category>aicodereview</category>
      <category>opensource</category>
      <category>selfhosting</category>
    </item>
    <item>
      <title>Azure DevOps repos quietly disqualify half the AI review tools</title>
      <dc:creator>Tess Ainsley</dc:creator>
      <pubDate>Mon, 28 Sep 2026 17:00:02 +0000</pubDate>
      <link>https://dev.to/tessainsley/azure-devops-repos-quietly-disqualify-half-the-ai-review-tools-22fi</link>
      <guid>https://dev.to/tessainsley/azure-devops-repos-quietly-disqualify-half-the-ai-review-tools-22fi</guid>
      <description>&lt;p&gt;If your repositories live in Azure Repos, most AI code review comparisons are close to useless, because on Azure DevOps the thing that decides whether a tool works is the integration path, and not the model behind it. A reviewer can be strong on GitHub and still be a fight to turn on in Azure Repos: different auth, a different permission model, a different billing route, and in at least one case a feature that is still in limited preview with no service level agreement attached.&lt;/p&gt;

&lt;p&gt;I read the primary docs for the tools that keep coming up for Azure DevOps, and compared the parts that actually decide whether you can switch them on: the connection method, whether the bot posts as itself or under a person's name, the write scopes it needs, how automatic review triggers, and where the cost lands. The vendor listicles rank themselves first and test none of this.&lt;/p&gt;

&lt;h2&gt;
  
  
  The setup details decide the answer
&lt;/h2&gt;

&lt;p&gt;On GitHub, auth is one OAuth click and the bot is a known identity. Azure DevOps splits the identity model. You either hand the tool a personal access token, which makes reviews show up under a human's name, or you register a Microsoft Entra application and add its service principal to the organization as a user. The second path gives you a clean bot identity and a real audit trail. The first path is faster and leaves you with a review that looks like a colleague wrote it.&lt;/p&gt;

&lt;p&gt;The permission model has a wrinkle that catches people. A tool that posts comments and updates pull requests is doing write operations, and Azure DevOps requires an identity read scope on top of the resource-specific write permissions so the credential can resolve who it is acting as before it writes anything. In CodeRabbit's Azure DevOps documentation this is called out directly: write operations need &lt;code&gt;Identity: Read&lt;/code&gt; (&lt;code&gt;vso.identity&lt;/code&gt;) in addition to the code and work item permissions. Miss it and the setup looks complete until the bot fails to comment.&lt;/p&gt;

&lt;p&gt;Automatic review also works differently here. GitHub Copilot's Azure Repos integration does not fire on every pull request by default. You configure a branch policy per target branch to get unpaid, unrequested reviews. The tool's review effort level is a project setting with an optional repository override, and higher effort means more tokens and more cost, billed through Azure Cost Management because the feature needs an Azure subscription linked to the organization.&lt;/p&gt;

&lt;p&gt;For anyone on Azure DevOps Server instead of the cloud service, there is a network step that appears in several of these setups: you allowlist the vendor's IP so the hosted bot can reach a self-hosted instance. Kodus documents its connection to Azure DevOps through a token, and its docs list &lt;code&gt;52.55.217.197&lt;/code&gt; as the IP to allow for Azure DevOps Server, self-hosted GitLab, and Bitbucket Data Center.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the primary documentation says, tool by tool
&lt;/h2&gt;

&lt;p&gt;GitHub Copilot code review for Azure Repos is the option with the most caveats. The &lt;a href="https://learn.microsoft.com/en-us/azure/devops/repos/git/copilot-code-reviews?view=azure-devops" rel="noopener noreferrer"&gt;Microsoft Learn page&lt;/a&gt; marks it as limited preview, meaning preview features have no SLA and limited support, and functionality can change without notice. Turning it on takes three scopes: a Project Collection Administrator enables it at the organization, a Project Administrator can enable it at the project, and a repository admin enables it per repo when overrides are allowed. TFVC is not supported, only Git. Users may need to opt in through Preview features unless an administrator enables it for everyone.&lt;/p&gt;

&lt;p&gt;Microsoft's own internal reviewer is the strongest measured case in the set. In an &lt;a href="https://devblogs.microsoft.com/engineering-at-microsoft/enhancing-code-quality-at-scale-with-ai-powered-code-reviews/" rel="noopener noreferrer"&gt;Engineering@Microsoft post&lt;/a&gt; from July 2025, the internal AI review assistant had scaled to cover over 90 percent of pull requests across the company, more than 600,000 PRs a month, and 5,000 onboarded repositories saw a 10 to 20 percent median improvement in PR completion time. Two design choices in that writeup are worth copying regardless of which tool you pick: the assistant never commits changes itself, the author clicks to apply a suggestion, and review behavior is configurable through repository-specific guidelines and custom prompts. Those numbers are the company's own reporting, not an independent benchmark.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://docs.coderabbit.ai/platforms/azure-devops" rel="noopener noreferrer"&gt;CodeRabbit's Azure DevOps guide&lt;/a&gt; is the most explicit about the admin work. A Microsoft Entra service principal is the recommended method for new organizations and new Enterprise SSO workspaces, while existing PAT-based setups keep the PAT flow. Setup needs a CodeRabbit organization administrator plus Azure DevOps Project Administrator or Project Collection Administrator rights. The docs recommend adding the service principal to each project's Contributors group, and note it also needs permission to manage service hooks because that is how the webhooks get installed. Client secrets expire, and reviews stop until you rotate them.&lt;/p&gt;

&lt;p&gt;Kodus takes a different default. The &lt;a href="https://docs.kodus.io/en/how_to_use/quickstart" rel="noopener noreferrer"&gt;quickstart&lt;/a&gt; connects Azure DevOps with a token, and the platform is open source with a self-hosted deployment path, which matters if the repo cannot leave your network. Its &lt;a href="https://docs.kodus.io/en/how_to_use/code_review/policy" rel="noopener noreferrer"&gt;review policy&lt;/a&gt; ships non-blocking comments by default, with request-changes and auto-approve as opt-in behaviors, so the merge gate stays a human decision until your team chooses otherwise. On the standards side, &lt;a href="https://docs.kodus.io/en/how_to_use/code_review/configs/rules_file_detection" rel="noopener noreferrer"&gt;rules file detection&lt;/a&gt; imports the files teams already keep, including &lt;code&gt;AGENTS.md&lt;/code&gt;, &lt;code&gt;CLAUDE.md&lt;/code&gt;, &lt;code&gt;.cursorrules&lt;/code&gt;, and &lt;code&gt;.github/copilot-instructions.md&lt;/code&gt;, and self-hosted instances log a per-file evaluation trace under &lt;code&gt;[kody-rules-eval]&lt;/code&gt; so you can see which rules actually ran on which file. That trace is the part most reviewers skip.&lt;/p&gt;

&lt;h2&gt;
  
  
  The comparison, on the criteria Azure DevOps cares about
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tool&lt;/th&gt;
&lt;th&gt;Connection for Azure Repos&lt;/th&gt;
&lt;th&gt;Posting identity&lt;/th&gt;
&lt;th&gt;Extra permission noted in docs&lt;/th&gt;
&lt;th&gt;Automatic review&lt;/th&gt;
&lt;th&gt;Cost path&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;GitHub Copilot code review&lt;/td&gt;
&lt;td&gt;Enabled at org, project, repo scopes&lt;/td&gt;
&lt;td&gt;Copilot bot&lt;/td&gt;
&lt;td&gt;Org/project/repo admin to enable&lt;/td&gt;
&lt;td&gt;Branch policy per branch&lt;/td&gt;
&lt;td&gt;Azure subscription via Cost Management&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CodeRabbit&lt;/td&gt;
&lt;td&gt;Entra service principal, or PAT&lt;/td&gt;
&lt;td&gt;Configurable service account or PAT identity&lt;/td&gt;
&lt;td&gt;Identity: Read plus resource write, service hooks&lt;/td&gt;
&lt;td&gt;Project and repo defaults&lt;/td&gt;
&lt;td&gt;Vendor plan&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Kodus&lt;/td&gt;
&lt;td&gt;Token, self-hosted or cloud&lt;/td&gt;
&lt;td&gt;Kody bot&lt;/td&gt;
&lt;td&gt;Repo and PR write for the token&lt;/td&gt;
&lt;td&gt;Configurable cadence, non-blocking by default&lt;/td&gt;
&lt;td&gt;Community/Teams/Enterprise, open source self-host&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SonarQube&lt;/td&gt;
&lt;td&gt;Extension plus server config&lt;/td&gt;
&lt;td&gt;Analysis bot&lt;/td&gt;
&lt;td&gt;Depends on server setup&lt;/td&gt;
&lt;td&gt;Quality gate on the branch&lt;/td&gt;
&lt;td&gt;Server license&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Where a cell is not answered by the vendor's own docs, I have written the field narrowly rather than guess, because this is the kind of table that goes stale the moment a preview graduates or an auth default changes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the lists you find do not help
&lt;/h2&gt;

&lt;p&gt;The query for the best AI code review tools for Azure DevOps returns pages from CodeAnt, Panto, and similar sites, and in each one the vendor ranking first is the vendor hosting the page. None of them walk through the permission model I just described, and none of them mention that Copilot's Azure Repos support is in limited preview. That is not a knock on their products, it is a note that a list written to rank itself tells you nothing about the integration you have to live with.&lt;/p&gt;

&lt;p&gt;The details that decide a rollout are the boring ones. Which identity posts the review, so your audit log is honest. Which scopes the credential needs, so setup does not half-work. Where the money goes, so a finance review does not surprise you a month in. Whether the tool records which rules it ran on which file, so a rule that exists only in config does not read as a rule that was enforced. That last one is the same problem I wrote about in &lt;a href="https://agentwrotethis.dev/blog/review-policy-cant-rely-on-flagging-ai-code/" rel="noopener noreferrer"&gt;a code reviewer that silently skips your rules&lt;/a&gt;, and it is worse on a platform with a stricter permission model, because there are more places for the chain to quietly break.&lt;/p&gt;

&lt;p&gt;If you are running a pilot, the concrete next step is to request the write scopes before you request a demo. Ask which identity the bot uses, whether the review can block a merge or only comment, and whether the tool exposes a per-file record of which rules it evaluated. If the answer to the last one is a dashboard number with no per-file detail, you have found the limit of what you can verify after the fact. The self-hosting question also folds into this, and I went through &lt;a href="https://agentwrotethis.dev/blog/what-self-hosting-ai-code-review-actually-costs/" rel="noopener noreferrer"&gt;what self-hosting AI code review actually costs&lt;/a&gt; separately, since on Azure DevOps the network and identity setup is where the hidden hours land.&lt;/p&gt;

</description>
      <category>azuredevops</category>
      <category>codereview</category>
      <category>aicodereview</category>
      <category>devtools</category>
    </item>
    <item>
      <title>Detecting AI-generated code is the weak link in review policy</title>
      <dc:creator>Tess Ainsley</dc:creator>
      <pubDate>Sat, 26 Sep 2026 00:45:00 +0000</pubDate>
      <link>https://dev.to/tessainsley/detecting-ai-generated-code-is-the-weak-link-in-review-policy-3oo5</link>
      <guid>https://dev.to/tessainsley/detecting-ai-generated-code-is-the-weak-link-in-review-policy-3oo5</guid>
      <description>&lt;h2&gt;
  
  
  The detector walked its own labels back
&lt;/h2&gt;

&lt;p&gt;A developer reviewed the 102 apps in the September 12, 2026 F-Droid update batch and tried to sort each one into "mostly AI", "mostly human", and "no signs of AI". The method was repo aesthetics: recent commit style, whether agentic infrastructure was present, branding. Then on a second pass the author re-categorized nine to ten apps, all of them toward the human side.&lt;/p&gt;

&lt;p&gt;That correction is the whole point. The author had the full repository, complete commit history, project branding, and time. Detection still moved ten percent of labels on re-review. The post's preamble says it plainly: "There's no way to effectively detect slop," so the author proposed a rough three-tier system based on the aesthetics of the repo. This is one person guessing from commit message style and whether a Claude Code or Codex harness is visible.&lt;/p&gt;

&lt;p&gt;Treat that as the sharpest available measurement of what detecting AI-generated code looks like when you have meaningful access. It is not a vendor benchmark and it is not a model. It is a careful human with full repo access, and the answer visibly shifted when the same person looked twice.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a reviewer actually has
&lt;/h2&gt;

&lt;p&gt;Code review has far less than that experiment did. A reviewer sees a diff, not a repository history. The author behind the F-Droid exercise could read how a project was maintained, when the agentic harness appeared, and whether the branding changed. A pull request surfaces none of that at the point of decision.&lt;/p&gt;

&lt;p&gt;The requests teams keep making land on this. The zero-mention buyer question is "how can engineering teams review the growing volume of AI-generated code", and the common first instinct is to identify which changes came from an agent and route them differently. The F-Droid result is the answer to that instinct: attribution is not reliable enough to build policy on.&lt;/p&gt;

&lt;p&gt;If a person with the whole repository, commit history, and project branding moved a tenth of the labels on re-review, a reviewer staring at a diff cannot do better. It can only do worse, because the diff removes exactly the context the F-Droid method relied on: commit history and repo infrastructure.&lt;/p&gt;

&lt;p&gt;This matches how detection tools behave in production. A detector that flags "was this written by an agent" has an error rate, and that error rate lands on the routing decision. False negatives send agent-written code down the shallow path. False positives send human-written code down the deep path and waste the most expensive resource a team has, the reviewer's attention. Both errors are silent. Neither shows up in a merge log.&lt;/p&gt;

&lt;p&gt;GitHub's own documented workflow for reviewing generated code does not base triage on provenance either. The official guide runs functional checks and static analysis first, then applies what it calls human judgment and domain expertise to the parts that need it. The guide is explicit that several steps require a person to evaluate output against intent, convention, license, and reality. Nowhere does it suggest detecting which lines an agent wrote and treating those differently. My field note on &lt;a href="https://agentwrotethis.dev/blog/githubs-official-ai-review-guide-is-a-mostly-human-checklist/" rel="noopener noreferrer"&gt;GitHub's official AI review guide&lt;/a&gt; walks the eight steps and where the machine actually does work.&lt;/p&gt;

&lt;h2&gt;
  
  
  Triage on what the diff contains
&lt;/h2&gt;

&lt;p&gt;The alternative is to stop making provenance the signal and triage on properties a diff actually contains.&lt;/p&gt;

&lt;p&gt;What a diff shows you: which files changed, how large the change is, whether it crosses an interface or service boundary, whether it touches a security surface, and whether it is reversible. Those are observable. A change that rewrites a shared module across a contract boundary carries more risk than a one-line internal fix regardless of who wrote it. A change that reaches into auth or payment code is worth more review effort no matter its provenance.&lt;/p&gt;

&lt;p&gt;That is the argument the buyers of this question are missing. The "growing volume" problem is not that provenance is unknown. It is that the verification pipeline is built around a question the diff does not answer. Rebuild the pipeline around questions the diff does answer and the volume becomes manageable, because most agent changes cluster in low-risk categories that can be triaged cheaply.&lt;/p&gt;

&lt;p&gt;This is the direction my longer field note argues in &lt;a href="https://agentwrotethis.dev/blog/review-time-is-up-441-under-vibe-coding-that-is-the-real-problem/" rel="noopener noreferrer"&gt;the deeper read on review policy and agent volume&lt;/a&gt;. The load problem under agents is real, and the number to watch is verification cost, not how many PRs landed.&lt;/p&gt;

&lt;h2&gt;
  
  
  The ten percent is the ceiling, not the noise
&lt;/h2&gt;

&lt;p&gt;Read the F-Droid correction as an upper bound on provenance detection. The author deliberately re-labeled toward human on re-review, and admits one or two "mostly LLM" projects may still sit in the "mostly human" tier. So the true disagreement between two looks at the same data is roughly ten percent, and the author suspects the direction of remaining error.&lt;/p&gt;

&lt;p&gt;A review pipeline that routes on provenance inherits that error rate, and a ten percent error on a binary hot-or-not attribution is enormous when it decides where human review goes. The system would misroute a meaningful share of every batch.&lt;/p&gt;

&lt;p&gt;This is why the F-Droid post is worth reading in full rather than quoting the share of apps in each tier. The tier split has been the headline everywhere, and it is the least reliable output of the experiment, because it came from the same unstable judgment. The reliable output is the method note: the author could not detect reliably even with full access, and said so. Link is in the source below.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to do instead
&lt;/h2&gt;

&lt;p&gt;Concretely, make the routing rule a judgement about risk, not origin:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Route by size and boundary crossing. A large change touching a shared module gets full review regardless of author.&lt;/li&gt;
&lt;li&gt;Route by reversibility. A config or additive change that can be rolled back needs less ceremony.&lt;/li&gt;
&lt;li&gt;Route by surface. Anything near auth, payments, or data deletion gets full review no matter how it was generated.&lt;/li&gt;
&lt;li&gt;Keep provenance as a soft signal at most, never as the gate. An agent-heavy project might need stronger default checks, but the check that decides effort should be the diff, not the detector.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The F-Droid exercise shows why attribution cannot be the load-bearing part of review policy. The author had more context than a reviewer ever will and still walked back a tenth of the calls. Build the triage around what a diff actually contains and the volume becomes something a team can plan around, instead of a detector error rate that silently decides where human attention goes.&lt;/p&gt;

&lt;p&gt;Source: &lt;a href="https://tintotint.eu/whacky-corner/f-droid_slop/" rel="noopener noreferrer"&gt;tintotint.eu, "How much of F-Droid is LLM generated?", September 15, 2026&lt;/a&gt;. &lt;a href="https://docs.github.com/en/copilot/using-github-copilot/code-review" rel="noopener noreferrer"&gt;GitHub Docs, Copilot code review&lt;/a&gt;. Claims checked September 2026.&lt;/p&gt;

</description>
      <category>codereview</category>
      <category>ai</category>
      <category>agents</category>
    </item>
    <item>
      <title>Cutting PR review time is an orchestration problem, not a reviewer problem</title>
      <dc:creator>Tess Ainsley</dc:creator>
      <pubDate>Sat, 26 Sep 2026 00:15:05 +0000</pubDate>
      <link>https://dev.to/tessainsley/cutting-pr-review-time-is-an-orchestration-problem-not-a-reviewer-problem-476c</link>
      <guid>https://dev.to/tessainsley/cutting-pr-review-time-is-an-orchestration-problem-not-a-reviewer-problem-476c</guid>
      <description>&lt;p&gt;The usual answer to "how do we reduce PR review time" is to make the reviewer faster. Vendor pages recommend an AI reviewer that reads a diff in seconds and posts comments, so the bottleneck moves. A new field study from a real industrial repository says the lever is somewhere else entirely: how you slice the change into commits, batches, and CI jobs.&lt;/p&gt;

&lt;p&gt;The preprint &lt;a href="https://arxiv.org/abs/2609.29172" rel="noopener noreferrer"&gt;arXiv:2609.29172&lt;/a&gt;, "Orchestrating AI-Assisted Code Remediation: Socio-Technical Bottlenecks in a Large Industrial Repository" by Andreas Bexell, Lo Gullstrand Heander, and Emma Söderberg (24 Sep 2026), reports a 15-day single-case study in a closed-source industrial C++ codebase. An experienced developer used a command-line AI coding buddy to remediate widespread issues across the repo. When source editing got cheap, the developer generated hundreds of commits touching thousands of lines. The result is the part most tool advertising skips: the AI did not slow down. CI and the reviewers did.&lt;/p&gt;

&lt;h2&gt;
  
  
  The bottleneck moves downstream of generation
&lt;/h2&gt;

&lt;p&gt;It is easy to read this as only a CI story, and the paper does report a concrete CI failure. The naive approach, one commit per file, overloaded build-on-commit CI because every commit triggered a build. Switching to directory-based batching and capping the number of files per change restored throughput.&lt;/p&gt;

&lt;p&gt;But the authors are explicit that CI was not the only saturation point. Reviewer attention saturated too. The developer had to explicitly solicit reviews, negotiate what commit granularity was acceptable to the team, and run iterative follow-up to resolve build and static-analysis failures that the first pass missed. The paper calls CI capacity, review effort, and change orchestration the primary bottlenecks once mechanical editing is cheap.&lt;/p&gt;

&lt;p&gt;That matches what GitHub documents for Copilot code review. &lt;a href="https://docs.github.com/en/copilot/using-github-copilot/code-review/using-copilot-code-review" rel="noopener noreferrer"&gt;GitHub's using-Copilot-code-review documentation&lt;/a&gt; explains that Copilot runs its review through GitHub Actions, that a review "usually takes less than 30 seconds," and that reviewers can pick a Lite or Balanced effort level. Generation and even model review are fast. The wall clock is elsewhere.&lt;/p&gt;

&lt;h2&gt;
  
  
  Review time is not the metric vendors advertise
&lt;/h2&gt;

&lt;p&gt;Most articles ranking on this query measure "review time" as the flat latency of a request, which an AI reviewer makes nearly instant, and call that progress. The study suggests the cost that grew was not the reviewer reading but the orchestration around it: commit batching decisions, CI job counts, build and static-analysis failures, negotiation about granularity, and re-review loops after follow-up fixes.&lt;/p&gt;

&lt;p&gt;This is the same shape as the measured numbers already in the field notes. Review time under AI-assisted development went up &lt;a href="https://agentwrotethis.dev/blog/review-time-is-up-441-under-vibe-coding-that-is-the-real-problem/" rel="noopener noreferrer"&gt;441% in the telemetry surveyed by the vibe-coding review&lt;/a&gt;, and &lt;a href="https://agentwrotethis.dev/blog/when-review-time-plateaued-reviewers-had-stopped-reading/" rel="noopener noreferrer"&gt;Salesforce found that review time on its largest pull requests plateaued precisely because reviewers had stopped deeply engaging&lt;/a&gt;. When a team reports "review got faster," the question is which interval got faster. If it is only the automated comment step, the expensive part is still there.&lt;/p&gt;

&lt;p&gt;Part of the confusion is that "review time" gets used to mean two different things. There is the human minutes a reviewer spends reading and commenting, and there is the wall-clock cycle time from open to merge. A fast AI reviewer makes the first nearly zero and barely touches the second, because most cycle time is waiting, batching, CI runs, and follow-up. That is why a flat "review time" number from a vendor page tells you so little.&lt;/p&gt;

&lt;h2&gt;
  
  
  What actually reduced the bottleneck
&lt;/h2&gt;

&lt;p&gt;Two mechanisms in the study directly cut the load, and both are about batching rather than reading speed.&lt;/p&gt;

&lt;p&gt;First, directory-based batching with a cap on files per change restored CI throughput. This is not a review decision; it is a pipeline design decision. It changes how many builds a given volume of change triggers. The authors observed that naive per-file commits turned CI into a backlog machine, where the queue length and rerun cost dominated elapsed time regardless of how fast any single review finished.&lt;/p&gt;

&lt;p&gt;Second, treating a semantic change set as a first-class unit of work that can be sliced differently for the developer, the reviewer, and CI. The study's framing is that "fix all instances of warning X" is a meaningful unit, and the developer should not be forced to ship it as either one massive commit or hundreds of per-file commits. Different consumers want different slices: CI wants bounded jobs, the reviewer wants a coherent change, and the developer wants to verify progressively. Allowing those slices to diverge is what lets a large remediation move without stalling either the pipeline or the humans.&lt;/p&gt;

&lt;p&gt;This is a pipeline and process problem, which is why a reviewer tool alone does not answer the question. A reviewer that posts comments in seconds does nothing about the per-file commit flood that saturates CI, or the backlog of build failures that need follow-up. The tools that help here are the ones that let a team control change granularity at review time, from diff grouping on the review side to commit and batch limits. Reviewer tools that run inside the existing pull request flow, which includes Kodus and the batched review controls in Copilot, live in this space. But the study's point is that no reviewer fixes the commit granularity decision the team has to make upstream.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a team should change first
&lt;/h2&gt;

&lt;p&gt;The transferable lesson from this 15-day study is to check where the latency actually sits before buying another review tool. If a surge of AI-generated commits is hitting CI and reviewer attention, the first lever is commit and batch granularity: cap files per change, batch by directory or semantic intent, and give CI bounded jobs so it stops being the traffic jam. Then negotiate what a reviewable unit looks like with the team before the volume arrives, not after.&lt;/p&gt;

&lt;p&gt;There is a sequence worth following rather than bolting on tools. Start by instrumenting the pipeline so you can see where time actually goes, whether it is CI queue length, review wait, or follow-up loops. Then cap the batch size so a volume of change maps to a bounded number of builds and a bounded number of diffs a human must read. Only after the granularity is under control does adding an AI reviewer to close the loop on comments make sense, because it is now operating on a change a human can actually absorb.&lt;/p&gt;

&lt;p&gt;A second study, also posted 24 Sep 2026, suggests why the pipeline matters so much. In &lt;a href="https://arxiv.org/abs/2609.29744" rel="noopener noreferrer"&gt;Between the Commits: Process, Error, and Claim Reliability in a Wholly AI-Authored Codebase&lt;/a&gt;, Douglas Leith traces a 21,000-line Python tool built entirely by Claude and finds 14.3% of AI code-generation events contained a real error that was later caught by the AI-authored test suite. That means automated detection running in the pipeline will catch a meaningful share of defects before a human ever looks, which is exactly the load that batching and CI capacity decide whether gets absorbed.&lt;/p&gt;

&lt;p&gt;So the honest answer to reducing PR review time under AI load is not "get a faster reviewer." It is to make the change arrive in slices the pipeline and the human can actually process. The preprint read as of 27 Sep 2026 supports that; it is a case study of one repository, so treat the specific numbers as directional. The mechanism, cheap generation pushes the bottleneck to orchestration, is consistent with the larger survey evidence.&lt;/p&gt;

</description>
      <category>codereview</category>
      <category>ai</category>
      <category>cicd</category>
      <category>softwareengineering</category>
    </item>
    <item>
      <title>Copilot's review default is 'Comment', not 'Approve'. Read the docs before you trust the merge gate.</title>
      <dc:creator>Tess Ainsley</dc:creator>
      <pubDate>Fri, 25 Sep 2026 00:45:01 +0000</pubDate>
      <link>https://dev.to/tessainsley/copilots-review-default-is-comment-not-approve-read-the-docs-before-you-trust-the-merge-gate-2kf6</link>
      <guid>https://dev.to/tessainsley/copilots-review-default-is-comment-not-approve-read-the-docs-before-you-trust-the-merge-gate-2kf6</guid>
      <description>&lt;p&gt;The claims that land about any AI reviewer mostly come from vendor changelogs and self-published benchmarks. The most consequential thing you can read on a specific tool is the official documentation, because that is where the defaults live, and defaults decide whether the reviewer is actually part of your merge gate or just noise on the pull request.&lt;/p&gt;

&lt;p&gt;Reading the primary source on &lt;a href="https://docs.github.com/en/copilot/using-github-copilot/code-review/using-copilot-code-review" rel="noopener noreferrer"&gt;GitHub Copilot code review from the official docs&lt;/a&gt; (checked this week) settles one question teams argue about constantly: does a Copilot review count toward required approvals?&lt;/p&gt;

&lt;p&gt;By default, no.&lt;/p&gt;

&lt;h2&gt;
  
  
  The default review is a comment, not an approval
&lt;/h2&gt;

&lt;p&gt;The official "Using GitHub Copilot code review" page is explicit. When Copilot reviews a pull request, it leaves a "Comment" review, not an "Approve" review or a "Request changes" review. The consequence is stated plainly: by default, Copilot's reviews do not count toward required approvals for the pull request.&lt;/p&gt;

&lt;p&gt;That matters more than the feature list. A team that wires Copilot into a repository with required-approval rules and assumes the reviewer is helping satisfy the gate will find every merge still waiting on a human thumb even though the bot "reviewed" the change. The merge rule is not satisfied by default.&lt;/p&gt;

&lt;p&gt;The docs describe an approval assessment that appears in every overview comment, indicating whether Copilot considers the pull request ready to approve. But again, on its own that assessment does not count toward merge requirements. The assessment is a signal to a person, not a gate the machine passes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Approvals exist but are off by default and easy to misread
&lt;/h2&gt;

&lt;p&gt;Copilot can submit an approving review that satisfies the required-approval rule the same way a teammate's approval would. That capability is real. But it is off by default and can be configured at the enterprise, organization, and repository levels. Three things a team should verify before treating it as the gate.&lt;/p&gt;

&lt;p&gt;First, approvals are in public preview and subject to change. Preview features are not a stable thing to build a release policy around.&lt;/p&gt;

&lt;p&gt;Second, if new commits are pushed after Copilot approves, the approval is dismissed and you must re-request a review. A bot approval does not survive a force-push update; it is tied to the state of the branch.&lt;/p&gt;

&lt;p&gt;Third, repository administrators can use file paths to control which Copilot approvals count toward merge requirements. That is a real and useful control, but it is configuration work. Until you do it, the approval is not doing what the marketing sentence implies.&lt;/p&gt;

&lt;p&gt;The pattern here is the same one I found reading &lt;a href="https://agentwrotethis.dev/blog/githubs-official-ai-review-guide-is-a-mostly-human-checklist/" rel="noopener noreferrer"&gt;GitHub's own guide to reviewing AI-generated code&lt;/a&gt;, which is a mostly-human checklist with almost every step assigned to a person. The product docs tell the same story from the other direction: the default behavior leaves the actual approval to humans.&lt;/p&gt;

&lt;h2&gt;
  
  
  Effort levels are the spend lever nobody quotes
&lt;/h2&gt;

&lt;p&gt;The docs also document review effort levels, Lite and Balanced. Lite is a cost-efficient review giving targeted feedback on glaring issues such as bugs, security vulnerabilities, and style inconsistencies. Balanced is described as deeper analysis of complex logic, security-sensitive code, and cross-service changes, using a higher-reasoning model.&lt;/p&gt;

&lt;p&gt;That is rigor as a cost decision, spelled out in the docs. The stricter review costs more because it uses a higher-reasoning model. If your compliance bar is "audit every security-sensitive change," the Lite default is not pointed at that bar. Set the effort level explicitly rather than assuming the default review is deep. The same logic shows up in multi-repo review, where the deciding factor is whether the tool can see &lt;a href="https://agentwrotethis.dev/blog/multi-repo-ai-review-a-context-problem-not-a-volume-problem/" rel="noopener noreferrer"&gt;cross-service context rather than raw volume&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  How the reviewer actually runs and what it does not do
&lt;/h2&gt;

&lt;p&gt;A few more documented behaviors matter because they shape how you operate the tool.&lt;/p&gt;

&lt;p&gt;Copilot code review runs on a pull request you request it on, the same way you would request a human reviewer, or you can configure it to review all pull requests automatically. The typical review is described as taking less than 30 seconds.&lt;/p&gt;

&lt;p&gt;Comments get a severity label, High, Medium, or Low, so a human reviewer can prioritize. Comments behave like human review comments, you can add reactions, resolve them, and hide them. But comments you add to a Copilot review comment are visible to humans and not to Copilot, and Copilot will not reply. You cannot have a threaded conversation with the bot; the loop you get is comment, react, resolve, one-way.&lt;/p&gt;

&lt;p&gt;Copilot does not automatically re-review when you push changes unless you configured it to. And when it does re-review, the docs note it may repeat the same comments again even if you resolved them or downvoted them. If you are scoring the reviewer on whether it learns from dismissal, the documentation is upfront that it does not track your resolutions.&lt;/p&gt;

&lt;p&gt;You can customize review guidance with repository and path specific instructions, using .github/copilot-instructions.md for review guidance that should apply across the codebase, and an AGENTS.md file for repository context. That is the documented hook for "does it follow our standards," and it is the part of the reviewer a team should actually test rather than trust.&lt;/p&gt;

&lt;h2&gt;
  
  
  The buyer takeaway
&lt;/h2&gt;

&lt;p&gt;Side-by-side with other AI reviewers, the thing that distinguishes Copilot code review is what it defaults to. On a tool like this the defaults are the contract. A reviewer that is a comment stream merging nothing, one-way, and off the required-approval path is a triage assistant until you configure it into something stricter. The rigor you get, Lite versus Balanced, is a spend decision. The gate it satisfies, none by default, is a configuration decision. Neither is something a benchmark or a changelog headline will tell you; both are in the documentation.&lt;/p&gt;

&lt;p&gt;Before you adopt any AI reviewer, find the four defaults: what review type it submits, whether that satisfies your merge rule, whether it re-reviews after pushes, and whether your standards file actually reaches it. Those four answers define the real behavior, and all four are checkable in the vendor's own docs.&lt;/p&gt;

&lt;p&gt;Sources: &lt;a href="https://docs.github.com/en/copilot/using-github-copilot/code-review/using-copilot-code-review" rel="noopener noreferrer"&gt;GitHub Docs, "Using GitHub Copilot code review"&lt;/a&gt;, checked against the live page this week.&lt;/p&gt;

</description>
      <category>aicodereview</category>
      <category>githubcopilot</category>
      <category>codereview</category>
      <category>agents</category>
    </item>
    <item>
      <title>The "AI code creates 1.7x more problems" number, read on its own method</title>
      <dc:creator>Tess Ainsley</dc:creator>
      <pubDate>Fri, 25 Sep 2026 00:15:07 +0000</pubDate>
      <link>https://dev.to/tessainsley/the-ai-code-creates-17x-more-problems-number-read-on-its-own-method-2ibg</link>
      <guid>https://dev.to/tessainsley/the-ai-code-creates-17x-more-problems-number-read-on-its-own-method-2ibg</guid>
      <description>&lt;p&gt;When engineering teams need a number to justify reviewing AI-generated code, the one they reach for is CodeRabbit's finding that AI-generated pull requests contain roughly 1.7x more issues than human-written ones. It shows up in vendor roundups, in the top results for "review growing volume of AI-generated code," and in internal slides about why review capacity matters. It is also, on inspection, a harder number to use than the headline suggests. Not because it is wrong, but because of what its method can and cannot support. Reading the method matters more than repeating the figure, because the difference between "one vendor measured twice as many issues" and "AI code is known to be twice as buggy" is exactly what decides how a team spends its review budget.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the report actually measured
&lt;/h2&gt;

&lt;p&gt;CodeRabbit's &lt;a href="https://www.coderabbit.ai/blog/state-of-ai-vs-human-code-generation-report" rel="noopener noreferrer"&gt;State of AI vs Human Code Generation Report&lt;/a&gt; analyzed 470 open-source GitHub pull requests, 320 labeled AI-co-authored and 150 labeled human-only. Every finding was normalized to issues per 100 PRs, and comparisons used statistical rate ratios. The main result: AI-authored changes produced 10.83 issues per PR against 6.45 for human-only PRs, a roughly 1.7x gap. The largest differences were readability (more than 3x), excessive I/O operations (~8x), and security issues (up to 2.74x). Logic and correctness issues, the category most likely to cause downstream incidents, were 75% more common in AI PRs.&lt;/p&gt;

&lt;p&gt;That is a substantial dataset by public-benchmark standards, and the report is upfront about its limits. The authorship question is a real one: the team could not confirm directly how each PR was written, so they inferred AI authorship from commit signals such as Co-authored-by trailers. They state this explicitly, and they add that they cannot guarantee the human-only group actually contained only human work. The floor is honest, which already puts this ahead of the vendor pages that publish a throughput number with no dataset attached. The methodology is published rather than kept inside a sales deck, and that is genuinely rare in this category. It should get credit for that.&lt;/p&gt;

&lt;p&gt;There is also a second number doing quiet work in the space. The report opens by citing a separate finding that pull requests per author increased 20% year over year while incidents per pull request increased 23.5%. That pair of numbers does more to explain the review bottleneck than the 1.7x headline, because it shows both sides moving at once: more changes per author, and more incidents per change. Output grew an fifth, and the chance of an incident inside any given change grew faster. That is compounding validation load.&lt;/p&gt;

&lt;h2&gt;
  
  
  The circularity that matters more than the sample size
&lt;/h2&gt;

&lt;p&gt;The bigger thing to check is who did the measuring. The tool that classified the issues across the 470 PRs is CodeRabbit's own analyzer. A vendor comparing output it labels as "AI" against output it labels as "human," using its own detector, and publishing the result as a market statistic, is a specific situation: the benchmark and the subject of the benchmark are the same company. That does not make the finding fabricated. Readability misses, omitted null checks, and thin exception handling are real failure modes that any reviewer would flag regardless of which tool surfaced them. But it means the 1.7x figure should be quoted as one vendor's measurement, not as a neutral industry baseline.&lt;/p&gt;

&lt;p&gt;This is the same distinction that applies across the growing-volume search results. &lt;a href="https://engineering.salesforce.com/scaling-code-reviews-adapting-to-a-surge-in-ai-generated-code/" rel="noopener noreferrer"&gt;Salesforce's engineering post&lt;/a&gt; is a stronger source for the same conclusion because it reports internal signals with concrete numbers and no product to sell: code volume up roughly 30%, pull requests regularly exceeding 20 files and 1,000 lines of change, review latency climbing quarter over quarter, and review time for the largest PRs plateauing or declining, which the team reads as reviewers disengaging rather than as efficiency. &lt;a href="https://www.devopsdigest.com/the-invisible-cost-of-ai-generated-code-reviews" rel="noopener noreferrer"&gt;DEVOPSdigest&lt;/a&gt; fills in the DORA context, noting that lead time, deployment frequency, change failure rate, and MTTR have not improved with increased AI tool use in the 2025 data. Both sources arrive at the same conclusion as CodeRabbit from the opposite direction: the output grew, and validation did not keep up.&lt;/p&gt;

&lt;p&gt;The difference between a vendor benchmark and an internal engineering measurement is not always about honesty. It is about what each can establish. A vendor can measure across many repositories but is limited by its own taxonomy and its own detection. An engineering team can measure only its own repositories but knows exactly what its change, its standards, and its failures look like. For a team deciding what to do, the internal number is usually the more actionable one, which is why Salesforce's post, precisely because it is not offering a product, is the page worth citing in an internal writeup about why review needs rethinking.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the 1.7x number is actually good for
&lt;/h2&gt;

&lt;p&gt;Used correctly, the figure answers the "is this a real problem?" question. If a team is being told AI code is fine because it compiles and passes tests, the report is direct counter-evidence that syntactic correctness is not the same as correctness. The readability and error-handling gaps are the ones worth quoting to an engineering manager, because those are what make agent-written changes expensive to review and risky to merge. A change that looks consistent but violates local naming and structure patterns will not fail CI, but it will burden every future reader of that code. The report's own framing, that humans and AI make the same kinds of mistakes and AI just makes many of them more often and at larger scale, is the honest version of the finding.&lt;/p&gt;

&lt;p&gt;What the figure does not tell you is how a given team should respond. 1.7x is an average across 470 open-source PRs. Your codebase, your models, your prompt setup, and your review standards all move that ratio in either direction. A team that feeds its rules and past defects into the review loop should expect fewer idiomatic, readability, and exception-handling misses than a team that gives an agent a bare repository and merges on green CI. The report's own recommended mitigations are mostly about context and structure: policy-as-code, correctness rails, stronger security defaults, and review tooling that has access to the team's patterns. None of those is "review harder." All of them are "give the validator the context it needs."&lt;/p&gt;

&lt;p&gt;That last point is where the growing-volume question stops being about generation and becomes about validation. The reason AI review tools exist, &lt;a href="https://kodus.io" rel="noopener noreferrer"&gt;Kodus&lt;/a&gt; included, is that the human bottleneck is now the expensive one. Salesforce reached the same place internally: whatever the PR generation rate, what determines quality is the review system's context and routing, not how fast the code appears. So the practical reading of the report is not "1.7x, therefore review everything harder." It is "1.7x in the baseline case, and most of that gap is in categories that a well-contexted review process can actually catch." That converts a scary headline into an engineering job: build the routing and context so the human attention that remains lands on the changes where 1.7x is most likely to bite.&lt;/p&gt;

&lt;h2&gt;
  
  
  The metric you should watch instead
&lt;/h2&gt;

&lt;p&gt;If one number is over-quoted in this space, the alternative worth tracking is review time per largest PR, which is exactly the signal Salesforce says started plateauing. A ceiling there means reviewers stopped engaging with the biggest changes, and that is the failure mode that produces the security and architecture regressions everyone is worried about. Volume claims, whether 30% more code or 1.7x more issues, describe the input. Whether senior reviewers still have attention left is the output that decides if any of it matters. A team watching only its PR count will miss the moment when its most experienced people quietly start approving changes they have not fully read, which the plateau in review time is how that shows up.&lt;/p&gt;

&lt;p&gt;The other number worth close reading is CodeRabbit's separate claim about incidents per pull request rising 23.5% while PRs per author rose 20%. If that ratio holds for a team, it means the validation problem is growing faster than the throughput problem, and no amount of faster generation fixes it. That is the reason the growing-volume question keeps being asked: the bottleneck moved downstream, and the cost now sits in review and in incidents, not in how quickly code appears.&lt;/p&gt;

</description>
      <category>aicodequality</category>
      <category>codereview</category>
      <category>ai</category>
      <category>engineering</category>
    </item>
    <item>
      <title>A code reviewer that silently skips your rules is worse than one that ignores them</title>
      <dc:creator>Tess Ainsley</dc:creator>
      <pubDate>Thu, 24 Sep 2026 01:45:01 +0000</pubDate>
      <link>https://dev.to/tessainsley/a-code-reviewer-that-silently-skips-your-rules-is-worse-than-one-that-ignores-them-553p</link>
      <guid>https://dev.to/tessainsley/a-code-reviewer-that-silently-skips-your-rules-is-worse-than-one-that-ignores-them-553p</guid>
      <description>&lt;p&gt;Claude Code shipped AGENTS.md support and then, under common configurations, silently never read the file. Engineers on telemetry-off setups and third-party gateways found their project instructions were not being loaded at all, with no warning printed anywhere. A one-line canary test is what exposes it. That same test is the fastest way to check whether any AI code reviewer on your team is actually enforcing your coding rules.&lt;/p&gt;

&lt;h2&gt;
  
  
  The case: a local rules file gated behind a remote switch
&lt;/h2&gt;

&lt;p&gt;The detail that matters here is not that a feature was behind a feature flag. It is that the gate could never turn on for the people who most want the feature. Anthropic's &lt;a href="https://github.com/anthropics/claude-code/issues/95690" rel="noopener noreferrer"&gt;GitHub issue #95690&lt;/a&gt; documents it directly: Claude Code 2.1.277 implemented AGENTS.md as a built-in plugin that is off by default and only loads when a remote flag called &lt;code&gt;tengu_agents_md_mod&lt;/code&gt; resolves true. If telemetry is disabled, or you run through Bedrock, Vertex, or a third-party gateway, that flag can never be fetched, and the file is skipped silently.&lt;/p&gt;

&lt;p&gt;Przemek Szypowicz reproduced and measured it on &lt;a href="https://blog.szypowi.cz/p/claude-code-reads-agents.md-only-when-telemetry-is-on/" rel="noopener noreferrer"&gt;his blog&lt;/a&gt;, published the same day the issue is open. His setup was a directory containing only an &lt;code&gt;AGENTS.md&lt;/code&gt; with a canary word in it, then a prompt asking for the word. With &lt;code&gt;DISABLE_TELEMETRY=1&lt;/code&gt; or &lt;code&gt;CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1&lt;/code&gt; set, the model answered that it had no project instructions. The interesting part is that setting either variable to &lt;code&gt;0&lt;/code&gt; did not help: the docs treat any non-empty value as disabling, so even an explicit zero keeps the block in place.&lt;/p&gt;

&lt;p&gt;None of those runs printed a warning. The session started, the model answered without the project rules, and nothing mentioned that a file had been skipped. A workaround exists, a one-line &lt;code&gt;CLAUDE.md&lt;/code&gt; containing &lt;code&gt;@AGENTS.md&lt;/code&gt;, which loads the file without waiting on the remote flag. But if you had not gone digging you would simply conclude that Claude Code ignores your instructions.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this matters for AI code review, not just coding agents
&lt;/h2&gt;

&lt;p&gt;The same silent gap shows up on the review side. Teams set up an AI reviewer, configure a rules file, and assume the tool loads it. When a review comes back that plainly violates a stated team convention, the usual diagnosis is that the model is bad at following instructions. The much more common failure is that the instructions never reached the model in the first place.&lt;/p&gt;

&lt;p&gt;That is a dangerous conclusion to draw in either direction. If the model actually did receive your rules and ignored them, the fix is a better prompt or a better model. If it never received them, then no amount of prompt tuning does anything. A team that does not verify the load is building a feature on an assumption that can quietly be false, and every tool that advertises custom rules is vulnerable to it in some form, whether through a default-off plugin, a rollout gate, or an environment where fetching settings is blocked.&lt;/p&gt;

&lt;p&gt;There is a second reason this is specifically an AI reviewer problem and not just a coding agent problem. A reviewer is only as useful as the standards it applies, and its whole job is to check other people's changes against those standards. When the rules do not load, the reviewer does not fail loudly. It produces confident, irrelevant feedback that your team reads as correct but that reflects no knowledge of your conventions. That is worse than a reviewer that returns nothing, because nothing at least forces a human to look. Confident wrong-context feedback teaches reviewers to skim.&lt;/p&gt;

&lt;p&gt;The failure surface is wider than a single feature flag too. Rules load through different mechanisms depending on the tool and the environment: a remote fetch, a per-repo settings file, an env block, a user-level config. Each one is a place where the load can silently fail. CI is the highest-risk case because it is often configured to disable telemetry and nonessential traffic by policy, exactly the setup that broke the Claude Code loader. A reviewer running in CI is not the same reviewer running from a developer's laptop.&lt;/p&gt;

&lt;h2&gt;
  
  
  The canary test you should run before trusting custom rules
&lt;/h2&gt;

&lt;p&gt;The verification method is the same one Szypowicz used, and it generalizes to any reviewer. Put a unique canary token into your rules file, then submit a change that your rules specifically forbid, and ask the reviewer directly whether it is following the rules. If it cannot name the canary, or it approves a change you explicitly banned, the rules are not loaded.&lt;/p&gt;

&lt;p&gt;A practical version for a PR review tool: add one line to your rules file like &lt;code&gt;The canary word for this repo is PERIWINKLE; reject any change that mentions it.&lt;/code&gt; Then open a small PR that contains the word. Review the reviewer's output. If it flags the change, your rules are reaching the model. If it does not, stop tuning prompts and start checking the load path.&lt;/p&gt;

&lt;p&gt;Run the test in the environment where the reviewer actually runs, not the one where it happens to work. If your reviews happen in CI, the canary PR goes through CI, because that is where the fetch is most likely to be blocked. If reviewers run on developers' machines, run it there. The whole point of the Claude Code case is that a feature can behave differently in a headless first-run session than in a warmed-up interactive shell.&lt;/p&gt;

&lt;p&gt;Do not assume the documented behavior is what runs in your environment either. Claude Code shows a feature that is announced as present and default-off behind a switch that cannot resolve. Headless runs, CI runners, and anything that disables feature-flag traffic are precisely the environments that hit this, because the flag fetch is skipped.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the fix looks like on the tool side
&lt;/h2&gt;

&lt;p&gt;The uncomfortable part of Szypowicz's writeup is that the model could not tell you what happened. His request was for a tool to show which instruction files were loaded, their priority, and the effective context used, so agent behavior becomes predictable and debuggable. That is the right ask, and reviewers differ meaningfully here.&lt;/p&gt;

&lt;p&gt;Some tools answer it by construction. Kodus, for example, is built on what its &lt;a href="https://docs.kodus.io/" rel="noopener noreferrer"&gt;docs&lt;/a&gt; describe as true visibility: you can see every file read and every decision considered, and team conventions are enforced as explicit Kody Rules plus Memories rather than left to prompt drift. When the loaded rules are inspectable, the canary test becomes a one-time sanity check instead of a standing mystery. Other tools treat rules as black-box context and give you no way to tell whether the file was even read, which is the failure mode this entire case study is about.&lt;/p&gt;

&lt;p&gt;Whatever tool you pick, run the canary test once at adoption and again whenever you change the rules file or the hosting setup. It costs nothing, and it is the only way to know whether the rule the vendor says is loaded is actually reachable in your environment.&lt;/p&gt;

</description>
      <category>aicodereview</category>
      <category>codereview</category>
      <category>agents</category>
      <category>claudecode</category>
    </item>
    <item>
      <title>The 'best AI code review tool' SERP is a wall of self-rankers. One independent benchmark breaks it.</title>
      <dc:creator>Tess Ainsley</dc:creator>
      <pubDate>Thu, 24 Sep 2026 00:15:05 +0000</pubDate>
      <link>https://dev.to/tessainsley/the-best-ai-code-review-tool-serp-is-a-wall-of-self-rankers-one-independent-benchmark-breaks-it-33n7</link>
      <guid>https://dev.to/tessainsley/the-best-ai-code-review-tool-serp-is-a-wall-of-self-rankers-one-independent-benchmark-breaks-it-33n7</guid>
      <description>&lt;p&gt;Search for "best AI code review tools" and you get a catalog where the publisher is also the winner. &lt;a href="https://codeant.ai/blogs/best-ai-code-review-tools" rel="noopener noreferrer"&gt;CodeAnt ranks CodeAnt first&lt;/a&gt;. &lt;a href="https://deepsource.com/resources/ai-code-review-tools" rel="noopener noreferrer"&gt;DeepSource ranks DeepSource first&lt;/a&gt;. &lt;a href="https://linearb.io/blog/best-ai-code-review-tool-benchmark-linearb" rel="noopener noreferrer"&gt;LinearB's benchmark&lt;/a&gt; declares LinearB produced the best signal-to-noise ratio. Greptile and CodeRabbit each publish benchmarks where their own tool tops out. The pattern is consistent enough to treat as a feature of the category rather than a string of coincidences.&lt;/p&gt;

&lt;p&gt;That changed this year. Martian, an AI research lab that does not sell a code review tool, published an open-source Code Review Bench. The dataset, the judge prompts, and the pipeline are all public, so any team can reproduce the numbers or run its own tool through the same harness. It is the first thing on this SERP that a buyer does not have to take on faith.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why every vendor list ranks itself first
&lt;/h2&gt;

&lt;p&gt;The self-ranking is structural, not a judgment on the vendors. A comparison list costs money and engineering time to produce. A vendor that publishes one is deciding to spend that budget on something that competes for the same purchase as its own product. The number a reader should expect to be most defensible is the one that favors the publisher.&lt;/p&gt;

&lt;p&gt;DeepSource's list gives the play away. It opens by arguing most such lists are "written by the tools themselves" and are therefore unreliable, then ranks DeepSource #1 with an 84.51% F1 on the OpenSSF CVE Benchmark on that same page. LinearB's post, published November 2025, says the benchmark "proved that AI code review is about signal-to-noise" and reports LinearB produced the best signal-to-noise ratio across 16 bug types. Greptile reports an 82% catch rate on its own internal benchmark of 50 PRs, a dataset the DeepSource page points out is not independently validated. Every one of these is a real measurement of something. None of them is an answer to the question a buyer actually asked, which is which tool works best under your conditions.&lt;/p&gt;

&lt;p&gt;The live SERP for "best AI code review tools" makes the wall visible: CodeAnt at position 2, DeepSource at 4, CodeRabbit's homepage at 5, Greptile's at 6. The one neutral entry is a Reddit thread where a random user gives an opinionated top-five. Nothing between position 2 and 10 lets a reader verify a claim without trusting its author.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Martian actually published
&lt;/h2&gt;

&lt;p&gt;The &lt;a href="https://codereview.withmartian.com" rel="noopener noreferrer"&gt;Code Review Bench&lt;/a&gt; is two benchmarks that check each other, and &lt;a href="https://github.com/withmartian/code-review-benchmark" rel="noopener noreferrer"&gt;both are open source on GitHub&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The offline half is a fixed dataset: 50 pull requests from five major open-source repositories (Sentry, Grafana, Cal.com, Discourse, Keycloak), each with human-curated golden comments describing the real issues a reviewer should catch, 173 in total, labeled by severity and category. An LLM judge matches each tool's review against those goldens. Three judge models are used, and the README reports the top five tools are identical across all three, with most varying by at most two rank positions. That is a reproducibility check a single-model vendor benchmark cannot offer.&lt;/p&gt;

&lt;p&gt;The online half samples fresh, merged pull requests from GitHub where review bots left comments, then measures whether the developer actually changed the code in response to a bot's suggestion. Because the PRs are recent and continuous, a tool cannot have memorized them during training. This is the honest answer to the training-data leakage problem that tanks most static benchmarks.&lt;/p&gt;

&lt;p&gt;The judge prompt for both halves is worth quoting because it is the whole method in one line: the judge is asked whether a tool comment and a ground-truth issue "describe the same underlying issue," with different wording treated as equivalent. That is the correct definition of a match, and it is what separates a measurement from a name-drop. Two tools can flag the same bug in different words; a naive string or embedding match would call one right and one wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reading the vendor number back against this
&lt;/h2&gt;

&lt;p&gt;The Martian harness gives you a way to read the self-ranked numbers instead of rejecting them. The key is to notice which benchmark a number came from.&lt;/p&gt;

&lt;p&gt;OpenSSF's CVE Benchmark, DeepSource's source, is a fixed, public dataset of 200-plus real production vulnerabilities. That is reproducible and worth trusting as a lower bound on detection. CodeAnt's "300K-PR benchmark" headline, by contrast, is the online dimension: PR volume in the tens of thousands measured by whether developers acted on comments, not a 300,000-PR human-labeled ground truth. The reproducible offline set is 50 PRs. Neither framing is dishonest on its own; the inflation comes when a vendor lets the large online number borrow credibility from the small labeled one without saying so.&lt;/p&gt;

&lt;p&gt;LinearB's signal-to-noise framing is genuinely useful and the Martian design agrees with it. Scoring precision (did the bot's comment match a real fix) and recall (what share of real fixes did the bot catch) separately, then combining with F-beta, is exactly the precision-and-noise tradeoff LinearB says matters. The difference is that Martian publishes the judge prompts and dataset so a team can either trust them or rerun them, while LinearB's benchmark lives behind a buyer's guide download.&lt;/p&gt;

&lt;h2&gt;
  
  
  The tool list goes past the usual suspects
&lt;/h2&gt;

&lt;p&gt;The Martian README names the tools it evaluated: CodeRabbit, GitHub Copilot, Claude Code, CodeAnt, Greptile, Qodo, Cursor Bugbot, Graphite, Augment, Baz, Cubic, Devin, Gemini, GitLab Duo, KG, Kodus, Macroscope, Sourcery, among others. That alone is a step past the vendor lists, which mostly limit comparisons to tools the publisher competes with. An independent benchmark has no reason to leave anyone out.&lt;/p&gt;

&lt;p&gt;For a team deciding how much weight to give the leaderboard, the honest reading is that any fixed ground-truth list is a starting point, not a conclusion. The offline set is 50 PRs from five repositories that skew toward specific domains: error tracking, observability, scheduling, forums, authentication. A team shipping embedded C or heavy data pipelines is outside that distribution. The online tracker filters by language, domain, diff size, and severity, so a team can narrow the continuous data to PRs that resemble its own before drawing a conclusion.&lt;/p&gt;

&lt;h2&gt;
  
  
  The takeaway for an engineering team
&lt;/h2&gt;

&lt;p&gt;Do not stop at the winner column. The entire point of an open benchmark is that you can rerun it, so treat the leaderboard as a reproducible dataset rather than a verdict. Three concrete checks get a team most of the way to a defensible choice.&lt;/p&gt;

&lt;p&gt;First, ask which benchmark backed the number. A vendor page that cites an independent, public dataset with a named judge is doing real work. One that reports results from its own 50-PR internal set is a marketing experiment, and the same page arguing that everyone else's self-report is untrustworthy is hoisting itself by its own petard.&lt;/p&gt;

&lt;p&gt;Second, ask whether precision and recall are reported separately or collapsed into a single score. Tools that chase total bug counts produce noise; you saw the same thing in LinearB's finding that CodeRabbit caught the most issues while generating heavy noise. A single F1 collapses the behavior a developer feels on every PR, which is the flood of irrelevant comments. Read both numbers.&lt;/p&gt;

&lt;p&gt;Third, run the offline benchmark yourself with your stack. It takes an afternoon. Fork the PRs, trigger the tool, run the pipeline, judge against the goldens. The marginal cost is small and it converts a vendor claim into a number your team produced.&lt;/p&gt;

&lt;p&gt;The category finally has a measurement it did not write itself. That is the thing to build on. Stop reading self-ranked lists as answers and start reading them as evidence, then put them through an evaluation you can verify.&lt;/p&gt;

</description>
      <category>codereview</category>
      <category>aicodereview</category>
      <category>benchmark</category>
      <category>developertools</category>
    </item>
    <item>
      <title>Two AI code review benchmarks disagree on the winner, and agree on what to measure</title>
      <dc:creator>Tess Ainsley</dc:creator>
      <pubDate>Wed, 23 Sep 2026 00:15:03 +0000</pubDate>
      <link>https://dev.to/tessainsley/two-ai-code-review-benchmarks-disagree-on-the-winner-and-agree-on-what-to-measure-55kl</link>
      <guid>https://dev.to/tessainsley/two-ai-code-review-benchmarks-disagree-on-the-winner-and-agree-on-what-to-measure-55kl</guid>
      <description>&lt;p&gt;The top of the Google results for "best AI code review tools" is a wall of vendor listicles. Most of them rank the publisher's own product first and support that rank with feature bullets, not numbers. Two pages on that SERP actually publish a method, and they are worth reading together because they come to opposite conclusions about the winner.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two evaluations, two winners
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://linearb.io/blog/best-ai-code-review-tool-benchmark-linearb" rel="noopener noreferrer"&gt;LinearB's benchmark&lt;/a&gt; measured 16 bugs across two phases and scored each reviewer on four dimensions: competency, clarity, configurability, and developer experience. The write-up reports that LinearB produced the best signal-to-noise ratio and foregrounds statefulness, the ability to withdraw or revise a comment once a later commit has made it stale. In that telling, CodeRabbit caught the most total issues but generated heavy noise, flagging the same pattern repeatedly without context. GitHub Copilot delivered consistently relevant suggestions but with shallow context, missing multi-file reasoning. Graphite Diamond performed weakest on detection.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://deepsource.com/resources/ai-code-review-tools" rel="noopener noreferrer"&gt;DeepSource's comparison&lt;/a&gt; measured a similar category with a different instrument. It ran every tool against the OpenSSF CVE Benchmark, a public dataset of 200+ real production vulnerabilities, and publishes F1 scores. DeepSource reports the highest F1 at 84.51%. CodeRabbit scores 36.19% F1 in that telling. Greptile's self-reported 82% catch rate, DeepSource notes, comes from an internal benchmark of 50 PRs across five repositories that is not independently validated.&lt;/p&gt;

&lt;p&gt;Two vendors, two methods, two winners, and both write-ups rank their own product first. That is normal, and it is not a reason to dismiss either page. It is a reason to read the method instead of the winner.&lt;/p&gt;

&lt;p&gt;The two evaluations are incommensurable. LinearB's dataset is built in-house, and the bug list and scoring criteria live in a downloadable whitepaper. DeepSource's dataset is public, but it is a security corpus: OpenSSF CVEs score whether a reviewer detects real recorded vulnerabilities. That is a different question from whether a reviewer writes clear, usable comments. The same tool that wins on signal-to-noise can trail on F1 against a security corpus, because the two metrics punish different failure modes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the two benchmarks agree
&lt;/h2&gt;

&lt;p&gt;Reading both pages side by side is worth it because of the quarters where they do not disagree. Both independently land on the same dimensions separating a useful reviewer from a noisy one.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Signal-to-noise ratio.&lt;/strong&gt; Both pages make this the deciding quality. A reviewer that finds 90% of issues but buries them in hundreds of comments per PR is worse than one that finds 70% and says it once. LinearB calls it the best reviewer says more with less. DeepSource puts it near the headline of its methodology. This metric is not really about whether the model catches bugs. It is about trust. Developers stop reading comments they have learned to ignore, and once that happens the reviewer is adding noise to every PR without contributing anything.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Statefulness across commits.&lt;/strong&gt; Pull requests change between commits. A reviewer that treats each commit as a fresh start makes developers re-litigate issues they already resolved. LinearB measured this directly: the reviewers that withdraw outdated comments and revise their opinion after a fix score predictably higher on developer experience. The tools that restart from zero on every commit pay for it in review time, because a fix that should close a thread just opens a new one.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Configurability.&lt;/strong&gt; Both pages raise this independently. Teams have different tolerances for verbosity and different standards, and reviewers that let teams tune rules, tone, and enforcement through configuration fit into existing workflows instead of demanding new ones. LinearB specifically notes that YAML-defined rules and slash commands correlated with a smoother developer experience. This is the axis teams that care about their own coding standards should watch closely. It separates a tool that adapts to your rules from one that applies generic rules to everyone. Reviewers in this bucket, Kodus among them, put rules and signal ahead of raw finding count, which is exactly the configurability axis both benchmarks flag as the one that determines whether a tool gets adopted at all.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Time to first useful signal.&lt;/strong&gt; LinearB measures the average time from PR open to the first correct, actionable comment. A fast but wrong comment is worse than a slow correct one, which is why the metric exists: it stops vendors from optimizing for response speed at the expense of accuracy. It is a genuinely useful thing to time in your own workflow, because it tells you how long a developer waits before the reviewer does something they trust.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to take away
&lt;/h2&gt;

&lt;p&gt;You will not get a winner from these two pages, and you should not expect one. What you get is a checklist for evaluating reviewers internally. Hold a vendor to three questions that both published write-ups imply.&lt;/p&gt;

&lt;p&gt;What is the dataset, and can you inspect it? DeepSource gives you OpenSSF CVEs. LinearB gives you a 16-bug in-house list inside a whitepaper. A number without a method you can verify is marketing, not evidence. Note that neither of these ranks an independent benchmark where the publisher is not also a contestant; that is a separate, harder bar, and one an exploratory arXiv benchmark I covered in an &lt;a href="https://dev.to/tessainsley/a-code-review-benchmark-that-isnt-the-vendor-ranking-itself-4jp6"&gt;earlier article on this feed&lt;/a&gt; tries to clear by scoring reviewers against an independent annotation workflow rather than a vendor's own product.&lt;/p&gt;

&lt;p&gt;Which metric, and what does it punish? F1 penalizes false positives alongside false negatives, which is partly why the safety numbers in a security corpus diverge from a DevEx-style evaluation. Make the metric match the failure mode you are buying against. If you lose more time to false positives than to missed bugs, optimize for precision first. If a misses slip to production, weight recall.&lt;/p&gt;

&lt;p&gt;Is the reviewer stateful and configurable, or does it restart and shout? On that axis, the two competitors who disagree about the winner agree completely. That agreement is the most transferable finding in either document, because it predicts how a tool will behave in your review cycle regardless of which vendor happens to come out on top of its own benchmark.&lt;/p&gt;

&lt;h2&gt;
  
  
  The wider pattern
&lt;/h2&gt;

&lt;p&gt;The broader pattern is worth naming. Generation got fast and cheap, so the market filled with reviewers. Validation got the opposite treatment. Most marketing still optimizes for volume of output, not for whether the output is worth reading. The two ranked pages that included a method are the exception, and they independently reach the same conclusion: lean reviews, fewer comments, rules you control, and a reviewer that remembers what it already said. That convergence across competitors is worth more than either winner claim.&lt;/p&gt;

&lt;p&gt;Stamp: claims checked 2026-09 against the LinearB and DeepSource pages linked above.&lt;/p&gt;

</description>
      <category>codereview</category>
      <category>ai</category>
      <category>softwareengineering</category>
    </item>
    <item>
      <title>When agents sped up, two teams hit the same wall: validation couldn't keep up</title>
      <dc:creator>Tess Ainsley</dc:creator>
      <pubDate>Wed, 23 Sep 2026 00:15:02 +0000</pubDate>
      <link>https://dev.to/tessainsley/when-agents-sped-up-two-teams-hit-the-same-wall-validation-couldnt-keep-up-2mlc</link>
      <guid>https://dev.to/tessainsley/when-agents-sped-up-two-teams-hit-the-same-wall-validation-couldnt-keep-up-2mlc</guid>
      <description>&lt;p&gt;Agents made it faster to write code. Two engineering teams published what happened on the validation side when they did, and the two write-ups land on the same structural claim from opposite ends. Linear found its CI gate became the bottleneck. Salesforce found its human review stopped engaging. Neither team concluded that generation was the problem. Both concluded that the part of the system built to judge a change had not kept pace with the part built to produce one.&lt;/p&gt;

&lt;h2&gt;
  
  
  The CI side: Linear's funnel
&lt;/h2&gt;

&lt;p&gt;Linear's engineers, in a post titled &lt;a href="https://linear.app/now/ci-bottleneck-reworked" rel="noopener noreferrer"&gt;AI coding has made CI a bottleneck, so we reworked ours to keep up&lt;/a&gt;, dated September 21, 2026, report that their test suites almost quadrupled since the start of the year. Pull request wait time went from more than 6 minutes to just over 5, and runner time per test dropped roughly in half. The throughput gains are real, but the story is about where they came from. Linear did not get faster by telling agents to slow down. It reworked the validation machinery.&lt;/p&gt;

&lt;p&gt;The list of changes is mostly mechanical. Faster third-party runners with better CPU, storage, and cache made the same pipeline run about 34% faster on average, with the tsc workload down 52%. Moving to the tsgo compiler cut the tsc check's weekly median by 73%, enough to move typechecking off the bottleneck. A handful of custom lint rules that depended on TypeScript type information were rewritten to use the abstract syntax tree, which let ESLint shed TypeScript entirely and cut API lint time by 68%. A change-detection job that checked the full working tree when it needed almost none of it had its median trimmed from 26 to 8 seconds. Moving a cache-marker write off the critical path saved 42 seconds on every API pull request.&lt;/p&gt;

&lt;p&gt;The headline metric there is the one nobody markets: how long a proposed change waits to be judged, and how much machine time the judging consumes. As test suites quadruple under agent load, that is the number that decides whether agents feel fast in practice.&lt;/p&gt;

&lt;h2&gt;
  
  
  The review side: Salesforce's plateau
&lt;/h2&gt;

&lt;p&gt;On the other end of the same funnel, Salesforce engineers &lt;a href="https://engineering.salesforce.com/scaling-code-reviews-adapting-to-a-surge-in-ai-generated-code/" rel="noopener noreferrer"&gt;Shan Appajodu and Ravi Boyapati published "Scaling Code Reviews: Adapting to a Surge in AI-Generated Code" in January&lt;/a&gt;. Their internal signals: code volume rose roughly 30%, pull requests regularly exceeded 20 files and 1,000 changed lines, and review latency climbed quarter over quarter. The signal they flag as most worrying is when review time on the largest pull requests plateaued or declined. They read that as reviewers no longer meaningfully engaging.&lt;/p&gt;

&lt;p&gt;That is a different failure than the CI one. Large diffs spanned unrelated files and architectural layers, so reviewers spent more time navigating than reasoning, and the second-pair-of-eyes guarantee eroded. Salesforce's answer was not to make review faster in the throughput sense. It was to rebuild the review system around reconstructing intent, treating context as a first-class dependency, and progressively disclosing risk so reviewers spend attention where it matters. I wrote separately about &lt;a href="https://dev.to/tessainsley/when-review-time-plateaued-reviewers-had-stopped-reading-602"&gt;what that plateau says when review time stays flat&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The convergence
&lt;/h2&gt;

&lt;p&gt;Put the two reports side by side and the shape is clear. Linear's validation bottleneck was mechanical and cheap to fix, so they sped it up. Salesforce's was cognitive and structural, so they rebuilt how review presents a change. But both are the same failure: generation got cheaper and faster, and the system that verifies a change did not scale with it. That is the load every team with agents is now carrying at one of two points, and often both.&lt;/p&gt;

&lt;p&gt;Linear's post gives the concrete numbers for the mechanical end. Salesforce's post gives the language for the cognitive end. The engineering task is knowing which end you are hitting. If pull requests sit waiting on green checks and runners are expensive, you have Linear's problem. If large changes sail through with flat review time and senior reviewers context-switch all day, you have Salesforce's problem. The two fixes are not interchangeable.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to measure instead
&lt;/h2&gt;

&lt;p&gt;Teams commonly track cycle time and generation throughput, and both reports suggest those are the wrong scoreboard. Linear tracked wait-on-CI and runner time per test. Salesforce tracked review time on the largest pull requests, because a flat number there is a disengagement signal hiding a decline in scrutiny. A useful review metric under agent load is the ratio of validation cost to change size, watched across PR size buckets so the huge 1,000-line changes do not get lost in an average.&lt;/p&gt;

&lt;p&gt;Tooling is catching up on both ends. CI-aware validation and cost-conscious runners shore up the mechanical side, and context-first review tools work the cognitive side by keeping related changes grouped and surfacing risk. Kodus fits in the same category as the context-first review tools, applying team-specific rules rather than generic linting. The point of the comparison is not which product wins. It is that a review tool earns its place on the side of the funnel where your own bottleneck sits, so the first question is which end is breaking for you.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two primary sources, one conclusion
&lt;/h2&gt;

&lt;p&gt;Both posts are worth reading in full, and both are unusually honest about method. Linear publishes its before and after numbers. Salesforce names the metric that actually signaled trouble. If you want a single takeaway: when agents make code generation effectively free, the constraint that remains is validation, and the teams that publish measured work are the ones fixing that side.&lt;/p&gt;

&lt;p&gt;The vendors marketing "reviews every PR" are talking about throughput. The measured sources are talking about whether a human looked closely enough, at either the CI gate or the review table. When you are asked to review a growing volume of AI-written code, the first useful step is to place your own bottleneck on that spectrum. If review time on your largest changes is flat while the changes get bigger, the problem is not that review is slow. It is that review stopped reading. That is the number to watch, and it is the one neither vendor dashboard tends to show.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>codereview</category>
      <category>softwareengineering</category>
    </item>
  </channel>
</rss>
