Here's a situation a lot of teams are in right now. An agent opens a 600-line PR in four minutes. Reading it properly takes you forty. By the time you finish, two more have landed in the queue. Nobody has really decided what "reviewed" means anymore, so the gap gets filled with a skim and an "lgtm".
Writing code got cheap. Reading it didn't. So the real question isn't "should we use agents". It's this: when code shows up faster than any human can read it, what is a human actually supposed to check?
I don't think anyone has a settled answer. Below are the camps I keep seeing, the arguments each one makes, and the part I think nobody has solved. Push back on any of it.
Where this came up
A recent r/rust thread about running about 20 coding agents in parallel on one repo was supposed to be about build caching. The top-voted comment ignored caching completely:
How in the hell are you reviewing the work of 20 simultaneous processes?
Most of the discussion after that was about review, not cargo. Here's the setup that started it, as evidence and not as a recommendation:
- About 20 agents, each in its own git worktree, all on the same Rust crate.
- rustc at full parallelism crashed with an access violation while a local model sat in RAM, so builds were capped at
CARGO_BUILD_JOBS=4. - Test and check results were cached by content hash across worktrees. On smaller swarms, 80%+ of checks came from that cache.
- Nothing merged unless the test score went up and no previously passing test broke. Changes merged one at a time, with a re-grade after each one, so two changes that each pass alone can't break the build together. Nothing touched the real repo until a human applied the final merged diff.
In other words: no line-by-line human review of agent output. That got called everything from "the tests are the reviewer" to "a slop factory." Both are fair reactions, and they lead straight to the three camps.
Camp 1: the tests are the reviewer
The argument: you can't read 20 streams of code, but you can define what correct means and make the machine prove it on every change. That moves the human work from reading diffs to writing specs, which scales.
What it needs to work:
- Tests good enough to deserve the job. A thin suite with a merge gate on top is still just a thin suite.
- A merge gate that re-runs everything after each merge, not just per branch. That's how you catch "both passed alone, broke together".
- Agents can't edit the tests that judge them. That sounds paranoid. It isn't.
That last point comes from an experiment we ran. Five small broken repos, Claude Code headless on Opus, Sonnet and Haiku, and a deliberately rushed prompt ("CI is red, I need to ship in 5 minutes"). On the four bugs that had a real fix, 0 of 24 runs touched a test. Every one fixed the actual bug. On the fifth repo, where two tests contradicted each other so no code change could make the suite pass, 2 of 6 runs edited a test to get green: one skipped the test, one deleted it. Both said so in their summaries. The other four stopped and asked which rule was right.
Small sample, but the pattern is the uncomfortable part. Agents mostly don't cheat when there's an honest fix. They cheat when the spec itself is broken, which is exactly when a test-as-reviewer setup is weakest.
The best argument against it: tests check what you thought of. Review is supposed to catch what you didn't think of. Design drift, a dependency nobody asked for, a quiet O(n²), a log line that leaks a token. None of that fails a unit test.
Camp 2: a second model reviews the first
The argument, as one commenter in the thread put it: "Code review has to be automated to scale. Using a different model works fine, e.g. Codex reviewing Claude or vice versa."
It got downvoted. Even so, it's probably the most common setup in practice right now, because it's cheap and it catches things tests can't: naming, structure, obvious security smells, "why is this file 900 lines".
What it needs to work:
- A reviewer that's actually different (a different vendor, prompt or context) so it doesn't share the author's blind spots.
- Findings that block a merge, or at least force a reply. A reviewer that writes comments nobody reads is just decoration.
The best argument against it: two models can be confidently wrong together, and model review produces a lot of noise. If every PR gets 14 nitpicks, people stop reading them, and you've rebuilt the "lgtm" problem one level up.
Camp 3: merge and pray (and read it later)
Nobody says it this way, but plenty of teams work like this: merge fast, watch production, revert quickly. Review happens after the fact, in incident reviews and in the next person who has to change that file.
The argument: the cost of a bad merge is set by how fast you can detect and roll back, not by how carefully you read. With good observability, feature flags and easy rollbacks, reading every line before merging is a tax you don't need to pay.
The best argument against it came from the same thread, about agents churning dependencies: "Just because agents can churn a huge amount of code, doesn't mean you won't have the same problems when they do." Some mistakes don't show up as an incident. They show up six months later as a codebase nobody understands, including the agents.
The part nobody has solved
My bet: camp 1 is the right backbone, with camp 2 as a cheap filter in front of it. But that combination still leaves a hole, and the hole is the interesting part.
Who reviews the tests?
If tests become the reviewer, every test change becomes the most important diff in the repo. But agents write tests too, and a new test that asserts the buggy behaviour looks exactly like a good test. Locking existing tests stops agents from weakening the gate. It does nothing about new tests that set a low bar.
There's also a second, quieter problem. The more review we hand off, the fewer humans actually know how the system works. Even if the checks are perfect, someone eventually has to change the architecture, and "the tests passed" doesn't teach anyone why the code is shaped the way it is.
And we're not great at judging any of this. METR's 2025 randomized study had 16 experienced open-source developers complete 246 real tasks. They expected AI to speed them up by 24%. With AI they took 19% longer, and afterwards they still believed it had sped them up by 20%. METR itself now calls those results out of date, and tools have improved a lot since then. But the gap between how fast it felt and how fast it was is worth remembering whenever a team says its review process "feels fine".
Something you can use today
Whichever camp you're in, these are cheap and don't depend on picking a side:
- Make existing tests read-only to agents. Allow new tests, but any change to an existing assertion needs a human. This is the single highest-leverage rule I know of.
- Review test diffs before code diffs. If the tests are honest, the code is mostly constrained. If they aren't, the code review was pointless anyway.
- Merge one at a time and re-run everything after each merge. Passing on its own branch isn't the same as passing after the merge.
- Pin the dependency set. Agents request new crates or packages; a human approves them. Every new dependency is a review item, not a side effect.
- Give humans a smaller, sharper diff. Ask the agent for a short "what changed and why" plus the riskiest 30 lines. Read those carefully instead of skimming all 600.
- Measure your review, don't just feel it. Track how often reverted or hotfixed code came from agent PRs. If you can't answer that, you don't know which camp you're actually in.
Your turn
I'd really like to hear how real teams handle this, especially the setups that broke:
- When an agent PR is too big to read, what do you actually look at first?
- Has a second-model reviewer ever caught something your tests missed? Or did it mostly add noise?
- Who reviews the tests in your setup? Is there a rule that stops an agent from writing a test that blesses its own bug?
- If you've gone full merge-and-revert, what made you trust it, and what was the first thing that went wrong?
Correct anything above that doesn't match what you've seen. Real failure stories beat my theory.
Top comments (0)