If you spend any time around coding agents in 2026, you've absorbed the advice: use multiple agents. One plans, one implements, one reviews. The implication runs one direction. Two agents beat one, and if two are good, three should be better.
Almost nobody has tested that claim with an oracle on the other end. It sounds true, because code review works for human teams. But "more reviewers help" and "reviewers who fail differently help" are two different claims, and the second can be true while the first is false.
RankEvolve: A Reliable Multi-Agent Auto-Research Harness for Evolving Ranking Models (Meta, August 2026) is nominally a paper about evolving a recommender model. The part I kept rereading is the benchmark buried inside it. ExecML turns the defects logged during that deployment into 96 oracle-graded tasks on each of two codebases, and then asks a question the multi-agent industry has mostly skipped: does composing two complete coding-agent products buy executable correctness, beyond spending the same budget on one?
The answer is yes. With a catch that matters more than the yes. On 96 tasks over the HSTU codebase, Claude Code → Codex → Claude Code hit 62.5% execution accuracy where Claude Code → Claude Code → Claude Code hit 45.8%, at the same budget. Meanwhile, doubling Claude Code netted roughly zero over a blind best-of-N baseline. More review and different review aren't the same thing.
This is the companion to my last post on harness engineering. That paper described how coding-agent harnesses are built. This one measures what happens when you use one harness to check another.
The Research Process
ExecML is a benchmark, not a leaderboard entry. The design:
- Two codebases. The public HSTU recommender implementation (Python/PyTorch) and LitGPT, a non-recommender training framework. HSTU carries the primary claim; LitGPT is a pre-specified transfer test.
- 96 private tasks per repository, seeded by real incidents logged during a twelve-iteration deployment. The tasks balance six incident families: data provenance and leakage, tensor routing, gradient flow, train/eval mode, metric semantics, and configuration wiring. Half are new implementations, half are repairs of deliberately introduced faults.
- A hidden executable oracle per task. A patch passes only if it simultaneously satisfies the regression suite, task-specific behavioral checks, scientific-safety invariants (leakage, dead gradients, wrong evaluation semantics), and evaluator-integrity checks. The primary endpoint is all-or-nothing execution accuracy (EA). The secondary endpoint is the silent critical-defect rate (CDR): oracle-confirmed faults that leave the patch runnable but capable of invalidating a scientific conclusion.
- Six conditions, budget-matched. Every flow gets the same three roles, per-node limits, tool permissions, action caps, and wall-clock and dollar caps from a frozen resource envelope. Timeouts and overruns count as failures and stay in the denominators.
- Products pinned. Claude Code ran Claude Opus 4.8; Codex ran GPT-5.6; both at maximum reasoning effort, in fresh sandboxes, randomized order, identical prompts and tool policies.
One caveat belongs at the top, not the bottom: "matched budget" means matched observable inference budget. Proprietary products don't expose provider-side compute, so this isn't equality of FLOPs. The authors are explicit about it, and I'll come back to it.
Key Findings
1. Heterogeneous composition wins at parity of budget
Here is the spine of the paper, three roles per condition, all measured on the same envelope:
| Condition | HSTU execution accuracy ↑ | HSTU silent defects ↓ | LitGPT execution accuracy ↑ | Cost/task |
|---|---|---|---|---|
| Claude Code, one pass (unmatched cheap reference) | 22.9% | 35.4% | 20.8% | $0.41 |
| Claude Code, full budget | 33.3% | 27.1% | 29.2% | $0.98 |
| Claude Code best-of-N, oracle-blind selection | 43.8% | 18.8% | 39.6% | $0.95 |
| Claude Code → Claude Code → Claude Code | 45.8% | 16.7% | 43.8% | $0.97 |
| Claude Code → Codex → Claude Code | 62.5% | 10.4% | 56.2% | $1.02 |
| Codex → Claude Code → Codex | 56.2% | 12.5% | 50.0% | $1.00 |
Look at the ladder beneath the winner. Spending more on one product moves accuracy from 22.9% to 33.3%. Adding independent candidates and an oracle-blind selection step reaches 43.8%. A same-product review chain adds almost nothing on top of that, inside the noise of the step below. Then swapping the reviewer for a different vendor's product jumps to 62.5%.
The paired contrast against the strongest budget-matched baseline is +16.7 points, 95% CI [6.6, 26.7]. The silent-defect rate falls from 16.7% to 10.4%. And it costs $1.02 per task against $0.97. A five-cent difference on a 16.7-point gap.
2. It replicated where it had no obligation to
The LitGPT columns are the pre-specified transfer test, and the authors committed in advance to publishing the result whatever it showed. It showed +12.5 points, 95% CI [3.0, 22.0], p = 0.008. The same direction, on a codebase that has nothing to do with recommenders.
This is the finding that blunts the obvious objection. Recommender code has particular shapes; HSTU has particular failure modes; the authors were from Meta Surely the effect is tuned to their home turf. A pre-registered, independent-domain replication doesn't eliminate that worry, but it puts most of it to rest.
3. The mechanism is not "diversity is good." It is complementary error.
This is the most useful part of the paper, and it's easy to skim past because it's definition-heavy. Strip the notation and the idea is simple.
For a pair where agent A implements and agent B reviews, two quantities matter:
- Complementarity (D): the probability mass where A fails and B succeeds. These are the errors B is able to rescue.
- Realized gain (G): what accrued to A's patch after review. Rescues minus the damage B's review introduced.
If B's errors track A's errors, D is small no matter how good B is in general: B is blind where A is blind. That's what the two measured pairs show:
- Claude Code reviewed by Codex: error correlation ρ = 0.21, complementarity D = 0.18, realized gain G = +0.06.
- Claude Code reviewed by Claude Code: error correlation ρ = 0.58, complementarity D = 0.10, realized gain G ≈ 0 once the reviewer's own mistakes are subtracted.
Two products that fail alike aren't a second opinion. They are the same opinion, twice. The paper formalizes the relationship and fits a line through it. The slope on complementarity is β̂₁ = 0.34, 95% CI [0.12, 0.56]. But the plain-English version is the one worth taking to work: composition pays when the second agent fails on the tasks the first one passes.
4. The fancy deployed topology's edge is budget, not topology
RankEvolve's production flow is a plan-merge topology: two planning lanes, Claude Code bound to one and Codex to the other, each lane revising its plan while reading the other's latest draft, then a merge, then a review–follow-up-plan dual. It posts the best raw numbers in the paper: 70.8% HSTU accuracy, 6.2% silent defects, 64.6% on LitGPT.
But it runs at $2.40 per task and 250 actions. Hold it to the matched budget and it ties Claude Code → Codex → Claude Code exactly: 62.5%, paired difference 0.0, 95% CI [-5.5, 5.5]. Against a contemporaneously re-run comparator, its paired lift is +8.3 points, but with McNemar p = 0.077 and CI [-0.7, 17.3]. Not statistically distinguishable.
The authors follow their own result: the confirmatory weight rests on the simple pairing, not the elaborate diagram. What the deployed flow buys is accuracy per task at whatever the extra spend costs, and the paper argues that spend is rational because averted defects protect far more expensive training runs. But don't attribute that to the topology. The topology is the packaging.
5. Order matters
The reversed pair still beats every homogeneous condition at 56.2%, but it trails the forward direction by 6.2 points. Which product implements and which reviews is part of the result.
What This Means for You
Stop doubling the same agent. This is the actionable core. If your setup runs two copies of the same product in an author/reviewer pattern, the paper's numbers suggest you are mostly buying a second invoice. Same-product review had ρ = 0.58 and netted approximately zero after its introduced errors. Pick a reviewer from a different lineage. Different vendor, different harness, different training corpus.
Estimate complementarity before you commit. You don't need an oracle to run a small version of this. Take 20–30 tasks where you already know the correct outcome, run your agent, run the second agent's review, and count two things: how many defects the reviewer caught that the author missed, and how many problems the review introduced. If the first number is small, the pairing isn't buying what you think.
Use the budget ladder deliberately. One pass → extended budget → blind best-of-N → cross-product review chain. Each rung fixes a different failure mode and costs a different amount. The paper's data says the biggest single jump comes from switching the reviewer's identity, not from adding passes to the same one.
The economics are lopsided in your favor. The conditions run at roughly $1 per task to gate training runs that consume hours of GPUs. The paper puts the ratio bluntly: an averted silent defect is worth on the order of 10³ times the inference spend that prevented it. If you're arguing for a second-reviewer budget line, that's the argument.
If you build agent products, this is a competitive surface. A different vendor's reviewer is now part of your product's quality story, and the ability to host other agents behind your interface stops being a curiosity and starts being measurable. Composite products are competing on whose reviewer catches what.
What I'd Want Verified
My last post ended with me opening Pi's docs and checking the paper's claims against the primary source. That move is closed off here: Claude Code, and the oracle are proprietary or private. So the audit shifts from implementation to claims. What I can say after reading the HTML version end to end:
- The headline survives its own intervals, but the intervals are wide. Each condition is a binary outcome over 96 tasks, so per-condition means carry roughly ±12-point clustered intervals. The paired +16.7 is the right number to quote, and the authors quote it that way. Paraphrase the result without the interval and you're overselling it.
- "Matched budget" is observable budget. Same envelope, same prompts, timeouts as failures. It isn't FLOP equality, and the authors say so.
- Oracles are incomplete. The paper reports mutation testing and audits to reduce false passes, which is the correct mitigation. It can't eliminate them. A passing patch is "not caught," not "proven correct."
- The mechanism is predictive, not causal. The complementarity slope's interval excludes zero and beats a marginals-only model on held-out data, but the paper explicitly doesn't claim that raising complementarity causes gain. Their words.
- The clean result and the messy one are different studies. The HSTU case study ran with a human operator at the protocol's gates. ExecML, where the 62.5% lives, had no human in the loop. Don't blend them.
- The model pins will age. Claude Opus 4.8 and GPT-5.6 are the products that ran this experiment. A future model that is better at reviewing its own family's errors would shrink the gap. The method is durable; the numbers are a snapshot.
- The replication I want: rerun the pair with open harnesses as author, reviewer, or both. Does the effect survive when the products share more DNA than two vendor stacks do?
Limitations
Condensed from the paper's own section, which is unusually candid. The evidence base is two Python/PyTorch codebases, so cross-domain transfer isn't established. The products are black boxes, so compute equivalence can't be shown. Oracles reduce but don't remove false passes. Complementarity is predictive, not causal. The knowledge layer was active during the deployment but never ablated, so its contribution is unmeasured. The cross-dataset transfer results are single runs. And checkpoint forking understates how far iterations diverge from one another.
Closing
"Two agents are better than one" is false as stated. The paper's numbers show a same-product review chain landing inside the noise of the baseline beneath it. The sentence the data supports is narrower and more useful: a different agent's review is worth more than a duplicate's, and you can measure the difference before you bet a workflow on it.
Call-to-Action
Read the paper yourself: RankEvolve: A Reliable Multi-Agent Auto-Research Harness for Evolving Ranking Models. The benchmark details are in Appendix C and the EOP in Appendix D; Table 4 in Section 4 is the part worth your time first.
Top comments (0)