We wanted to answer a simple question:
How well do today’s leading AI code review tools actually reason about bugs that cross module boundaries?
So we gave Codzee, Greptile, CodeRabbit, and CodeAnt the exact same pull request.
The PR contained 10 planted defects across 172 lines and 2 modules, with unit-related bugs involving kilometres, metres, miles, pence, pounds, seconds, minutes, and more.
The result?
Three tools tied at the top. But Codzee found something nobody else did.
Press enter or click to view image in full size
On paper, it looks like a three-way tie.
But one finding changed the story.
The Bug Nobody Had Found Before
One of the planted defects involved banker’s rounding in a fare calculation.
The code used Python’s round() behavior in a situation where commercial rounding was expected.
For a fare of 1,250 pence, Python rounds it to:
1,200 pence — not 1,300.
Codzee caught it.
But it didn’t stop at saying “rounding may be incorrect.”
It reasoned through the behavior of Python’s standard library and explained the actual commercial consequence: half-pound fares can round differently depending on whether the pound value is even or odd.
Greptile, CodeRabbit, and CodeAnt all reviewed the same function.
None of them caught it.
And this wasn’t just a one-off.
This rounding defect had appeared in previous benchmark exercises. Across 16 benchmarked pull requests and roughly 500 findings, no tool had reported it.
Until Codzee.
That was the moment the benchmark became more interesting than a simple scorecard.
The Full Defect Breakdown
Press enter or click to view image in full size
The benchmark isn’t about claiming perfection. It’s about understanding where each tool succeeds, where it fails, and why.
What the Developer Running the Benchmark Noticed
From the developer’s perspective, another result stood out.
Codzee didn’t just review the implementation.
It also looked at the tests.
Two of its findings identified tests that were effectively validating incorrect behavior including a massively inflated fare and an unrealistic journey duration.
That matters because a bad test can make a bad implementation look correct.
A useful AI reviewer needs to question both.
And Then We Found an Even Bigger Lesson
The most interesting result wasn’t actually which tool won.
It was the 90% detection rate for unit-related defects.
These same types of defects had survived nine previous rounds across the tools.
This time, both modules explicitly stated their units in their opening docstrings.
That small change made a huge difference.
The lesson?
Give the reviewer context, and the reviewer can reason better.
For developers, documenting units is almost free.
For AI code review systems, it provides the context needed to connect values across functions and modules.
So, What Did We Learn?
Codzee
9/10 — tied for first and the only tool to catch the banker’s-rounding defect.
Greptile
9/10 — strong reasoning and consistently concrete findings.
CodeRabbit
9/10 — particularly strong at connecting implementation issues with incorrect tests.
CodeAnt
7/10 — good critical-severity detection, but missed the most severe fare calculation defect.
And Codzee?
Two consecutive strong rounds.
After winning Round 6 with 8/13, Codzee tied for first in Round 7 with 9/10.
The Takeaway
The benchmark started with a simple goal:
Find the bugs.
But the more interesting question became:
Can an AI code reviewer actually understand why the code is wrong?
The banker’s-rounding result suggests Codzee is getting closer to that level of reasoning.
It understood the code → calculation → consequence.
And that’s ultimately what we want from an AI code reviewer
Not more comments. Better understanding.



Top comments (0)