A case study from six rounds of testing Codzee against Greptile, CodeRabbit and CodeAnt
It started with a question we couldn't answer
We build an AI code review tool called Codzee. At some point we realized we couldn't answer a simple question: is it actually good?
Our demos looked great, and so did every competitor's. Every landing page promises "context-aware" and "low-noise" review. None of that tells you what happens when a tool meets a real pull request with a bug it can't find by pattern matching.
So we stopped reading feature lists and ran an experiment. We built pull requests with bugs planted in them, gave the same PR to our tool and three competitors, and checked what each one caught.
We've now done this six times, and each round was harder than the last. This is the story of round six, the round that taught us the most, and the first one Codzee won.
The rules we set for ourselves
Because we were testing our own product, we didn't trust ourselves. We set two rules:
- The answer key comes first. We wrote down every planted bug before the pull request was opened, so nobody could redefine what counts as a "catch" afterward.
- Publish the bad news too. If our tool stumbled, it goes in the write-up.
Here's the setup for round six:
| PR | #7 in a test repository |
| Size | 192 lines across two modules |
| Planted bugs | 13 |
| Cross-file bugs | 6 |
| Domain | Guest billing: invoices, payments, tax, cancellations |
| Tools | Codzee, Greptile, CodeRabbit, CodeAnt |
| How we scored | Pulled every finding from the GitHub REST API and matched them to bugs by hand |
Qodo was excluded from this benchmark for billing reasons.
The trap we built
Most AI review demos show a tool catching a bug in one function. That's easy mode, so we designed round six around something harder.
The two modules were tied together by a written contract. It said things like:
- The invoice total is pre-tax.
- The balance reads a key called
amount_pence. - The due date is a date, not a datetime.
- Invoice numbers are unique for life.
Six of our 13 bugs were cross-file defects: each module looks correct on its own, and the error exists only in how they fit together. You can't find them by reading either file. You have to read both against the contract.
The centerpiece was this one:
# Illustrative sketch, not copied from the PR
invoice = build_invoice(booking) # built...
charge = rates.compute(booking) # ...ignored, then recomputed
gateway.charge(guest, charge) # guest sees one number, pays another
The payment path builds an invoice, throws it away, and recomputes the charge from another module. The guest is billed a figure that appears nowhere on their invoice, with tax and any extra lines dropped. Nothing crashes, and every function does its own job correctly.
The first thing we noticed: speed
Codzee finished in 57 seconds. CodeAnt took about two minutes. Greptile took 3 minutes 43 seconds. (We didn't record a time for CodeRabbit.)
We wouldn't read much into this, because the timings may not be directly comparable. But it was the first visible difference, before anyone looked at the findings.
The second thing we noticed: nobody was wrong
We pulled every finding from the API and matched them against the answer key. There were zero false positives across all four tools. Every finding anyone posted was real, so in this round the differences weren't about who was wrong. They were about who saw more, and where they put what they saw.
The seam
Here's how the four tools did on the six cross-file bugs:
| The bug | Codzee | Greptile | CodeRabbit | CodeAnt |
|---|---|---|---|---|
| Payment ignores the invoice, recomputes the charge | ✓ | ✓ | ✗ | ✗ |
| Tax is never collected | ✓ | ✓ | ✓ | ✗ |
| Balance reads a key the invoice never writes | ✓ | ✓ | ✓ | ✓ |
| Due date compared against a datetime | ✓ | ✓ | ✓ | ✓ |
A timedelta passed where a day count is expected |
✓ | ✓ | ✓ | ✓ |
| Invoice numbers restart, payment records collide | ✓ | ✓ | ✓ | ✓ |
| Total | 6/6 | 6/6 | 5/6 | 4/6 |
Four of the six were caught by everyone. These were visible mismatches, such as a wrong key, a wrong type or a counter that resets. A tool sees a value in one place and the wrong shape of it in another.
The top two rows, the money bugs, were different. They split the field.
Only Codzee and Greptile found the recomputed charge, and in our assessment Codzee's description was the more complete of the two. It was also the one bug Codzee marked Critical, which felt right, since it's the only defect in the PR that silently bills a guest the wrong amount.
The most interesting result belonged to CodeRabbit. It commented on the tax problem in the same function but never noticed that the function was also ignoring the invoice it had just built. It saw the symptom and not the cause sitting a few lines away. We can't see inside the tool, so we can't say why. It was exactly the gap we'd designed the round to expose.
CodeRabbit did something else well, though. Its fix guidance on the tax bug was the clearest of the round: include tax in both the amount due and the gateway charge, with both call sites named.
The plot twist: the bugs we didn't plant
This is the part of the round that changed how we think about our own tool.
Our answer key is limited by our imagination. Real code has bugs nobody planted, so we also counted genuine findings outside the key.
For the first five rounds, this was the criticism we couldn't answer: Codzee had found zero of them. Greptile had found eleven.
In round six, Codzee found three. Two were serious, and no other tool reported either one:
The declined card that counted as paid. The payment path never checks the gateway receipt. A declined card is recorded as a full payment, so the invoice shows as settled and no money was taken.
The retry that charges twice. After a timeout, a retry builds a new invoice with a new number. That changes the idempotency key, and the guest gets charged twice.
Both require following control flow across steps, not reading a single line.
Greptile found two of its own, and both are about durability. An unsupported country code silently inherits the 20% GB rate, so an invoice looks plausible with a tax figure nobody chose. And payment records live only in an in-memory dictionary, so a restart wipes the history that balance and refund calculations rely on.
CodeAnt found one, in keeping with its style, which was precise and focused on a single function. Its findings included an unsynchronized read-and-increment on the invoice counter, unvalidated payment amounts, and a getter that hands back the live internal list so callers can edit the ledger.
The test file nobody opened
Three tests in the diff asserted the wrong behavior.
That's a nasty class of bug, because a passing test makes a defect look intentional. A reviewer who sees green reasonably assumes the behavior is correct.
Only Codzee opened the test file. This is now the fourth round out of six where it's been the only tool to do so. It posted one finding per bad test, each naming the correct expected value:
- The due-date test should expect 15 June, not 1 July.
- An unpaid invoice should show 22,300 pence outstanding, not zero.
- Asserting that amount due equals the total just restates the implementation.
Greptile, CodeRabbit and CodeAnt had nothing to say about the test file.
The scoreboard
| Codzee | Greptile | CodeRabbit | CodeAnt | |
|---|---|---|---|---|
| Bugs found (of 13) | 8 | 7 | 6 | 6 |
| Cross-file (of 6) | 6 | 6 | 5 | 4 |
| Total findings | 18 | 9 | 6 | 7 |
| Real bugs we didn't plant | 3 | 2 | 0 | 1 |
| Findings on test file | 3 | 0 | 0 | 0 |
| False positives | 0 | 0 | 0 | 0 |
Codzee found the most bugs, had a perfect cross-file score and was the only tool with test-file findings. But the margin over Greptile was one defect on one PR. Greptile matched us on the hardest dimension with half as many findings.
What none of us caught
Three bugs defeated all four tools:
- A mutable default argument, used in two functions
- A function-local import that hid the recompute from view
- A truncation that quietly loses a penny on every invoice
The first is a textbook lint finding. In round one, nobody caught banker's rounding or an unreachable return either. As far as we can tell, none of these tools runs a linter, though we haven't verified how each works internally.
We don't know why every tool missed these. Our hypothesis, and it is only a hypothesis, is that these tools put their effort into semantic reasoning and less into mechanical checks and subtle numeric behavior.
The practical lesson: an AI reviewer doesn't replace your linter.
The part where our own tool let us down
If we stopped here, this would be a flattering story, and we promised it wouldn't be.
Problem 1: almost half the findings weren't on the code. Of Codzee's 18 findings, only 10 were posted inline on the diff. The other 8 sat in the summary, including all three test findings, the ones that most need to sit next to the code they describe. Across six rounds, our inline rate has been 50%, 100%, 0%, 75%, 25%, 56%, with no trend. Every other tool anchored its findings inline.
Someone reading the pull request sees less of the win than the analysis shows.
Problem 2: eighteen findings and no verdict. Codzee flagged a Critical that bills guests incorrectly, and the check conclusion was still neutral. A neutral check can't block a merge. Greptile found fewer bugs and still reported a confidence score.
There was one thing we got right, though. In round five, Codzee had graded a worthless test suite as Minor, which was poor severity calibration. In round six it used three levels (1 Critical, 12 Major, 5 Minor), the Critical went to the right bug, and the bad tests were graded Major. Running the same experiment repeatedly let us confirm that a problem we found actually got fixed.
The six-round story
One round is an anecdote, so here's the full ramp:
| Tool | R1 | R2 | R3 | R4 | R5 | R6 |
|---|---|---|---|---|---|---|
| Greptile | 5/6 | 10/12 | 3 | 6/6 | 9/9 | 7/13 |
| Codzee | 4/6 | 9/12 | 2 | 3/6 | 8/9 | 8/13 |
| CodeRabbit | 4/6 | 8/12 | 1 | 5/6 | 8/9 | 6/13 |
| CodeAnt | 1/6 | 8/12 | 3 | 4/6 | 6/9 | 6/13 |
The honest reading:
- Greptile is the most consistent tool in the series. It has never had a weak round and has found 13 unplanted bugs across six.
- Codzee is rising. It got its first win, its first unplanted finds and a fixed severity problem, but it started behind.
- CodeRabbit is steady and precise, and consistently the lowest-volume tool.
- CodeAnt is strong wherever a defect is visible in one function and weaker wherever it requires comparing two things.
What we learned
1. The real test is reasoning distance. Every tool can read a function. The ones that stand out can read a function against a contract in another file.
2. The hardest bugs are absences. An ignored invoice and a missing receipt check are things the code didn't do, and they're harder to flag than a line that's visibly wrong.
3. More comments aren't a better review. Greptile got six of six cross-file from nine findings. We posted twice as many and found one more bug.
4. Read the tests. Three bad tests were quietly vouching for three bugs, and only one tool noticed.
5. The best findings are often unplanted. The declined card and the double charge weren't in our key, and they may be the two most serious findings in the PR.
6. Placement is part of quality. A finding that isn't next to the code is easy to miss.
Steal this checklist
You can use it even if you never touch Codzee.
- Plant a cross-file bug. Have one module quietly ignore what another returns.
- Plant a bad test. See whether the tool opens the test file.
- Count findings, then count what you'd act on.
- Check where comments land, inline or buried in a summary.
- Check severity. Does "Critical" mean critical?
- Check the merge signal. Can it block a merge or only comment?
- Read the fix guidance. Does it name call sites and expected values?
- Track unplanted finds.
- Keep your linter.
- Test what we didn't: security depth, CI/CD integration, customization, scale and team adoption were outside this experiment.
Limitations
This was an observational comparison, not a statistically significant benchmark.
- It was one PR in this round: 192 lines, one domain, and probably one language.
- We wrote the answer key. Codzee's team designed the bugs and has a stake in the result.
- Matching findings to bugs was done by hand, which involves judgment.
- Tool versions and configurations may differ across tools.
- Time figures may not be directly comparable.
- Tools change fast. These results describe the day of the run.
Nothing here shows that any tool is universally better.
What happens next
Round seven pushes reasoning distance further. The same quantity will be metres in one module and kilometres in another, the classic units-crossing-a-boundary bug. Every tool missed that class in an earlier benchmark.
We expect it to be hard, and we'll publish the results either way.
The takeaway
We started this series wanting to know whether our tool was any good. After six rounds, the answer is more useful than yes or no.
It's good at some things, such as reading across files, reviewing tests and finding bugs nobody planted. It's still weak at others, namely inline placement and merge verdicts. And it's competing with tools that are better in places: Greptile on consistency, CodeRabbit on precision, CodeAnt on single-function hazards.
If you're choosing an AI code review tool, don't trust a feature list, ours included. Hide a bug in the gap between two files, put one wrong test in the diff, and see who notices.
We build Codzee and ran this benchmark ourselves. The methodology and limitations are above. We'd rather you check our work than trust it.
What kinds of bugs do you find AI reviewers consistently miss? Tell us in the comments.
Top comments (0)