DEV Community

Who Reviews the Reviewers? Benchmarking AI Agents on PR Audits

What I Benchmarked
I benchmarked the ability of LLMs to conduct automated pull request (PR) auditing without falling into the "helpful hallucination" trap.

When building and orchestrating AI workflows for code reviews, a recurring failure mode constantly pops up: models often hallucinate phantom security flaws or logic errors in perfectly clean code just to have something to say. It’s as if the model feels obligated to leave a comment to prove it did the work.

I set out to measure this explicitly. I created a dataset of 100 PRs: 50 containing subtle, real-world vulnerabilities (like race conditions, improper state management, or injection risks) and 50 perfectly optimized, clean code snippets. The goal wasn't just to see if they could catch real bugs, but to measure their false-positive rate—quantifying exactly how often these agents "cry wolf" and hallucinate an issue just to justify a "Request Changes" review.

Models Tested
I ran this benchmark against Gemini 1.5 Pro, GPT-4o, Claude 3.5 Sonnet, and Llama 3.1 70B.

I chose this lineup because they represent the current frontier of models typically selected for multi-agent developer tools and automated CI/CD pipelines. Since my day-to-day involves integrating agentic AI into GitHub pipelines using orchestration platforms like n8n and LangChain, I needed to see how these specific models behave when acting as autonomous reviewers rather than just conversational coding assistants. I wanted to see which model is actually safe to trust with a repository's main branch without overwhelming developers with fake alerts.

Findings
The results completely changed how I approach agent evaluation and orchestration.

The "Yes-Man" Reviewer: Most models exhibited a strong bias toward leaving a critical comment. While Claude 3.5 Sonnet and Gemini 1.5 Pro were exceptionally accurate at identifying genuine security flaws, they still occasionally nitpicked clean code. Llama 3.1 70B struggled the most with false positives, frequently flagging standard, clean implementations as problematic.

Context Window Distraction: When I padded the PRs with irrelevant boilerplate files (simulating a messy, real-world pull request), the models' accuracy in spotting the actual bug dropped significantly. This proved that simply passing an entire repository's context into an agent drastically degrades its auditing precision.

The Zero-Shot Penalty: Prompting models with a simple "Review this code" led to a 40% higher hallucination rate compared to providing a strict system prompt like, "Only flag critical security or logical errors; if the code is clean, output 'APPROVE'."

What I'd measure next: Moving forward, I want to benchmark a multi-agent debate architecture—having an "Auditor" agent flag issues and a separate "Defender" agent verify them against the actual codebase before the webhook ever posts a comment to the developers.

Top comments (0)