This is a submission for the Kaggle Benchmarking Challenge
What I Benchmarked
I keep seeing posts about using AI to review pull requests before they get merged. That made me want to test one specific thing: can an LLM catch a security bug when nobody tells it to look for one?
So I built a benchmark called "Plausible PR." It has 10 pull request diffs. Each one is written to look like a harmless cleanup: collapsing an if-statement, swapping a comparison, "simplifying" a query. Every single diff actually removes something that was protecting the app: a permission check, a rate limit, a constant-time comparison, an input boundary check.
The prompt is always the same: "Is this PR safe to merge?" I never tell the model to look for security issues. I just want to know if it notices on its own, the way a careful human reviewer would.
Models Tested
I ran the benchmark against four models:
- Claude Sonnet 5 (Anthropic)
- GPT-5.5 (OpenAI)
- Gemini 3.7 Flash (Google)
- DeepSeek-R1 (DeepSeek)
I picked one model from each major lab, plus an open-weight model, so the comparison covers the range people actually choose between when they wire an LLM into a review workflow.
Findings
Here is the full board, pass or fail per task per model:
| # | Task | Claude Sonnet 5 | DeepSeek-R1 | Gemini 3.7 Flash | GPT-5.5 |
|---|---|---|---|---|---|
| 1 | Permission check collapse | PASS | PASS | PASS | PASS |
| 2 | Auth comparison crash | FAIL | FAIL | FAIL | FAIL |
| 3 | Retry on every exception | PASS | PASS | PASS | PASS |
| 4 | SQL injection | PASS | PASS | PASS | PASS |
| 5 | CORS wildcard | PASS | FAIL | PASS | FAIL |
| 6 | JWT algorithm confusion | PASS | FAIL | FAIL | FAIL |
| 7 | Timing-unsafe token compare | PASS | PASS | PASS | FAIL |
| 8 | Rate limit bypass | PASS | PASS | FAIL | PASS |
| 9 | Path traversal | PASS | PASS | PASS | PASS |
| 10 | Secrets logged | PASS | PASS | PASS | PASS |
| Total /10 | 9 | 7 | 7 | 6 |
Claude Sonnet 5 came out on top with 9/10. GPT-5.5 trailed at 6/10.
The result that surprised me most is task 2. It is a one-line change: a token check switches from comparing bytes to comparing a plain string with secrets.compare_digest. That function throws a TypeError on non-ASCII text instead of returning False. So a malformed login attempt no longer gets a clean 401, it crashes the request instead. All four models, including the one that scored 9/10 on everything else, called this a safe simplification and approved it.
The pattern I noticed: models are good at catching bugs that have a keyword to react to. "SQL" and string interpolation together set off an obvious alarm, so every model caught the SQL injection (task 4) and the path traversal (task 9). But bugs that only show up when you trace what happens to a slightly unusual input, like a malformed header or a non-ASCII string, get waved through even by the strongest model. Nothing about the diff "looks dangerous." It just quietly changes what happens on an edge case nobody wrote a test for.
If I ran this again, I would add a second round: give the model the same crash bug, but this time ask "does this handle all input types correctly," and see if a more targeted prompt catches what the open-ended "is this safe to merge" prompt misses. That would tell me whether this is a knowledge gap or a hint-dependence problem.
My Benchmark
Full benchmark on Kaggle: Plausible PR: Security Bugs Hidden in Clean Refact
Each task includes the full diff, the model's response, and the judge criteria used to score it, so you can see exactly why something passed or failed.

Top comments (1)
Task 2 is a great find: secrets.compare_digest raising TypeError on non-ASCII str is exactly the kind of behaviour you only know from the docs or from getting burned, and no keyword in the diff hints at it.
One thing I'd add before round two: every diff in the set is unsafe, so "No, don't merge" is the right answer ten times out of ten. A model that refuses everything would score 10/10 and look like the best reviewer. Mixing in genuinely harmless refactors of the same style, say ten of them, would let you report both halves: how often each model catches the bug, and how often it blocks a clean PR. In a real review workflow that second number decides whether people keep reading the bot's comments.
On the ranking: the models all saw the same ten tasks, so the fair comparison is paired. Sonnet vs GPT-5.5 differ on only three tasks (5, 6, 7), all in Sonnet's favour, and with three disagreements an exact sign test gives p = 0.25. That's suggestive, not a gap yet. Running each task a few times per model would help too, since a single sample per cell mixes the model's ability with the luck of one generation.