DEV Community

Kudzai Murimi
Kudzai Murimi Subscriber

Posted on

Just published a benchmark testing whether AI code reviewers catch security bugs nobody tells them to look for.
I ran 10 "harmless-looking" refactors past Claude Sonnet 5, GPT-5.5, Gemini 3.7 Flash, and DeepSeek-R1 — all four missed the exact same one-line

Sign in to view linked content

Top comments (0)