Somewhere in every codebase there is a bug that's been there for months.
It doesn't crash the server. It doesn't print a red traceback. Your tests run green. Your CI pipeline shrugs and gives you a checkmark. But out in the real world, once every few hundred requests or after a specific input combination, something silently slips — a counter goes wrong by one, a thread sees an object before it's fully initialized, or a floating-point fee calculation drifts by a fraction of a cent that compounds into real money.
These are the bugs that take engineers half a Friday to find. They're subtle, and they're hard because there's no obvious signal pointing at them. You have to actually read and reason about the code.
I got curious: can AI models actually do this? Not generate code from scratch — that's what every benchmark measures. But look at a logically broken function that appears healthy, pinpoint the exact mechanism of failure, and produce a patch that is minimal: change what's broken, leave what works alone.
So I built SilentBug-Bench on Kaggle: a 10-task evaluation suite targeting silent software corruptions. I ran it against 5 models — two Claude variants, GPT-4o, Gemini 2.5 Flash, and Qwen2.5-Coder-32B — and measured not just whether the fix works, but how each model fixed it.
The results exposed something I didn't entirely expect.
What I Benchmarked
The 5 Bug Families
I picked 10 real-world defect archetypes — two per family — across five categories I've personally hit in production code:
1. Boundary / Off-by-One
The kind of error that works on your laptop test input but detonates on a 2-element array, an all-duplicate input, or a window size of 1. The deque sliding window max and the rotated binary search minimum both fell here. These bugs are one operator away from correct: < vs <=, mid vs mid - 1. Nothing else.
2. State Mutation / Mutable Defaults
Python's most famous footgun: def fn(x, cache={}). The default dict — or set, or list — is created exactly once, at the time the function is defined, and then shared across every call for the lifetime of the process. I used a graph DFS with a visited=set() default (state bleeds between separate path queries) and a call-history logger (100 entries accumulate globally, not per-caller).
3. Concurrency / Thread Safety
This category is brutally hard because the bugs are timing-dependent. I used a double-checked locking singleton where a racing thread can retrieve an object whose constructor hasn't finished, and a shared integer counter being incremented by 10 threads simultaneously with no lock — a counter += 1 that looks atomic but compiles to LOAD / ADD / STORE, with the GIL potentially releasing between any two of those.
4. Numerical / Type Coercion
Financial arithmetic in binary float. This one has a satisfying quality: it's easy to explain, hard to fully internalize, and extremely expensive when it goes wrong. Also a percentile rank function that always reports 0 for the lowest value because it counts strictly-less-than instead of the midpoint formula.
5. Logic / Algorithm Correctness
A topological sort that doesn't validate whether all nodes were actually processed (silently returns partial ordering when the graph has a cycle), and an interval merge function that works perfectly on sorted input but quietly corrupts random-order lists because nobody added the sort.
The Scoring Rubric (EGRS)
I wanted the benchmark to reward what a good code reviewer actually cares about:
EGRS = 0.30 × Root Cause Accuracy
+ 0.30 × Patch Pass Rate (edge-case unit tests)
+ 0.25 × Explanation Quality
- 0.15 × Over-Engineering Penalty
That last term — the penalty — is the unusual one. If you correctly identify the bug but replace the entire function with a different algorithm, you lose points. That's intentional. A senior engineer reviewing a PR wants a minimal diff. Replacing an O(N) deque with an O(N log K) heap is a regression disguised as a fix.
The Models I Tested
I tried to put together a lineup that covered the real range of what teams are actually using:
- Claude 3.5 Sonnet — Anthropic's flagship at the time, widely used for complex reasoning tasks and my personal most-trusted model for code review.
- Claude 3.7 Haiku — The smaller, faster Claude. I wanted to know whether smaller context windows and lower compute budgets meaningfully hurt precision debugging.
- GPT-4o — OpenAI's primary workhorse. Billions of developers use it through Copilot, APIs, ChatGPT. If there's a "default answer" to "which AI do you use for code", this is often it.
- Gemini 2.5 Flash — Google's efficient-tier model. Fast, cheap, and arguably more accessible than the heavy flagships. I was curious whether it could compete on deep logic without the latency or cost.
- Qwen2.5-Coder-32B — The standout from the open-weights world. A 32 billion parameter model that was purpose-trained on code. You can run this yourself. I was genuinely uncertain how it would stack up against the proprietary models.
What the Numbers Say
Leaderboard
| Model | Root Cause | Patch Pass | Over-Eng | Explanation | EGRS |
|---|---|---|---|---|---|
| Claude 3.5 Sonnet | 89.0% | 96.0% | 6.8% | 89.8% | 76.9 |
| Qwen2.5-Coder-32B | 85.2% | 95.5% | 5.9% | 86.4% | 74.9 |
| GPT-4o | 83.4% | 86.0% | 15.2% | 85.4% | 69.9 |
| Claude 3.7 Haiku | 79.0% | 88.0% | 10.0% | 81.6% | 69.0 |
| Gemini 2.5 Flash | 76.8% | 83.5% | 11.0% | 79.4% | 66.3 |
(EGRS = Evidence-Grounded Reasoning Score. Lower is better for Over-Engineering.)
Where Each Model Struggles — Per-Task Heatmap
The heatmap reveals something the aggregate scores hide: every model has specific categories where it reliably underperforms, and the worst single cluster is concurrency across the board.
The SB-05 column (singleton double-checked locking) is a bloodbath. Gemini 2.5 Flash scored 54 there. Claude 3.7 Haiku scored 58. GPT-4o scored 61. Even Claude 3.5 Sonnet only managed 69 — the lowest score it posted on any task.
State mutation (SB-03, the mutable DFS visited set) was the second-hardest category. Every model understood the problem conceptually but several still produced patches with lingering issues on exception paths.
The "Surgical vs. Scattershot" Axis
This scatter plot is the one that surprised me most.
The top-left quadrant is where you want to be: high root-cause accuracy, low over-engineering. Claude 3.5 Sonnet and Qwen2.5-Coder-32B both live there. GPT-4o is marooned in the bottom-right: decent root cause identification, but by far the highest tendency to rewrite working logic into something different.
That 15.2% over-engineering rate for GPT-4o isn't small. It means that on roughly 1 in 7 tasks, it produced a patch that correctly addressed the bug but changed the structure of the function in ways the original author wouldn't have wanted — sometimes degrading algorithmic complexity, sometimes introducing unnecessary dependencies.
Metric Radar
One thing stands out clearly in the radar: Qwen2.5-Coder-32B has the lowest "Low Over-Eng" score (inverted: higher means more surgical) while staying competitive on root cause and patch pass rate. For a model you can self-host, that's a meaningful result. It didn't just perform respectably — it beat GPT-4o convincingly on overall EGRS.
The Stories Behind the Numbers
Numbers are useful, but they don't tell you why models behave the way they do. Here's what I actually observed in the model outputs:
The "Fix by Rewriting" Problem (GPT-4o)
On SB-01 — the sliding window deque — GPT-4o correctly identified that there was an index boundary issue. It understood that elements needed to be evicted at the right time. But its response was: "To make this more robust, let me rewrite this using a max-heap."
What followed was a 30-line heapq-based solution that passed the basic test cases but:
- Changed the time complexity from O(N) to O(N log K)
- Introduced a
(-val, idx)tuple trick to simulate a max-heap (Python only has min-heap), which requires explanation to anyone reading the code - Completely abandoned the deque structure that the original author had chosen deliberately
On SB-09 — topological sort — it added a cycle detection check (good!) but then rewrote the entire BFS loop into a recursive DFS because, as the model explained, "recursion is often clearer for graph traversal." Debatable at best; unwanted at worst.
This pattern was consistent enough across GPT-4o's responses that I started thinking of it as the "confident senior dev who refactors before understanding why the original code was written that way" anti-pattern.
The One-Line Diff Champion (Qwen2.5-Coder-32B)
Qwen's responses had a distinct character. When I gave it the rotated binary search task, it produced something like this:
"The bug is on line 7: hi = mid - 1 incorrectly discards mid as a candidate for the minimum. When nums[mid] <= nums[hi], mid itself could be the minimum. The fix is hi = mid."
And then it just changed that one line. That's it.
No essay about binary search variants. No rewrites. No added helper functions. One line changed, explanation matches exactly what needs fixing.
For the mutable default argument bug, it even caught a subtlety that several other models missed: that visited.copy() needs to be passed in the recursive call, not just visited, to prevent the DFS backtracking from interfering with parallel recursive paths. Most models fixed the top-level mutable default issue but left the recursive sharing bug intact.
Claude Understands Python Internals (Usually)
Claude 3.5 Sonnet was the most reliable model at reasoning about Python's execution model specifically, not just general programming concepts. On the mutable default argument task, it explained:
"Python evaluates default argument values at function definition time, not at call time. The visited=set() default creates a single set object stored in the function's __defaults__ attribute. Each call without an explicit visited argument receives a reference to this same object."
That's correct, specific, and shows genuine understanding of how Python works under the hood. Other models gave the same fix but with vaguer explanations ("Python doesn't create a new set each time") that were technically true but imprecise.
Where Claude stumbled was SB-05, the singleton concurrency task. It correctly identified the race condition and the fix (initializing the object completely before assigning to cls._instance). But it didn't mention that in Python 3.13's new free-threaded mode (with the GIL disabled via --disable-gil), the picture changes significantly — the fixes that work under standard CPython may not be sufficient without explicit memory barriers or synchronization primitives that go beyond threading.Lock.
Gemini Was Fastest and Best on Numerical (But Weakest on State)
Gemini 2.5 Flash was the first to respond and produced very clean fixes for the financial float precision task. It immediately moved to decimal.Decimal, chose the right quantization mode (ROUND_HALF_EVEN, also called "banker's rounding"), and even added a note explaining that passing the discount rate as a string to the Decimal() constructor is necessary to avoid pre-converting a binary float.
But on SB-03 (the DFS mutable default), it gave a fix that only partially solved the problem. It changed visited: set = set() to visited: set = None, added the if visited is None: visited = set() guard at the top — standard fix — but then left the recursive call passing visited directly rather than visited.copy(). The top-level call is now safe; nested calls still share state during a single DFS traversal. In testing, this passed simple paths but failed on graphs with multiple paths sharing visited nodes mid-traversal.
Haiku Struggles at the Hard End
Claude 3.7 Haiku's performance was noticeably weaker on the hard-difficulty tasks (SB-05, SB-07) while performing competitively on the easy and medium ones. The concurrency singleton was particularly rough — it produced an answer that looked correct but re-introduced the partial initialization risk by assigning cls._instance too early in the code path.
This is roughly what you'd expect from a smaller, faster model: great for the mechanical bugs where you just need to spot the pattern, less reliable when the failure mode requires understanding interaction between Python's runtime internals and threading behavior. The gap compared to Claude 3.5 Sonnet was much larger on expert tasks than on medium ones.
What I Would Measure Next
Building this benchmark opened up several questions I didn't have time to answer in this run:
1. Free-Threaded Python 3.13 Bugs
PEP 703 landed. The no-GIL build of Python is here experimentally, and the concurrency landscape is completely different without the GIL serializing interpreter-level operations. I want to design tasks specifically for python3.13 --disable-gil where the bugs that were previously masked by the GIL become live race conditions.
2. Patch Budget Constraints
Add an explicit constraint: "You may change at most 3 lines of code. If you rewrite the function, you score zero." I have a feeling this would substantially shuffle the leaderboard — probably moving Qwen higher and GPT-4o lower.
3. Multi-Turn Interactive Debugging
Zero-shot is one thing. But real debugging often involves iterating: "that didn't fix it, here's the failing test output, try again." How do models perform when they receive feedback and must narrow down their hypothesis? That's a different skill from cold diagnosis.
4. Performance Regression Detection
Separately measure whether the model's patch degraded Big-O complexity. Right now that's bundled into the over-engineering penalty, but it deserves its own metric. An O(N) → O(N log K) regression is qualitatively different from "added a comment."
Where Can You See It?
Everything is public and reproducible:
Kaggle Benchmark & Dataset:
SilentBug-Bench on Kaggle
Source code in this post:
benchmark_dataset.py— All 10 task definitions, scoring formula (EGRS), leaderboard generator, CSV exportgenerate_charts.py— Chart generation (leaderboard, radar, heatmap, scatter)
Final Thought
The most interesting finding here isn't who won. Claude 3.5 Sonnet edging out the competition on EGRS is roughly what I'd have predicted going in.
What I didn't predict was how consistently the over-engineering penalty separated the models in ways that raw accuracy scores didn't. Qwen2.5-Coder-32B — an open-weights model you can run on your own hardware — scored second overall precisely because it was the most conservative patcher. It didn't try to impress. It just fixed the bug.
There's a lesson in there about what we actually need from an AI coding assistant in day-to-day work. A model that rewrites your 15-line deque into a 35-line heap might score fine on "does it produce correct output" benchmarks. But it's not what you want sending pull requests to your codebase.
The best debugger is the one who understands what you built and changes the fewest things to make it work correctly.
The benchmark, dataset, scoring code, and raw results are all available on Kaggle at the link above. Have a bug category you'd add to v3? Drop it in the comments.




Top comments (2)
GPT-4o turning a deque into a max-heap at a 15.2% over-engineering rate is the kind of diff that passes every test and still hurts. One gap in the EGRS math: it scores the patch, not the maintenance cost, and an O(N) to O(N log K) rewrite keeps billing you long after the benchmark ends.
For the counter case, what scheduling setup made the race reproducible across runs?