This is a submission for the Kaggle Benchmarking Challenge.
Introduction
As a C++ learner, I wanted to explore how well AI models can identify common logical errors in code and suggest appropriate fixes.
For the DEV × Kaggle Benchmarking Challenge, I created a small benchmark on Kaggle called C++ Logical Bug Detection Benchmark.
The goal is to test whether AI models can recognize common programming mistakes, explain why they occur, and suggest corrections.
What Does the Benchmark Test?
The benchmark contains three tasks:
1. Sum Accumulator
This task checks whether a model can identify incorrect variable usage in a summation loop and suggest the correct fix.
2. Off-by-One Error Detection
This task evaluates whether a model can identify incorrect loop boundaries and indexing that may cause an off-by-one error.
3. Array Indexing Error Detection
This task checks whether a model can detect incorrect array index usage and explain how to correct it.
AI Models Tested
I tested three AI models:
- Qwen 3 Coder 480B
- GPT-5.4 mini
- Gemini 3.7 Flash
I selected these models to compare how different AI systems handle simple C++ logical bug-detection tasks.
Results
All three models achieved a score of 100.00 on the benchmark.
| Model | Score |
|---|---|
| Qwen 3 Coder 480B | 100.00 |
| GPT-5.4 mini | 100.00 |
| Gemini 3.7 Flash | 100.00 |
All three models passed the three tasks, giving a total of 9 successful model-task evaluations out of 9.
Key Findings and Limitations
The results show that all three models handled the three simple bug-detection tasks successfully.
However, this is a small benchmark with only three tasks. It does not establish how well these models perform on complex C++ debugging, large codebases, or real-world software projects.
In the future, I would like to expand the benchmark with more challenging logical errors, nested loops, pointer-related bugs, and edge cases.
What I Learned
Building this benchmark helped me understand how AI models can be evaluated using specific programming tasks and measurable results.
It also encouraged me to think about how to design better tests that go beyond simple examples.
Try the Benchmark
You can explore my benchmark here:
C++ Logical Bug Detection Benchmark on Kaggle
I would appreciate feedback and suggestions for adding more challenging C++ debugging tasks.
Top comments (2)
All three models hitting 100% on 3 tasks says more about the benchmark's difficulty than about the models, and you call that out yourself, which is the right instinct. Pointer aliasing bugs or off by one errors buried inside nested loop bounds, not just the loop itself, would probably start telling them apart.
Thanks for the thoughtful feedback! I agree that the 100% scores highlight the simplicity of the current benchmark more than the models' capabilities. Pointer aliasing and off-by-one errors hidden in nested loops are great ideas for making the next version more challenging. I'll explore adding these cases to better evaluate the models. Appreciate the suggestions!