DEV Community

Mahender Pratap
Mahender Pratap

Posted on

Building a C++ Logical Bug Detection Benchmark on Kaggle: Testing 3 AI Models

This is a submission for the Kaggle Benchmarking Challenge.

Introduction

As a C++ learner, I wanted to explore how well AI models can identify common logical errors in code and suggest appropriate fixes.

For the DEV × Kaggle Benchmarking Challenge, I created a small benchmark on Kaggle called C++ Logical Bug Detection Benchmark.

The goal is to test whether AI models can recognize common programming mistakes, explain why they occur, and suggest corrections.

What Does the Benchmark Test?

The benchmark contains three tasks:

1. Sum Accumulator

This task checks whether a model can identify incorrect variable usage in a summation loop and suggest the correct fix.

2. Off-by-One Error Detection

This task evaluates whether a model can identify incorrect loop boundaries and indexing that may cause an off-by-one error.

3. Array Indexing Error Detection

This task checks whether a model can detect incorrect array index usage and explain how to correct it.

AI Models Tested

I tested three AI models:

  • Qwen 3 Coder 480B
  • GPT-5.4 mini
  • Gemini 3.7 Flash

I selected these models to compare how different AI systems handle simple C++ logical bug-detection tasks.

Results

All three models achieved a score of 100.00 on the benchmark.

Model Score
Qwen 3 Coder 480B 100.00
GPT-5.4 mini 100.00
Gemini 3.7 Flash 100.00

All three models passed the three tasks, giving a total of 9 successful model-task evaluations out of 9.

Key Findings and Limitations

The results show that all three models handled the three simple bug-detection tasks successfully.

However, this is a small benchmark with only three tasks. It does not establish how well these models perform on complex C++ debugging, large codebases, or real-world software projects.

In the future, I would like to expand the benchmark with more challenging logical errors, nested loops, pointer-related bugs, and edge cases.

What I Learned

Building this benchmark helped me understand how AI models can be evaluated using specific programming tasks and measurable results.

It also encouraged me to think about how to design better tests that go beyond simple examples.

Try the Benchmark

You can explore my benchmark here:

C++ Logical Bug Detection Benchmark on Kaggle

I would appreciate feedback and suggestions for adding more challenging C++ debugging tasks.


Top comments (2)

Collapse
 
respect17 profile image
Kudzai Murimi •

All three models hitting 100% on 3 tasks says more about the benchmark's difficulty than about the models, and you call that out yourself, which is the right instinct. Pointer aliasing bugs or off by one errors buried inside nested loop bounds, not just the loop itself, would probably start telling them apart.

Collapse
 
mahenderpratap profile image
Mahender Pratap •

Thanks for the thoughtful feedback! I agree that the 100% scores highlight the simplicity of the current benchmark more than the models' capabilities. Pointer aliasing and off-by-one errors hidden in nested loops are great ideas for making the next version more challenging. I'll explore adding these cases to better evaluate the models. Appreciate the suggestions!