DEV Community

Shaban Umar
Shaban Umar

Posted on

I showed AI the buggy code. Its tests started protecting the bug.

Kaggle Benchmarking Challenge Submission

This is a submission for the Kaggle Benchmarking Challenge

What I Benchmarked

I've been a QA engineer since 2023. The most dangerous test suite I've ever seen was a green one: every check passed, the build shipped, and the bug went out with it.

So when people say LLMs "write good unit tests", my question is: good at what? Passing is easy. A suite that asserts nothing passes. The job of a test is to fail when the code is wrong.

For the Kaggle Community Benchmarks challenge I built Green Tests, Shipped Bugs, a benchmark that measures exactly that, and it turned up the most useful thing I've learned about AI-written tests:

When a model can see the code, it often tests what the code does instead of what it should do, and a buggy function gets a test suite that protects the bug.

One sentence in the prompt mostly fixes it. And the best models don't need the sentence at all.

How it works

I wrote 12 small functions of the kind that sit in every codebase: pagination, URL slugs, money rounding, interval merging, first-occurrence binary search, business-day counting, IPv4 validation, duration parsing, string truncation, percent change, chunking and leap years. Each one has:

  • a precise docstring, which is the contract;
  • one correct implementation;
  • one planted bug, a realistic mistake someone could ship: page < 0 instead of page < 1, float rounding on money, a binary search that finds an occurrence instead of the first;
  • 2–3 extra mutants, meaning more small, plausible bugs.

The model writes a pytest suite. A sandbox runs it against every version. This is mutation testing:

  • Invalid: the suite fails on the correct code. It cries wolf, so it scores zero.
  • Caught: a valid suite fails on a buggy version.
  • Bug-locking test: a test that fails on the correct code but passes on the buggy code. It asserts the bug as if it were the spec.

Every mutant is checked against a hand-written reference suite, so each one can be caught by a test that simply follows the docstring.

Three conditions

condition the model sees
spec_only signature + docstring
code_shown the full implementation, which contains the planted bug
code_warned the same buggy code, plus: "this hasn't been reviewed and may contain bugs; test the docstring, not the code"

code_shown is how people actually use these tools: paste a function, type "write tests for this". Nothing in the prompt says there's a bug.

Models Tested

9 models through Kaggle's model proxy, with identical prompts: GPT-6 Astra, Claude Opus 5, Claude Sonnet 5, Claude Haiku 4.5, Gemini 3.1 Pro, Gemini 3.5 Flash-Lite, DeepSeek-R1, Qwen3-Coder 480B and GLM-5.

I picked them to spread across three axes: frontier vs. small (Opus 5 and GPT-6 Astra next to Haiku 4.5 and Flash-Lite), general vs. code-specialised (Qwen3-Coder), and closed vs. open-weight (DeepSeek-R1, Qwen3-Coder, GLM-5). That spread turned out to matter: the habit of copying the bug into the tests doesn't simply follow model size.

12 functions × 3 conditions = 36 suites per model, capped at 40 test cases and 10,000 output tokens each. (I planned a tenth, gpt-oss-120b, but the proxy answered "model is currently experiencing heavy load" on three separate days.)

Findings

Results

Planted bug caught, by condition, and bug-locking tests summed over all 36 suites, sorted by mutation score (the leaderboard metric):

model spec only code shown code + warning bug-locking tests mutation score
GPT-6 Astra 100% 100% 100% 0 1.00
Gemini 3.1 Pro 100% 92% 100% 0 0.97
Claude Opus 5 100% 100% 92% 0 0.97
Gemini 3.5 Flash-Lite 58% 75% 83% 1 0.72
Claude Sonnet 5 75% 33% 100% 53 0.69
DeepSeek-R1 75% 42% 75% 46 0.63
Claude Haiku 4.5 50% 75% 42% 16 0.56
GLM-5 58% 42% 58% 23 0.51
Qwen3-Coder 480B 25% 50% 42% 67 0.39

Across all models:

spec only code shown code + warning
planted bug caught 71% 68% 77%
bug-locking tests (total) 13 158 35
suites that lock the planted bug in 0 14 1

What I learned

1. Seeing the code multiplied the bug-locking tests

The headline isn't the catch rate, which only dips (71% → 68%). It's what the tests assert. With only the docstring, the models wrote 13 bug-locking tests between them. Given the same docstring plus the buggy code: 158. In 14 suites the planted bug itself was locked in: the suite passes on the buggy code and fails on the fix.

The model reads the implementation, works out what it returns, and writes that down as the expected value, faithfully including the bug.

It isn't universal, and the split is the interesting part. The three strongest models (GPT-6 Astra, Claude Opus 5 and Gemini 3.1 Pro) wrote zero bug-locking tests in any condition. Shown the buggy code, they tested the docstring anyway. Every one of the 206 bug-locking tests came from the other six. Flash-Lite and Haiku even caught more bugs with the code in front of them. So this is not "AI can't write tests"; it's a habit some models have, and you won't know which kind you're using unless you check.

2. Some models do it on purpose

The clearest case is Claude Sonnet 5, a strong model that caught 75% of planted bugs from the docstring alone and 33% once it saw the code, with 53 bug-locking tests, all of them in code_shown. The pagination function's planted bug lets page=0 through instead of rejecting it. Sonnet's suite contained:

test_page_zero_is_not_rejected_and_yields_empty_list
Enter fullscreen mode Exit fullscreen mode

The test's name describes the bug. The test requires it. The same pattern showed up across functions: truncate tests expecting a result one character longer than max_len, and slugify tests expecting "naïve résumé" → "na-ve-r-sum-", which is exactly the planted bug's output.

That's a real technique, characterization testing: pin down what legacy code does today before you refactor it. But "write tests for this function" doesn't say which one you want. Some models assume pin the current behaviour, and a pinned bug is a bug nobody will ever fix, because "the tests say it's correct".

3. One sentence fixes it (for most models)

Add "this may contain bugs; test the docstring, not the code" and the catch rate goes to 77%, higher than spec-only, with bug-locking tests down from 158 to 35. Sonnet 5 went from 33% to 100% with that one sentence. The model can do the right thing. It needs to be told which job it's doing.

The exception: Qwen3-Coder wrote 28 bug-locking tests even when warned. It mostly ignored the instruction.

4. The coverage is there; the expected values are wrong

If I throw away each suite's wrong tests (the ones that fail on correct code) and score what's left, the average kill rate goes from 72% to 94%. The models do exercise the right cases. What breaks the suite is the expected values: arithmetic slips (percent_change(-100, 100), weekend counts in business_days, Haiku claiming 1600 and 2000 aren't leap years) and values copied from the buggy code. Even Opus 5's one miss was of this kind: in the warned condition it expected business_days to count backwards when the end date comes first, where the docstring says to return 0. The two failure types feel different in practice: an arithmetic slip makes a suite fail on correct code and gets noticed; a copied bug makes a suite green and doesn't.

5. More tests ≠ better tests

Qwen3-Coder wrote the most tests per suite (45 on average, over the 40-case cap) and had the lowest score. GPT-6 Astra wrote 37 and was perfect. Without a cap, some frontier models happily generate several hundred parametrized cases for a 10-line function; an earlier run of mine hit the 10,000-token output limit, which is why the cap exists.

6. The models found bugs in my code

Honest bonus. My first run had bugs in three of my "correct" reference implementations: the money rounder crashed on 32-digit amounts, the duration parser accepted Arabic-Indic digits, and the IPv4 check crashed on a 5,000-digit octet despite a docstring promising it never raises. Several models' tests failed my "correct" code, and they were right. I fixed the references, tightened three docstrings and re-ran everything.

Then the final run found a fourth. GPT-6 Astra wrote test_rounding_is_independent_of_callers_decimal_context: it turns on strict decimal settings globally and checks that round_money still returns "12345.13". My version inherited the caller's settings and crashed. A money function shouldn't depend on global state, so I fixed it and re-graded that suite (no other suite was affected). The reference implementation in a benchmark needs testing too, and here the models did it.

What to do with this

If you ask an assistant for tests:

  1. Give it the contract, not just the code. Docstring, ticket, spec.
  2. Say which job you want. "Find bugs against this spec" and "pin current behaviour" are opposite instructions.
  3. Tell it the code may be wrong. It's one sentence, and in this benchmark it cut bug-locking tests from 158 to 35.
  4. Run the tests against a known-good version when you can. A test that fails on correct code is a bug report about the test, or, sometimes, about your code.

What I'd measure next

  • Harder mutants. The best models caught almost everything; off-by-one bugs in less familiar code (date maths across time zones, Unicode normalisation) would separate them further.
  • The repair loop. Show the model its own failing tests on the correct code and ask it to fix them. Does it fix the test, or "fix" the code to match the test?
  • Property-based tests. Do models that reach for Hypothesis lock in fewer bugs?

My Benchmark

The leaderboard score is the mutation score from a separate, fresh run of the task for each model, so it won't match the tables above exactly (those come from the per-condition analysis runs in the notebook). The leaderboard also has a tenth model, Gemini 3.7 Flash (0.88), which I only ran there.

leaderboard score
Gemini 3.1 Pro 1.00
GPT-6 Astra 1.00
Claude Opus 5 0.94
Gemini 3.7 Flash 0.88
Claude Sonnet 5 0.80
Gemini 3.5 Flash-Lite 0.72
DeepSeek-R1 0.67
GLM-5 0.67
Claude Haiku 4.5 0.61
Qwen3-Coder 480B 0.54

With 36 suites per model, a single run moves some models by 10–15 points (GLM-5 went from 0.51 to 0.67, Qwen3-Coder from 0.39 to 0.54), while the top three and the bottom of the ranking stayed put. If you build on this, run each model more than once. Adding a problem takes a docstring, a correct function and a few buggy variants. I'd love to see harder mutants, and more cases where the reference implementation is the one that's wrong.

Top comments (1)

Collapse
 
coldboothq profile image
Coldboot •

The split between "caught the bug" and "locked the bug in" is the useful part. A catch rate alone would have hidden most of what you found.

One thought on the conditions: in a coding agent, code_shown is the default. The agent has read the file before it writes a test, and the contract is often in a ticket and not in a docstring. A fourth condition might be worth trying: code shown, with the intended behaviour given separately as a short issue description. That is closer to how the request arrives in practice.

On harder mutants, one from a real app: I had tests that failed only in timezones ahead of UTC. A suite written and run in UTC never sees that class, so running the sandbox once with TZ set ahead of UTC could show it.

Did the bug-locking tests cluster in particular functions, or were they spread evenly across the 12?