DEV Community

Nikolay Chernev
Nikolay Chernev

Posted on

Your AI Wrote Tests With 100% Coverage. They Catch 4% of Real Bugs.

That's not a hypothetical split — it's a documented test suite: 100% line coverage, 4% mutation score. Every line executed. 96% of seeded bugs slipped through anyway. Source.

Coverage tells you which lines ran. It says nothing about whether the test would notice if the logic were wrong. That gap is exactly where AI-generated tests tend to live: fluent, green, and checking almost nothing.

A fresh data point: what happens when a second AI checks the first one's work

Two weeks ago, First Mate Technologies published results from building QueueMate, a restaurant queue app, with a deliberate builder/checker split: one model writes the implementation, a separate model designs test cases and reviews the code before it merges. The checker produced 554 test cases. The resulting sweep found 38 defects — including a severity-one bug in session handling that the builder's own tests had missed. (developer-tech.com)

That's the whole thesis of verification-first testing in one case study: the model that wrote the code is a bad judge of whether the code is right. You need a second, independent check — and ideally not just "another AI," but something that actually probes behavior.

Also this past week: Momentic launched "Mo," an agent that drives an app the way a human would and reports reproduction steps and video evidence rather than asserting against its own output (SiliconANGLE, Sep 28) — a sign the industry is converging on the same answer: separate the generator from the grader.

The number that exposes the gap: mutation score

Here's the practical tool, and it's not new or exotic — mutation testing has been around for decades, it's just newly relevant:

  1. A mutation testing tool (Stryker, PIT, mutmut, depending on your stack) automatically changes a small piece of your code — flips a > to >=, deletes a line, swaps a boolean.
  2. It reruns your test suite against that mutated code.
  3. If a test fails, the mutant is "killed" — your suite caught it. If every test still passes, the mutant "survives" — your suite didn't actually verify that logic, no matter what the coverage report says.

Gartner's February 2026 insights abstract on this put it plainly: relying on AI to expand test coverage leads to missed defects, because AI-generated tests often lack depth and precision — and recommended mutation testing specifically as the corrective layer (Gartner, Sushant Singhal & Erin Khoo).

A mutation score isn't a vibe. It's a number that tells you, mutant by mutant, whether your tests would actually notice if the logic broke.

Seeing it without installing a mutation framework

You don't need a full mutation testing setup to feel this gap in under ten minutes. I built a free, runnable lab that does the same thing by hand: four bugs seeded into working-looking code.

  • node run-lab.js weak → the AI-generated test suite passes 3/3. On broken code.
  • node run-lab.js verify-candidate → an independent verification suite fails, and catches all 4 seeded bugs.

Same mechanism as a mutation score, minus the tooling setup: deliberately broken logic, and a direct answer to "would my tests have caught this?" Repo: https://github.com/chernevnikolay86-wq/ai-test-verification-lab

What to actually do about it

  • Don't grade an AI's tests by whether they're green. Grade them by whether they'd fail on a plausible wrong version of the code.
  • If you can run a mutation tool for your stack, run it on AI-generated test files specifically — the survival rate is the honest coverage number.
  • If you can't, write (or have a second, independent pass write) 2-3 deliberately broken variants of the function under test and check whether the existing suite catches any of them. That's mutation testing with a manual crank.
  • Treat "the AI's own tests pass" as the start of verification, not the end of it — the question that matters is "real bug, or model failure?", and only an independent check answers it.

This is chapter-one territory in AI-Assisted Software Testing in Practice, a field guide for testers and SDETs who already know the fundamentals and now need a working method for verifying AI-written tests rather than trusting them. No certification, no platform claims — a runnable kit and the reasoning behind it.

Make AI-generated tests you can actually trust. Don't trust it — run it.

Top comments (0)