Your AI assistant just generated tests until the coverage bar turned green. 92%. Ship it?
Here's the trap: coverage measures which lines ran, not whether a test would notice when they're wrong. AI made that gap free — you can generate 90% coverage in a minute, all of it happy-path, none of it discriminating. That's coverage theatre: it looks like safety, it measures activity, it proves almost nothing.
Why coverage lies
A line is "covered" the moment a test executes it. But executing a line and asserting the right thing about it are different claims. An AI-generated test that calls applyDiscount(order) and asserts typeof result === 'number' covers the function — and would stay green if the discount math were completely wrong.
Coverage answers "did this code run in a test?" The question that earns trust is "would a test fail if this code were wrong?" Those are not the same, and AI output routinely nails the first while skipping the second.
The antidote is old and boring: break the code on purpose
If a suite can't tell correct behavior from a real defect, its coverage number is decoration. The cheapest check is mutation: change > to >=, flip a boolean, swap two lines — then see if anything goes red. If the suite stays green on broken code, it isn't protecting you.
I put the smallest runnable version of this in a free lab (no install, Node 20+):
node run-lab.js all
It hands the AI's own suite a pricing function with four seeded defects, then runs an independent verification suite against the same broken code:
WEAK TESTS (the AI's suite) .... PASS 3/3
STRONG VERIFICATION ........... FAIL 8/12 <- catches all 4 defects
MUTATION GATE ................. PASS 4/4
LAB GATE: PASS
The AI's tests pass against code that is wrong. Verification is what fails — and that failure is the signal coverage never gave you.
What to do instead of chasing the coverage bar
- Before you trust a generated test, ask: what realistic broken version of this code would still make it pass? If you can picture one, the test isn't discriminating.
- Add a mutation or seeded-defect check to CI. One flipped operator per critical function exposes a suite full of decoration.
- Keep the thing that writes tests separate from the thing that approves them. The generator proposes; an independent reviewer (a different model or session) checks; a human owns the release call.
- Report mutation score, not just coverage %. Coverage tells you what ran; mutation tells you what your tests can actually catch.
Coverage isn't useless — low coverage is a real red flag. But high coverage is not evidence of anything on its own, and AI made it trivially easy to manufacture. Treat the green bar as a starting question, not an answer.
Try it, then go deeper
Free, self-contained lab: https://github.com/chernevnikolay86-wq/ai-test-verification-lab
I turned the full method into a book + runnable kit — the operating loop, governance, and labs across unit, web/API, instrument-cluster and CAN/UDS domains.
- Kindle on Amazon ($9.99): https://www.amazon.com/dp/B0HL4SC5PX
- Full bundle (PDF + EPUB + labs) on Gumroad ($19): https://nikolaychernev.gumroad.com/l/ai-testing-in-practice
Honest caveat: the lab systems are simulators with declared defects — they prove the method catches faults, not that any product is certified. That's the point: don't trust things that merely look right.
Top comments (0)