Your AI just generated a test. Before you trust it, run this 15-second check — no tools, just one question.
"What broken version of this code would still make this test pass?"
If you can picture even one realistic wrong implementation that stays green, the test isn't checking behavior. It's decoration — delete it or fix it.
Why this one question works
A test earns trust only if it fails when the behavior is wrong. AI-generated tests reliably do the two easier things: they run, and they agree with the code the model just wrote — bugs included. The 15-second question forces the third, the one that matters: discrimination.
Take a generated test like this:
test('applies discount', () => {
expect(typeof applyDiscount(order)).toBe('number')
})
What broken applyDiscount would still pass? Almost any of them — wrong rate, wrong rounding, no discount at all. It returns a number every time. Covered, green, worthless.
Now the discriminating version:
test('10% over $2000, exactly at the boundary', () => {
expect(applyDiscount({ total: 2000 })).toBe(2000) // boundary: not discounted
expect(applyDiscount({ total: 2001 })).toBe(1800.9) // 10% off
})
A > where the spec says >= now turns this red. That's a test.
Make the habit concrete (free, 10 min)
I built a tiny lab (no install, Node 20+) that runs the whole failure mode:
node run-lab.js all
It hands the AI's own suite a pricing function with four seeded defects. The AI's tests pass 3/3 against the broken code; an independent verification suite fails and catches all four:
WEAK TESTS (the AI's suite) .... PASS 3/3
STRONG VERIFICATION ........... FAIL 8/12 <- catches all 4 defects
Open the flawed file and try to write the smallest test that catches each defect before you read the verification suite. That exercise is the whole method.
The loop around the question
The 15-second test is the pocket version of a bigger habit:
- Design the discriminating test first, then generate.
- Keep the generator separate from the approver — a different model or session reviews it.
- Prove the suite catches faults with seeded defects and mutation.
- A human owns the release call.
Try it, then go deeper
Free, self-contained lab: https://github.com/chernevnikolay86-wq/ai-test-verification-lab
I turned the full method into a book + runnable kit:
- Kindle on Amazon ($9.99): https://www.amazon.com/dp/B0HL4SC5PX
- Full bundle (PDF + EPUB + labs) on Gumroad ($19): https://nikolaychernev.gumroad.com/l/ai-testing-in-practice
Honest caveat: the lab systems are simulators with declared defects — they prove the method catches faults, not that any product is certified. Don't trust it — run it.
Top comments (0)