Test-driven development and coding agents fit together unusually well. A test
is a precise specification the agent can check its own work against. The
red-green-refactor cycle keeps each change small enough to review. And the
"run the tests, read the output, try again" rhythm is exactly what these tools
are good at.
The guidance on how to do it has matured quickly, and most of it is sound. But
after working this way across different codebases and languages, with the agent
writing both the tests and the code, I kept finding tests that passed whether
the code worked or not.
Most of those were preventable. A handful of rules, written once into the
agent's instructions, stop the common shapes before they are written. The rest
needed a check after green, because no rule anticipates every way a test can
fail to discriminate.
This post is both: the rules worth writing down, and the thirty-second check
for what they miss.
The work behind this post used IBM Bob as the coding agent. Nothing here is
specific to it. Bob routes tasks across several frontier models rather than a
single fixed one, so a single session may not even have used the same model
throughout — and the same patterns turned up regardless of language, codebase
or tool.
Everything in this post is reproducible: sample project on GitHub
— six test shapes and a runner that mutates one line at a time.
Where the practice stands
It's worth tracing how the advice has developed, because each stage solved a
real problem.
Keep the tests in human hands. The early guidance was that if a model
writes both the code and the tests, the same assumption ends up in both and
they agree with each other. So humans wrote the tests and the AI wrote the
code. That division worked, and it's still good advice when the spec matters
more than the speed.
Write the discipline down. Agents can produce twenty tests in the time you
write one, and most teams took that trade. The guidance shifted to enforcing
the cycle through whatever instruction mechanism the tool offers — an
AGENTS.md in the project root now works across several of them, alongside
tool-specific rules files like .bob/rules for IBM Bob, CLAUDE.md for Claude
Code and .cursor/rules for Cursor — carrying the project's conventions and
the test-first rule, one behaviour per cycle, a human gate before
implementing, and, in the stronger versions, committing the red test so the
failing state lives in git history.
Kent Beck, who originated TDD, has published a system prompt along these lines: always
follow the cycle, write the simplest failing test first, implement the minimum
needed to pass.
Enforce it mechanically. The newest approach doesn't rely on the agent
following instructions at all. TDD Guard hooks every file write and blocks it
unless the process was honoured: a failing test exists, you're on one test at
a time, you're working outside-in. Under the covers it spins up a second model
as a judge on each edit, because "did this follow TDD?" is a fuzzy question
that is easier to ask a model than to encode in rules.
Each stage is a genuine improvement, and the progression is the right one.
Prompting alone tends to produce what one practitioner aptly called "test
first, not test-driven" — all the tests, then all the code, in two large
steps. Instruction files make the cycle stick more often. Hooks make it stick unless
the judge misses something.
Every tool I looked at governs sequence: was the test written first, is it
one behaviour, did red precede green. That's the hard part to enforce, and they
enforce it well.
The step I'm suggesting governs something different: whether the resulting
test can fail at all.
The rules: what prevents most of it
These go in whatever instruction mechanism your tool offers, and they are the
higher-value half of this post. Written once, they make most of the shapes in
the next section much less likely, though no instruction is followed perfectly:
- Only an assertion failure counts as red. A compile error or a missing symbol is not a failing test.
- Every negative assertion first proves the code ran. Assert the call count, then the absence.
- Assert the specific item, not that a collection is non-empty.
- Tests go through production construction — the real factory or constructor, never a hand-built object graph.
- The code under test and the test's own client never share an injected dependency. Separate instances, separate recorders.
- Report from the artifact, not from memory. Quote the changed lines or paste the raw test output.
- Plans make decisions. No "either/or", no placeholders, no pre-ticked checklists, and project terms quoted from their source.
- When a test's coverage is disputed, settle it with a mutation, not an argument. This is the one rule that needs the check, and it is the one that caught what the others missed.
The check, for what rules don't catch
Those rules prevent most of the shapes below. What follows is for the cases
they cannot anticipate.
After green, before moving on:
Break one production line on purpose. Say in advance which test should fail.
Run it. Confirm it fails on an assertion. Then restore the line.
That's mutation testing, done by hand, one line at a time, at review time
rather than in CI. Note the direction: the code is correct, you introduce a
deliberate defect, and the test failing is the good outcome — it's the test
doing its job.
A test that can't fail isn't a new problem, and it isn't specific to AI.
Mutation testing has existed for decades precisely because of it, and most TDD
guides mention it — usually a line in a metrics section beside coverage
thresholds, pointing at Stryker or mutmut. That framing makes it sound like a
quarterly exercise. Used per change, it's smaller and more immediate: the one
question that separates a test from a decoration.
What's different with agents is that TDD normally guards against this, and the
guard gets weaker. The red step is meant to be the proof — you watched the test
fail, so it can fail. That holds when a person writes one test and sees it go
red for the reason they expected. It holds less well when an agent produces the
test: the red often comes from the function not existing yet rather than from
the assertion, and the volume means few of them get inspected closely.
Better instructions narrow this a long way, and you should write them. But
instructions produce better tests; they do not produce proof that a given test
can fail. That distinction is the whole reason the check exists.
So the mutation asks what the red phase no longer reliably answers. A red test
proves it failed before the code existed — when everything failed. This asks:
now that the code exists, would this test notice if it broke?
Most of the time the answer is yes, it takes thirty seconds, and you move on.
Occasionally it isn't, and those are the cases worth writing about.
What that looks like
Two agent-written tests for the same behaviour: the service forwards the
caller's correlation ID to a downstream supplier. Both green.
Break the one production line they exist to protect:
- new SupplierClient(supplierTransport, { forwardCorrelation: true });
+ new SupplierClient(supplierTransport, { forwardCorrelation: false });
The service no longer forwards the header. It still compiles, which matters —
a mutation that doesn't compile tells you nothing. Then run the tests:
✓ hollow: correlation ID reaches the supplier
× fixed: correlation ID reaches the supplier
AssertionError: expected undefined to be 'abc'
41| expect(supplierSide.callCount).toBe(1);
42| expect(supplierSide.last()?.headers[HEADER]).toBe("abc");
Restore the line, confirm green, move on. Total cost: about thirty seconds.
The reasoning behind it: two passing tests tell you the tests and the code
agree. That's true when the code is right — and equally true when the test
can't tell the difference. From a green suite, those two situations look
identical.
Introducing a defect separates them. A test that genuinely checks the behaviour
has to notice, because its result depends on that behaviour. The hollow one
passed with forwarding switched on and with it switched off, which means its
result never depended on forwarding at all.
Six ways a passing test can't fail
I've rebuilt each of these on a small fictional service in TypeScript so you
can run them (link at the end). Each has a hollow version that passes and
a fixed version that also passes. The difference only shows under mutation.
None of these come from carelessness. They're the kind of test a competent
developer writes and a reviewer approves.
Four of the six are preventable by a rule. Two are not, and those are the ones
that justify the check. I've marked each.
1. The fixture makes failure impossible
Preventable by rule 5.
The service should forward a caller's correlation ID to a downstream supplier.
The test records outbound requests and checks for the header:
const shared = new Recorder();
const service = buildService(new Transport(shared));
const caller = new Transport(shared, (req) => service.handle(req)); // same recorder
caller.send({ path: "/orders", headers: { [HEADER]: "abc" }, body: ORDER });
expect(shared.requests.some((r) => r.headers[HEADER] === "abc")).toBe(true);
The test's own client and the code under test share one recorder. The inbound
request already carries the header, so the recorder always sees it — whether
or not the service forwarded anything. Break the forwarding and the test stays
green.
The assertion is fine. The fixture defeats it. The fix is separate recorders
and an assertion on the side that matters:
expect(supplierSide.callCount).toBe(1);
expect(supplierSide.last()?.headers[HEADER]).toBe("abc");
This is the one I'd least expect to find by reading. The fixture looks
correct; only the mutation shows it isn't.
2. The test builds its own object graph
Preventable by rule 4.
The test constructs the objects by hand, with the right settings, instead of
using the factory production uses. When the factory stops passing the setting,
the test doesn't notice — it never calls the factory.
Convincing in review: real code, real requests, real assertions. Just not the
code that ships.
3. "Nothing bad was sent" when nothing was sent
Preventable by rule 2.
expect(rec.requests.some((r) => HEADER in r.headers)).toBe(false);
some() over an empty array is false, so if the code path never runs, this
passes. "Nothing wrong happened" and "nothing happened" look identical. One
line fixes it: prove the calls were made first.
4. "Not empty", about something that's never empty
Preventable by rule 3.
An audit trail always contains the inbound entry, so asserting it isn't empty
can't tell you whether the reservation was recorded. Assert the specific entry.
5. The assertion that looks redundant
Not preventable by a rule. Nothing in the test looks wrong; the argument for
removing the assertion is reasonable until a mutation answers it.
A log's sequence number must advance only after a write succeeds. After a
failed write, the test checks the store is empty — which it always is, because
the write is all-or-nothing.
There's a reasonable argument for dropping the second assertion, that the
counter hasn't moved: the store is empty, so what could be wrong? The mutation
answers it. Advance the counter before the write, and the store is still
empty, but:
expected 2 to deeply equal +0
The counter has moved past entries that were never written. Only the
"redundant" assertion sees it.
6. The downstream system hides the behaviour
Not preventable by a rule. The test is correct; the behaviour is invisible
from where it is looking, and you only find that out by breaking the code.
Only the first request of an order should carry a parent ID. The integration
test checks what the downstream system recorded — but that system keeps the
value from the first call and ignores it afterwards. Sending it every time
changes nothing observable, and every integration test passes either way.
The integration test isn't wrong. It's looking where the behaviour is
invisible. A unit test on the outbound requests can see it.
What to do when a mutation survives
The first instinct is to strengthen the assertion. That's usually the wrong
move, and it cost me two rewrites before I stopped reaching for it. A surviving
mutation means the test's result didn't depend on the behaviour — so the
question is why not, and the answer is often somewhere other than the
assertion.
Four things to check, in order:
- Did the code path run at all? Add a call count, or print it. If the answer is no, the test was passing vacuously and the assertion was never reached.
- Could the expected value have arrived by another route? A shared fixture, a global, a fallback default. If the thing you're asserting on can be supplied by something other than the code under test, the assertion has nothing to discriminate.
- Is the assertion pointed at something that's always true? A collection that's never empty, a status that's set elsewhere, a field with a default.
- Is this the object production builds? If the test assembled its own, the mutation changed something the test never touches.
Then fix the cause, not the symptom. In the example above, the fix was
separating the recorders so the assertion had two distinguishable sides — not
a sharper assertion on a fixture that couldn't tell them apart.
And re-run the mutation afterwards. A rewrite isn't finished until it goes red;
more than once I found that my "fixed" version still survived, for a second
reason I hadn't spotted.
That happened to the sample project for this post, too. One of the "fixed"
tests turned out to be partly hollow: the assertion was right, but the fixture
only exercised a single order line, so a bug that dropped every line after the
first went unnoticed. It was caught by running a mutation against examples
written specifically to demonstrate this failure mode, by someone actively
looking for it. Knowing the shape isn't the same as proving the test can
fail.
What counts as a failing test
One detail worth making explicit, because it's easy to get wrong in both
directions.
A mutation must be a valid wrong implementation: it compiles, it runs,
it's just incorrect. If you change a function name to one that doesn't exist,
every test touching that code fails — the hollow ones included. The build
going red proves the line runs, not that any assertion checks it.
The same applies to red phases generally. A test that fails because the
function doesn't exist yet is a weaker signal than one that fails on an
assertion. Sometimes that's unavoidable early in a cycle; it's worth noting
when it happens and proving the test properly once the code is there.
In the sample project, the runner enforces both rules: a mutation has to
type-check, and only an assertion failure counts as a kill.
The same question, one step earlier
Most agents can work plan-first: the agent produces a plan and nothing changes
in the code until the plan is agreed. It's worth doing, and worth structuring
the plan around Red / Green / Refactor per behaviour, with every test named and
its assertions stated. That forces a commitment to what red looks like before
anything is written, which is most of the value of TDD arriving before the
first line of code.
It also moves the same question one stage earlier. A plan is an artifact too,
and it comes back with recognisable shapes: a decision left open with
"either… or", a placeholder where a test should be named, a checklist item
ticked before any work exists — sometimes against a project-specific rule the
agent defined for itself rather than looking up.
The useful move is the same one: verify the artifact, not the report of it.
MODE: Plan. VERIFY ONLY. Report, do not fix.
Read the plan file. Do not rely on your memory of the edits.
For each check, report PASS or FAIL and quote the exact lines
with line numbers as evidence.
Seventeen checks on one plan; eight failed on the first pass. Every fix was a
one-line edit once identified.
Every stage of this work produces an artifact the agent will also report on —
a test suite, a red commit, a plan, a checklist. Both checks in this post come
from the same instinct: read the artifact against something that could have
come out differently.
Patterns worth watching for
These come up often enough to be worth naming, and they're much easier to
catch when you're expecting them:
- A finding gets softened on repeat. A gap described as "not a blocker", or an assertion described as redundant. Useful response: settle it with a mutation rather than a discussion.
- A constraint gets read more strictly than it was written. "No version bumps" treated as "no new dependencies", with a design built around the stricter reading. Worth restating constraints precisely.
- A result gets reported from memory. "All tests pass" when they weren't run after the last edit. Asking for the raw output rather than a summary removes the ambiguity entirely.
- A fix creates a gap elsewhere. One change replaced a call that had been quietly supplying a default. Worth asking, on any replacement, what the old one provided.
- Decisions don't survive a context reset. A settled choice gets re-argued in a new session from the current code, without the history of why. Handoff notes that restate decisions fix this.
When to use it, and when not to
It's worth being clear about what this does and doesn't buy you.
A mutation check proves a test can fail for one change. It doesn't prove the
suite is complete, and it won't surface problems a fix causes elsewhere — a
change that's correct in isolation and wrong for the code around it still needs
a reviewer.
It also isn't worth doing everywhere. On small, visible work — a new field, a
validation rule, a copy change — reading the diff tells you everything the
mutation would. Reach for it where the code has indirection: injected
dependencies, factories, downstream systems, async paths, anything you can't
verify by eye. Those are the same places a hollow test is invisible in review,
which is not a coincidence.
If you take one thing from this, take the rules — they cost nothing once
written, and they prevent most of the shapes above. The check is for the
residue: the test that looks right, that a rule would not have caught, and that
few would question in review.
The loop gives you sequence: test first, one behaviour, red before green. The
rules deal with the shapes that recur. And then, occasionally, one more
question: could this have failed?
Sample project with all six test shapes and a runner that mutates one line at
a time: https://github.com/thasnim-fluxone/tests-that-cannot-fail.
npm run mutate shows every hollow test surviving and every fixed one caught.
Top comments (1)
Writing the rules before the agent fills them in is the right order. For signup routes I'd add cases a shape check passes: a well-formed email on a domain with no MX, and a national phone without a country, so the agent cannot delete those rules on a rewrite.