<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Thasnim</title>
    <description>The latest articles on DEV Community by Thasnim (@thasnimfluxone).</description>
    <link>https://dev.to/thasnimfluxone</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4152046%2F9422b51c-048a-4cf2-8160-2a37a3ee0af2.jpeg</url>
      <title>DEV Community: Thasnim</title>
      <link>https://dev.to/thasnimfluxone</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/thasnimfluxone"/>
    <language>en</language>
    <item>
      <title>TDD With Coding Agents: Write the Rules, Then Check They Held</title>
      <dc:creator>Thasnim</dc:creator>
      <pubDate>Wed, 30 Sep 2026 11:08:24 +0000</pubDate>
      <link>https://dev.to/thasnimfluxone/tdd-with-coding-agents-write-the-rules-then-check-they-held-29ak</link>
      <guid>https://dev.to/thasnimfluxone/tdd-with-coding-agents-write-the-rules-then-check-they-held-29ak</guid>
      <description>&lt;p&gt;Test-driven development and coding agents fit together unusually well. A test&lt;br&gt;
is a precise specification the agent can check its own work against. The&lt;br&gt;
red-green-refactor cycle keeps each change small enough to review. And the&lt;br&gt;
"run the tests, read the output, try again" rhythm is exactly what these tools&lt;br&gt;
are good at.&lt;/p&gt;

&lt;p&gt;The guidance on how to do it has matured quickly, and most of it is sound. But&lt;br&gt;
after working this way across different codebases and languages, with the agent&lt;br&gt;
writing both the tests and the code, I kept finding tests that passed whether&lt;br&gt;
the code worked or not.&lt;/p&gt;

&lt;p&gt;Most of those were preventable. A handful of rules, written once into the&lt;br&gt;
agent's instructions, stop the common shapes before they are written. The rest&lt;br&gt;
needed a check after green, because no rule anticipates every way a test can&lt;br&gt;
fail to discriminate.&lt;/p&gt;

&lt;p&gt;This post is both: the rules worth writing down, and the thirty-second check&lt;br&gt;
for what they miss.&lt;/p&gt;

&lt;p&gt;The work behind this post used IBM Bob as the coding agent. Nothing here is&lt;br&gt;
specific to it. Bob routes tasks across several frontier models rather than a&lt;br&gt;
single fixed one, so a single session may not even have used the same model&lt;br&gt;
throughout — and the same patterns turned up regardless of language, codebase&lt;br&gt;
or tool.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Everything in this post is reproducible: &lt;a href="https://github.com/thasnim-fluxone/tests-that-cannot-fail" rel="noopener noreferrer"&gt;sample project on GitHub&lt;/a&gt;&lt;br&gt;
— six test shapes and a runner that mutates one line at a time.&lt;/em&gt;&lt;/p&gt;


&lt;h2&gt;
  
  
  Where the practice stands
&lt;/h2&gt;

&lt;p&gt;It's worth tracing how the advice has developed, because each stage solved a&lt;br&gt;
real problem.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Keep the tests in human hands.&lt;/strong&gt; The early guidance was that if a model&lt;br&gt;
writes both the code and the tests, the same assumption ends up in both and&lt;br&gt;
they agree with each other. So humans wrote the tests and the AI wrote the&lt;br&gt;
code. That division worked, and it's still good advice when the spec matters&lt;br&gt;
more than the speed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Write the discipline down.&lt;/strong&gt; Agents can produce twenty tests in the time you&lt;br&gt;
write one, and most teams took that trade. The guidance shifted to enforcing&lt;br&gt;
the cycle through whatever instruction mechanism the tool offers — an&lt;br&gt;
&lt;code&gt;AGENTS.md&lt;/code&gt; in the project root now works across several of them, alongside&lt;br&gt;
tool-specific rules files like &lt;code&gt;.bob/rules&lt;/code&gt; for IBM Bob, &lt;code&gt;CLAUDE.md&lt;/code&gt; for Claude&lt;br&gt;
Code and &lt;code&gt;.cursor/rules&lt;/code&gt; for Cursor — carrying the project's conventions and&lt;br&gt;
the test-first rule, one behaviour per cycle, a human gate before&lt;br&gt;
implementing, and, in the stronger versions, committing the red test so the&lt;br&gt;
failing state lives in git history.&lt;/p&gt;

&lt;p&gt;Kent Beck, who originated TDD, has published a system prompt along these lines: &lt;em&gt;always&lt;br&gt;
follow the cycle, write the simplest failing test first, implement the minimum&lt;br&gt;
needed to pass.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Enforce it mechanically.&lt;/strong&gt; The newest approach doesn't rely on the agent&lt;br&gt;
following instructions at all. TDD Guard hooks every file write and blocks it&lt;br&gt;
unless the process was honoured: a failing test exists, you're on one test at&lt;br&gt;
a time, you're working outside-in. Under the covers it spins up a second model&lt;br&gt;
as a judge on each edit, because "did this follow TDD?" is a fuzzy question&lt;br&gt;
that is easier to ask a model than to encode in rules.&lt;/p&gt;

&lt;p&gt;Each stage is a genuine improvement, and the progression is the right one.&lt;br&gt;
Prompting alone tends to produce what one practitioner aptly called "test&lt;br&gt;
first, not test-driven" — all the tests, then all the code, in two large&lt;br&gt;
steps. Instruction files make the cycle stick more often. Hooks make it stick unless&lt;br&gt;
the judge misses something.&lt;/p&gt;

&lt;p&gt;Every tool I looked at governs &lt;strong&gt;sequence&lt;/strong&gt;: was the test written first, is it&lt;br&gt;
one behaviour, did red precede green. That's the hard part to enforce, and they&lt;br&gt;
enforce it well.&lt;/p&gt;

&lt;p&gt;The step I'm suggesting governs something different: &lt;strong&gt;whether the resulting&lt;br&gt;
test can fail at all.&lt;/strong&gt;&lt;/p&gt;


&lt;h2&gt;
  
  
  The rules: what prevents most of it
&lt;/h2&gt;

&lt;p&gt;These go in whatever instruction mechanism your tool offers, and they are the&lt;br&gt;
higher-value half of this post. Written once, they make most of the shapes in&lt;br&gt;
the next section much less likely, though no instruction is followed perfectly:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Only an assertion failure counts as red.&lt;/strong&gt; A compile error or a missing
symbol is not a failing test.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Every negative assertion first proves the code ran.&lt;/strong&gt; Assert the call
count, then the absence.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Assert the specific item, not that a collection is non-empty.&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tests go through production construction&lt;/strong&gt; — the real factory or
constructor, never a hand-built object graph.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The code under test and the test's own client never share an injected
dependency.&lt;/strong&gt; Separate instances, separate recorders.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Report from the artifact, not from memory.&lt;/strong&gt; Quote the changed lines or
paste the raw test output.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Plans make decisions.&lt;/strong&gt; No "either/or", no placeholders, no pre-ticked
checklists, and project terms quoted from their source.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;When a test's coverage is disputed, settle it with a mutation, not an
argument.&lt;/strong&gt; This is the one rule that needs the check, and it is the one
that caught what the others missed.&lt;/li&gt;
&lt;/ol&gt;


&lt;h2&gt;
  
  
  The check, for what rules don't catch
&lt;/h2&gt;

&lt;p&gt;Those rules prevent most of the shapes below. What follows is for the cases&lt;br&gt;
they cannot anticipate.&lt;/p&gt;

&lt;p&gt;After green, before moving on:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Break one production line on purpose. Say in advance which test should fail.&lt;br&gt;
Run it. Confirm it fails &lt;strong&gt;on an assertion&lt;/strong&gt;. Then restore the line.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That's mutation testing, done by hand, one line at a time, at review time&lt;br&gt;
rather than in CI. Note the direction: the code is correct, you introduce a&lt;br&gt;
deliberate defect, and the test failing is the good outcome — it's the test&lt;br&gt;
doing its job.&lt;/p&gt;

&lt;p&gt;A test that can't fail isn't a new problem, and it isn't specific to AI.&lt;br&gt;
Mutation testing has existed for decades precisely because of it, and most TDD&lt;br&gt;
guides mention it — usually a line in a metrics section beside coverage&lt;br&gt;
thresholds, pointing at Stryker or mutmut. That framing makes it sound like a&lt;br&gt;
quarterly exercise. Used per change, it's smaller and more immediate: the one&lt;br&gt;
question that separates a test from a decoration.&lt;/p&gt;

&lt;p&gt;What's different with agents is that TDD normally guards against this, and the&lt;br&gt;
guard gets weaker. The red step is meant to be the proof — you watched the test&lt;br&gt;
fail, so it can fail. That holds when a person writes one test and sees it go&lt;br&gt;
red for the reason they expected. It holds less well when an agent produces the&lt;br&gt;
test: the red often comes from the function not existing yet rather than from&lt;br&gt;
the assertion, and the volume means few of them get inspected closely.&lt;/p&gt;

&lt;p&gt;Better instructions narrow this a long way, and you should write them. But&lt;br&gt;
instructions produce better tests; they do not produce proof that a given test&lt;br&gt;
can fail. That distinction is the whole reason the check exists.&lt;/p&gt;

&lt;p&gt;So the mutation asks what the red phase no longer reliably answers. A red test&lt;br&gt;
proves it failed &lt;em&gt;before the code existed&lt;/em&gt; — when everything failed. This asks:&lt;br&gt;
now that the code exists, would this test notice if it broke?&lt;/p&gt;

&lt;p&gt;Most of the time the answer is yes, it takes thirty seconds, and you move on.&lt;br&gt;
Occasionally it isn't, and those are the cases worth writing about.&lt;/p&gt;
&lt;h3&gt;
  
  
  What that looks like
&lt;/h3&gt;

&lt;p&gt;Two agent-written tests for the same behaviour: the service forwards the&lt;br&gt;
caller's correlation ID to a downstream supplier. Both green.&lt;/p&gt;

&lt;p&gt;Break the one production line they exist to protect:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight diff"&gt;&lt;code&gt;&lt;span class="gd"&gt;- new SupplierClient(supplierTransport, { forwardCorrelation: true });
&lt;/span&gt;&lt;span class="gi"&gt;+ new SupplierClient(supplierTransport, { forwardCorrelation: false });
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The service no longer forwards the header. It still compiles, which matters —&lt;br&gt;
a mutation that doesn't compile tells you nothing. Then run the tests:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;✓ hollow: correlation ID reaches the supplier
× fixed:  correlation ID reaches the supplier

AssertionError: expected undefined to be 'abc'
  41|   expect(supplierSide.callCount).toBe(1);
  42|   expect(supplierSide.last()?.headers[HEADER]).toBe("abc");
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Restore the line, confirm green, move on. Total cost: about thirty seconds.&lt;/p&gt;

&lt;p&gt;The reasoning behind it: two passing tests tell you the tests and the code&lt;br&gt;
agree. That's true when the code is right — and equally true when the test&lt;br&gt;
can't tell the difference. From a green suite, those two situations look&lt;br&gt;
identical.&lt;/p&gt;

&lt;p&gt;Introducing a defect separates them. A test that genuinely checks the behaviour&lt;br&gt;
has to notice, because its result depends on that behaviour. The hollow one&lt;br&gt;
passed with forwarding switched on and with it switched off, which means its&lt;br&gt;
result never depended on forwarding at all.&lt;/p&gt;


&lt;h2&gt;
  
  
  Six ways a passing test can't fail
&lt;/h2&gt;

&lt;p&gt;I've rebuilt each of these on a small fictional service in TypeScript so you&lt;br&gt;
can run them (link at the end). Each has a &lt;strong&gt;hollow&lt;/strong&gt; version that passes and&lt;br&gt;
a &lt;strong&gt;fixed&lt;/strong&gt; version that also passes. The difference only shows under mutation.&lt;/p&gt;

&lt;p&gt;None of these come from carelessness. They're the kind of test a competent&lt;br&gt;
developer writes and a reviewer approves.&lt;/p&gt;

&lt;p&gt;Four of the six are preventable by a rule. Two are not, and those are the ones&lt;br&gt;
that justify the check. I've marked each.&lt;/p&gt;
&lt;h3&gt;
  
  
  1. The fixture makes failure impossible
&lt;/h3&gt;

&lt;p&gt;&lt;em&gt;Preventable by rule 5.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The service should forward a caller's correlation ID to a downstream supplier.&lt;br&gt;
The test records outbound requests and checks for the header:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;shared&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Recorder&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;service&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;buildService&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Transport&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;shared&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;caller&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Transport&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;shared&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;service&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;handle&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt; &lt;span class="c1"&gt;// same recorder&lt;/span&gt;

&lt;span class="nx"&gt;caller&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;send&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;path&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;/orders&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;HEADER&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;abc&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="na"&gt;body&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;ORDER&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="nf"&gt;expect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;shared&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;requests&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;some&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;r&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;HEADER&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;abc&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)).&lt;/span&gt;&lt;span class="nf"&gt;toBe&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The test's own client and the code under test share one recorder. The inbound&lt;br&gt;
request already carries the header, so the recorder always sees it — whether&lt;br&gt;
or not the service forwarded anything. Break the forwarding and the test stays&lt;br&gt;
green.&lt;/p&gt;

&lt;p&gt;The assertion is fine. The fixture defeats it. The fix is separate recorders&lt;br&gt;
and an assertion on the side that matters:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="nf"&gt;expect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;supplierSide&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;callCount&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;toBe&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="nf"&gt;expect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;supplierSide&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;last&lt;/span&gt;&lt;span class="p"&gt;()?.&lt;/span&gt;&lt;span class="nx"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;HEADER&lt;/span&gt;&lt;span class="p"&gt;]).&lt;/span&gt;&lt;span class="nf"&gt;toBe&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;abc&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is the one I'd least expect to find by reading. The fixture looks&lt;br&gt;
correct; only the mutation shows it isn't.&lt;/p&gt;
&lt;h3&gt;
  
  
  2. The test builds its own object graph
&lt;/h3&gt;

&lt;p&gt;&lt;em&gt;Preventable by rule 4.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The test constructs the objects by hand, with the right settings, instead of&lt;br&gt;
using the factory production uses. When the factory stops passing the setting,&lt;br&gt;
the test doesn't notice — it never calls the factory.&lt;/p&gt;

&lt;p&gt;Convincing in review: real code, real requests, real assertions. Just not the&lt;br&gt;
code that ships.&lt;/p&gt;
&lt;h3&gt;
  
  
  3. "Nothing bad was sent" when nothing was sent
&lt;/h3&gt;

&lt;p&gt;&lt;em&gt;Preventable by rule 2.&lt;/em&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="nf"&gt;expect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;rec&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;requests&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;some&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;r&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;HEADER&lt;/span&gt; &lt;span class="k"&gt;in&lt;/span&gt; &lt;span class="nx"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;)).&lt;/span&gt;&lt;span class="nf"&gt;toBe&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;some()&lt;/code&gt; over an empty array is &lt;code&gt;false&lt;/code&gt;, so if the code path never runs, this&lt;br&gt;
passes. "Nothing wrong happened" and "nothing happened" look identical. One&lt;br&gt;
line fixes it: prove the calls were made first.&lt;/p&gt;
&lt;h3&gt;
  
  
  4. "Not empty", about something that's never empty
&lt;/h3&gt;

&lt;p&gt;&lt;em&gt;Preventable by rule 3.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;An audit trail always contains the inbound entry, so asserting it isn't empty&lt;br&gt;
can't tell you whether the reservation was recorded. Assert the specific entry.&lt;/p&gt;
&lt;h3&gt;
  
  
  5. The assertion that looks redundant
&lt;/h3&gt;

&lt;p&gt;&lt;em&gt;Not preventable by a rule. Nothing in the test looks wrong; the argument for&lt;br&gt;
removing the assertion is reasonable until a mutation answers it.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;A log's sequence number must advance only after a write succeeds. After a&lt;br&gt;
failed write, the test checks the store is empty — which it always is, because&lt;br&gt;
the write is all-or-nothing.&lt;/p&gt;

&lt;p&gt;There's a reasonable argument for dropping the second assertion, that the&lt;br&gt;
counter hasn't moved: the store is empty, so what could be wrong? The mutation&lt;br&gt;
answers it. Advance the counter before the write, and the store is still&lt;br&gt;
empty, but:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;expected 2 to deeply equal +0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The counter has moved past entries that were never written. Only the&lt;br&gt;
"redundant" assertion sees it.&lt;/p&gt;
&lt;h3&gt;
  
  
  6. The downstream system hides the behaviour
&lt;/h3&gt;

&lt;p&gt;&lt;em&gt;Not preventable by a rule. The test is correct; the behaviour is invisible&lt;br&gt;
from where it is looking, and you only find that out by breaking the code.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Only the first request of an order should carry a parent ID. The integration&lt;br&gt;
test checks what the downstream system recorded — but that system keeps the&lt;br&gt;
value from the first call and ignores it afterwards. Sending it every time&lt;br&gt;
changes nothing observable, and every integration test passes either way.&lt;/p&gt;

&lt;p&gt;The integration test isn't wrong. It's looking where the behaviour is&lt;br&gt;
invisible. A unit test on the outbound requests can see it.&lt;/p&gt;


&lt;h2&gt;
  
  
  What to do when a mutation survives
&lt;/h2&gt;

&lt;p&gt;The first instinct is to strengthen the assertion. That's usually the wrong&lt;br&gt;
move, and it cost me two rewrites before I stopped reaching for it. A surviving&lt;br&gt;
mutation means the test's result didn't depend on the behaviour — so the&lt;br&gt;
question is &lt;em&gt;why not&lt;/em&gt;, and the answer is often somewhere other than the&lt;br&gt;
assertion.&lt;/p&gt;

&lt;p&gt;Four things to check, in order:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Did the code path run at all?&lt;/strong&gt; Add a call count, or print it. If the
answer is no, the test was passing vacuously and the assertion was never
reached.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Could the expected value have arrived by another route?&lt;/strong&gt; A shared
fixture, a global, a fallback default. If the thing you're asserting on can
be supplied by something other than the code under test, the assertion has
nothing to discriminate.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Is the assertion pointed at something that's always true?&lt;/strong&gt; A collection
that's never empty, a status that's set elsewhere, a field with a default.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Is this the object production builds?&lt;/strong&gt; If the test assembled its own,
the mutation changed something the test never touches.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Then fix the &lt;em&gt;cause&lt;/em&gt;, not the symptom. In the example above, the fix was&lt;br&gt;
separating the recorders so the assertion had two distinguishable sides — not&lt;br&gt;
a sharper assertion on a fixture that couldn't tell them apart.&lt;/p&gt;

&lt;p&gt;And re-run the mutation afterwards. A rewrite isn't finished until it goes red;&lt;br&gt;
more than once I found that my "fixed" version still survived, for a second&lt;br&gt;
reason I hadn't spotted.&lt;/p&gt;

&lt;p&gt;That happened to the sample project for this post, too. One of the "fixed"&lt;br&gt;
tests turned out to be partly hollow: the assertion was right, but the fixture&lt;br&gt;
only exercised a single order line, so a bug that dropped every line after the&lt;br&gt;
first went unnoticed. It was caught by running a mutation against examples&lt;br&gt;
written specifically to demonstrate this failure mode, by someone actively&lt;br&gt;
looking for it. Knowing the shape isn't the same as proving the test can&lt;br&gt;
fail.&lt;/p&gt;


&lt;h2&gt;
  
  
  What counts as a failing test
&lt;/h2&gt;

&lt;p&gt;One detail worth making explicit, because it's easy to get wrong in both&lt;br&gt;
directions.&lt;/p&gt;

&lt;p&gt;A mutation must be a &lt;strong&gt;valid wrong implementation&lt;/strong&gt;: it compiles, it runs,&lt;br&gt;
it's just incorrect. If you change a function name to one that doesn't exist,&lt;br&gt;
every test touching that code fails — the hollow ones included. The build&lt;br&gt;
going red proves the line runs, not that any assertion checks it.&lt;/p&gt;

&lt;p&gt;The same applies to red phases generally. A test that fails because the&lt;br&gt;
function doesn't exist yet is a weaker signal than one that fails on an&lt;br&gt;
assertion. Sometimes that's unavoidable early in a cycle; it's worth noting&lt;br&gt;
when it happens and proving the test properly once the code is there.&lt;/p&gt;

&lt;p&gt;In the sample project, the runner enforces both rules: a mutation has to&lt;br&gt;
type-check, and only an assertion failure counts as a kill.&lt;/p&gt;


&lt;h2&gt;
  
  
  The same question, one step earlier
&lt;/h2&gt;

&lt;p&gt;Most agents can work plan-first: the agent produces a plan and nothing changes&lt;br&gt;
in the code until the plan is agreed. It's worth doing, and worth structuring&lt;br&gt;
the plan around Red / Green / Refactor per behaviour, with every test named and&lt;br&gt;
its assertions stated. That forces a commitment to what red looks like before&lt;br&gt;
anything is written, which is most of the value of TDD arriving before the&lt;br&gt;
first line of code.&lt;/p&gt;

&lt;p&gt;It also moves the same question one stage earlier. A plan is an artifact too,&lt;br&gt;
and it comes back with recognisable shapes: a decision left open with&lt;br&gt;
"either… or", a placeholder where a test should be named, a checklist item&lt;br&gt;
ticked before any work exists — sometimes against a project-specific rule the&lt;br&gt;
agent defined for itself rather than looking up.&lt;/p&gt;

&lt;p&gt;The useful move is the same one: verify the artifact, not the report of it.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;MODE: Plan. VERIFY ONLY. Report, do not fix.
Read the plan file. Do not rely on your memory of the edits.
For each check, report PASS or FAIL and quote the exact lines
with line numbers as evidence.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Seventeen checks on one plan; eight failed on the first pass. Every fix was a&lt;br&gt;
one-line edit once identified.&lt;/p&gt;

&lt;p&gt;Every stage of this work produces an artifact the agent will also report on —&lt;br&gt;
a test suite, a red commit, a plan, a checklist. Both checks in this post come&lt;br&gt;
from the same instinct: read the artifact against something that could have&lt;br&gt;
come out differently.&lt;/p&gt;




&lt;h2&gt;
  
  
  Patterns worth watching for
&lt;/h2&gt;

&lt;p&gt;These come up often enough to be worth naming, and they're much easier to&lt;br&gt;
catch when you're expecting them:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A finding gets softened on repeat.&lt;/strong&gt; A gap described as "not a blocker",
or an assertion described as redundant. Useful response: settle it with a
mutation rather than a discussion.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A constraint gets read more strictly than it was written.&lt;/strong&gt; "No version
bumps" treated as "no new dependencies", with a design built around the
stricter reading. Worth restating constraints precisely.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A result gets reported from memory.&lt;/strong&gt; "All tests pass" when they weren't
run after the last edit. Asking for the raw output rather than a summary
removes the ambiguity entirely.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A fix creates a gap elsewhere.&lt;/strong&gt; One change replaced a call that had been
quietly supplying a default. Worth asking, on any replacement, what the old
one provided.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Decisions don't survive a context reset.&lt;/strong&gt; A settled choice gets re-argued
in a new session from the current code, without the history of why. Handoff
notes that restate decisions fix this.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  When to use it, and when not to
&lt;/h2&gt;

&lt;p&gt;It's worth being clear about what this does and doesn't buy you.&lt;/p&gt;

&lt;p&gt;A mutation check proves a test can fail for &lt;em&gt;one&lt;/em&gt; change. It doesn't prove the&lt;br&gt;
suite is complete, and it won't surface problems a fix causes elsewhere — a&lt;br&gt;
change that's correct in isolation and wrong for the code around it still needs&lt;br&gt;
a reviewer.&lt;/p&gt;

&lt;p&gt;It also isn't worth doing everywhere. On small, visible work — a new field, a&lt;br&gt;
validation rule, a copy change — reading the diff tells you everything the&lt;br&gt;
mutation would. Reach for it where the code has indirection: injected&lt;br&gt;
dependencies, factories, downstream systems, async paths, anything you can't&lt;br&gt;
verify by eye. Those are the same places a hollow test is invisible in review,&lt;br&gt;
which is not a coincidence.&lt;/p&gt;

&lt;p&gt;If you take one thing from this, take the rules — they cost nothing once&lt;br&gt;
written, and they prevent most of the shapes above. The check is for the&lt;br&gt;
residue: the test that looks right, that a rule would not have caught, and that&lt;br&gt;
few would question in review.&lt;/p&gt;

&lt;p&gt;The loop gives you sequence: test first, one behaviour, red before green. The&lt;br&gt;
rules deal with the shapes that recur. And then, occasionally, one more&lt;br&gt;
question: &lt;em&gt;could this have failed?&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Sample project with all six test shapes and a runner that mutates one line at&lt;br&gt;
a time: &lt;a href="https://github.com/thasnim-fluxone/tests-that-cannot-fail" rel="noopener noreferrer"&gt;https://github.com/thasnim-fluxone/tests-that-cannot-fail&lt;/a&gt;.&lt;br&gt;
&lt;code&gt;npm run mutate&lt;/code&gt; shows every hollow test surviving and every fixed one caught.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>testing</category>
      <category>ai</category>
      <category>tdd</category>
      <category>typescript</category>
    </item>
  </channel>
</rss>
