DEV Community

aicoding-guide
aicoding-guide

Posted on Originally published at aicoding-guide.com

Make a coding agent write the failing test before the implementation

Originally published at https://aicoding-guide.com.

Ask a coding agent to "implement this feature" and the implementation arrives fast, but the verification often ends at "it should work." The most effective fix is to make it write a failing test first, confirm the failure, and only then implement.

An agent with a concrete goal — "make this test pass" — can iterate on its own. This article covers the prompts that work across all three tools, and the tool-specific ways to enforce the rule rather than request it.

Key point
What you will learn

  • The four steps of the workflow and the prompt for each
  • What to put in your context file (CLAUDE.md / AGENTS.md / GEMINI.md)
  • How to stop the agent from editing the tests, and what to do when it goes wrong

The four steps

Step What happens What to tell the agent
1. Write the test A failing test derived from the spec Say "don't write the implementation yet"
2. Confirm the failure Run it and see it fail Have it confirm the failure is "not implemented"
3. Implement Change the implementation until the test passes Say "don't change the test"
4. Tidy up Remove duplication, improve names, add cases Have it confirm tests still pass afterwards

The important part is that steps 1 and 3 are separate instructions. Roll them into one ("write tests and implement it") and the agent tends to write the implementation first and back-fill a test that fits it.

The prompts

Step 1: tests only

I want a calculateTotal function in `src/lib/pricing.ts` that returns the discounted
total from a quantity and a unit price.
Spec:
- 5% off at 10 or more, 10% off at 100 or more
- Throw if the quantity is 0 or less
- Round fractions down

First, write only the tests for this spec in `src/lib/pricing.test.ts`.
Do not write the implementation yet. Once the tests are written, run
`pnpm test pricing`, confirm they all fail because the function does not exist,
and report back.
Enter fullscreen mode Exit fullscreen mode

Steps 2–3: implement

Implement calculateTotal so every test in pricing.test.ts passes.
Do not modify the test file. If you think a test is wrong, report why
without changing it.
Run `pnpm test pricing` afterwards and paste the result.
Enter fullscreen mode Exit fullscreen mode

Step 4: tidy up

Review the implementation and the tests, and clean up duplication or unclear names.
Add boundary cases if any are missing (9, 10, 99, 100).
Finally run `pnpm test pricing` and `pnpm lint` and report the results.
Enter fullscreen mode Exit fullscreen mode

Glossary
Boundary cases: tests at the points where the spec changes behavior (10, 100) and just either side of them (9, 11). Agents tend to test "representative" values only, so asking for boundaries explicitly pays off.

Put it in the context file

If repeating the instructions gets tiring, write the workflow into your context file. The same text works for all three tools:

## How to implement (test first)
1. For a new feature or a fix, write a failing test first and run it to confirm the failure
2. Implement until the test passes. Do not change the test file while implementing
3. If you judge a test to be wrong, report why without changing it and wait for instructions
4. Include the test command you ran and its result (passed / failed counts) in your report
5. If there is no test setup, propose setting one up first
Enter fullscreen mode Exit fullscreen mode

For keeping one copy of that text across the three files, see CLAUDE.md, AGENTS.md and GEMINI.md differ only in how they load.

Claude Code: enforce it with a hook

In Claude Code you can make this a mechanism rather than a request. The documentation is explicit that a context file is "context, not enforced configuration" and points at a PreToolUse hook to block an action regardless of what Claude decides.

A hook that blocks edits to test files during the implementation phase:

#!/usr/bin/env bash
# .claude/hooks/protect-tests.sh
# Only blocks test-file edits when PROTECT_TESTS=1
[ "${PROTECT_TESTS:-0}" = "1" ] || exit 0
file=$(jq -r '.tool_input.file_path // empty')
case "$file" in
  *.test.*|*.spec.*|*/__tests__/*)
    echo "Test files can't be changed during the implementation phase: $file. If a test is wrong, report why." >&2
    exit 2 ;;
esac
exit 0
Enter fullscreen mode Exit fullscreen mode
{
  "hooks": {
    "PreToolUse": [
      {
        "matcher": "Edit|Write|NotebookEdit",
        "hooks": [{ "type": "command", "command": "bash .claude/hooks/protect-tests.sh" }]
      }
    ]
  }
}
Enter fullscreen mode Exit fullscreen mode

Exit code 2 on PreToolUse blocks the tool call, so starting with PROTECT_TESTS=1 claude makes tampering with the tests physically impossible rather than merely discouraged. For hook basics, see Run lint and format on save with Claude Code hooks.

Fix the procedure in a skill: writing the four steps into a /tdd <spec> slash command keeps every run on the same track (Create your own slash commands). Delegating the test run to a subagent keeps long test output out of your main context.

Codex and Gemini CLI

Neither has an equivalent of hooks, so you run this on instructions plus inspection.

  • Codex: keep the implementation phase on --sandbox workspace-write, and check with /diff that the test files are unchanged before you approve. Put the workflow in AGENTS.md.
  • Gemini CLI: turn checkpointing on, and if a test file was modified, roll back with /restore. Put the workflow in GEMINI.md.

When it goes wrong

Symptom What to do
It writes the implementation during the test step Repeat "don't write the implementation yet" at the end of the prompt too. In Claude Code, use plan mode to get a plan only
It loosens the test when it won't pass State "don't change the test" and block it with a hook
The tests are all mocks and prove nothing Say "mock external dependencies if you must, but never mock the logic of the function under test"
The test run takes too long Put a narrowing command in the prompt (pnpm test pricing)

Summary

  • Make "write the test" and "implement" separate instructions, with a confirmed failure in between
  • State "don't change the test file" and "include test results in your report" every time, or put them in your context file
  • Claude Code can block test-file edits with a PreToolUse hook and fix the procedure in a skill
  • With Codex and Gemini CLI, use /diff or checkpointing to check the tests weren't tampered with

Top comments (0)