In my last post I told you about a test that passed and asserted nothing. It clicked a button, waited 3 seconds, and checked that the page still existed. Green tick, zero value.
The AI did not do anything strange there. It did what coding agents do by default: it stopped when the work looked finished. Anthropic's Claude Code docs say it plainly. If the agent has no check it can run, "looks done" is the only signal it has, and you become the verification loop.
So I stopped asking how to prompt better and started asking a different question: what has to be true before the agent is allowed to say done?
This post is the setup I built around that question. It is not a tool review. Almost all of it is plain files in your repo and one shell command, so it does not depend on which agent you use.
TL;DR
- One rule: nothing is done until a check that can fail has passed.
-
Six pieces: a short
AGENTS.md, oneverifycommand, a plan before non-trivial code, tests that go red first, small sessions, and a review gate. - One habit: every repeated mistake becomes a rule, a test or a hook.
- What I left out matters as much as what I kept. There are two studies on that below, and they disagree.
ticket -> PLAN.md -> failing test -> fix -> verify -> fresh-context review -> my review -> CI
The one rule
Nothing is done until a check that can fail has passed.
Not "the agent says it is done". Not "the diff looks reasonable". A command that returns pass or fail, and that I have personally seen fail.
Anthropic draws a useful line in Building effective agents: workflows run LLMs and tools through predefined code paths, while agents direct their own process. My setup is a bet on that line. Let the agent roam where the work is open-ended, like reading code and proposing a fix. Put everything predictable in code: the verify script, the hook, CI.
1. A short AGENTS.md that only holds what the agent cannot guess
AGENTS.md is a README for agents: one predictable file at the repo root that many coding agents read. Claude Code reads CLAUDE.md instead, and its memory docs suggest a CLAUDE.md that imports the shared file:
@AGENTS.md
## Claude Code
- Use plan mode for anything that touches more than two files.
Here is the shape of mine. Swap in your stack.
# AGENTS.md
## Commands (run from repo root)
- Install: `npm ci`
- Verify everything: `npm run verify`
- One test: `npx playwright test tests/cart.spec.ts -g "coupon"`
## Definition of done
1. `npm run verify` exits 0. Paste the last 20 lines of its output in your reply.
2. New behavior has a test that failed before your change.
3. The diff only touches files the task needs. No drive-by refactors.
## Things the code does not tell you
- Locators: `getByRole` or `getByLabel`. No CSS selectors, no `waitForTimeout`.
- Test data comes from `tests/fixtures/`.
## Ask before editing
- `playwright.config.ts`, `.github/workflows/`, lockfiles
It is about 20 lines, on purpose. There are two studies on context files and they do not agree. A team at ETH Zurich (arXiv 2602.11988) found that context files did not generally improve task success and raised inference cost by over 20% on average. The agents followed the instructions fine. Repo overviews just did not help. A smaller study (arXiv 2601.20404, 10 repos and 124 pull requests) found roughly 29% lower median runtime and 17% fewer output tokens with an AGENTS.md.
Different measures, different samples. I read them together like this: the file earns its place for what the agent cannot infer, like commands, quirks and your definition of done, and it costs you for everything else. My test for each line comes from the Claude Code docs: would removing it cause a mistake? If not, it goes. GitHub's look at 2,500+ agents.md files points the same way: one real code example beats three paragraphs of description.
2. One verify command, and a hook that enforces it
The most important line in that file is npm run verify. One command, one exit code. Lint, types, and a small smoke suite, fast enough that nobody is tempted to skip it.
{
"scripts": {
"lint": "eslint .",
"typecheck": "tsc --noEmit",
"test:smoke": "playwright test --grep @smoke",
"verify": "npm run lint && npm run typecheck && npm run test:smoke"
}
}
If you need a starting point for the smoke part, here is how I structure smoke and regression suites (Flutter, but the idea carries over). The CI side is in my test automation notes.
A line in AGENTS.md is a request. A hook is a guarantee. In Claude Code, a Stop hook runs when the agent tries to finish, and exiting with code 2 sends your message back and keeps it working. Add this to .claude/settings.json:
{
"hooks": {
"Stop": [
{ "hooks": [ { "type": "command", "command": "bash .claude/hooks/stop-gate.sh" } ] }
]
}
}
And this is .claude/hooks/stop-gate.sh:
#!/usr/bin/env bash
input="$(cat)"
# Already continuing because of this hook? Let it stop. No loops.
echo "$input" | grep -Eq '"stop_hook_active"[[:space:]]*:[[:space:]]*true' && exit 0
# Nothing changed, nothing to verify.
[ -z "$(git status --porcelain)" ] && exit 0
if ! output="$(npm run verify --silent 2>&1)"; then
echo "verify failed. Fix this before you finish:" >&2
echo "$output" | tail -n 40 >&2
exit 2
fi
Not on Claude Code? CI plays the same role, just later.
Then sabotage the gate itself. Break one assertion on purpose and confirm verify goes red. made this point in the comments on my last post: the checker can be the empty assertion. He ran a writing linter for three weeks before discovering two of its limits did nothing. A gate that cannot fail is worse than no gate, because it makes you feel safe.
3. Plan before code, unless the diff fits in one sentence
The Claude Code docs recommend four phases: explore, plan, implement, commit. They also say the part most guides skip: if you can describe the diff in one sentence, skip the plan. Planning is overhead for a typo and a bargain for a change across five files.
For anything bigger, the agent writes a plan and I read it before any code exists:
Read the code around the cart total and the coupon logic.
Do not change anything yet. Write PLAN.md with:
1. What you found: files, functions, how the total is calculated
2. Your approach, and one alternative you rejected
3. Files you will touch, and files you will not
4. How we will verify it. Name the failing test you will write first.
Then stop.
I read the plan like a pull request. The first thing I look for is files I did not expect. That is cheap to fix here and expensive to fix in a 400 line diff. Then I add the edge cases myself, the ones nobody asks for, like the payment webhook arriving twice. The agent will not volunteer those, and that is where most of my real bugs came from. For a deeper version of plan-first.
4. Red first, then sabotage
This piece came straight from the comments on my last post, so credit where it is due.
Red first. rule: a generated test must be red against the unfixed code before anyone reads the fix. A test that cannot go red first is probably the "clicked a button, checked the page exists" kind. So the agent works in two steps:
Bug: the cart total does not update when the last item is removed while a coupon is applied.
Step 1: write ONE Playwright test that reproduces this. Do not touch application code.
Step 2: run it and show me the failure output. Then stop.
I run it myself and read the failure. It has to fail for the right reason. A timeout or a typo is not a reproduction. Only then:
Now fix the bug so that test passes. Run `npm run verify` and paste the end of the output.
Sabotage. version: after the fix, remove the code under test (the handler, the flag, the stub), run again, and confirm the test goes red. If it stays green, it checks that the page exists, not that the feature works. Two minutes, no new tools. added a cheap pre-filter: look through CI history for tests that have never failed. A test that has passed every run since the day it was added either covers something nobody touches or asserts nothing.
What about Playwright's own agents? Playwright ships a planner, generator and healer (npx playwright init-agents --loop=vscode). The planner writes a Markdown test plan, the generator turns it into tests, and the healer repairs failing ones. Its documented outcome is a passing test, or a skipped one if it believes the app is broken. Last time I called self-healing locators mostly marketing, and I still think the run-time kind is. The difference is where it runs. A healer at dev time hands me a diff to review. A locator that silently repairs itself in CI stops telling me the UI changed. described that failure: a fallback hopped to a different element with the same label, and the suite stayed green across two deploys while form submissions were broken.
5. Small sessions, fresh context, written handoff
Most of the Claude Code best practices come from one constraint: the context window fills up fast, and performance drops as it fills. Anthropic's context engineering post calls it a finite attention budget. So I spend it carefully.
-
One task per session. Unrelated task? Fresh session (
/clearin Claude Code). - Two failed corrections means restart. By then the context is full of failed attempts. A clean session with a better prompt usually beats a long one full of corrections. That one is straight from the docs.
- Investigations go to a subagent or a separate session, so reading fifty files does not eat the room I need for the fix. When work runs longer than one session, I use a handoff file. Anthropic's post on long-running agents rests on the same idea: every new session starts with no memory, so a progress file and the git history carry the state. Mine is four lines:
Before you stop, append to PROGRESS.md under today's date:
- Done: one line each
- Next: the single next step
- Gotchas: anything that cost you time
Then commit with a message that says what changed and why.
6. A review gate that is not just me
Three layers, in this order.
A fresh-context reviewer. The agent that wrote the code is the worst reviewer of it. A second session or subagent sees only the diff and the plan:
Review this diff against PLAN.md. You have not seen the work behind it.
Report only: missing requirements, missing tests for the listed edge cases,
and changes outside the plan. No style opinions. If nothing qualifies, say "no gaps".
The last line matters. A reviewer asked to find problems will usually find some even when the work is sound, and chasing every finding leads to over-engineering. The Claude Code docs warn about exactly this.
Me, reading the diff. For test code I still do what I said last time: I own the assertions. A generated assertion checks that something exists. A written one checks that something is correct.
CI, running the same npm run verify. Same command on my laptop, in the hook and in the pipeline. If it passes in one place and fails in another, that is a bug in the setup and I fix that first.
The habit that keeps it short
Every time the agent makes the same mistake twice, I ask where the fix belongs:
- A line in AGENTS.md. Cheapest, and the weakest, because it is advisory.
- A lint rule or a test. It fails loudly and nobody has to remember it.
- A hook. For things that must happen every time, no exceptions. I prefer 3 over 2 over 1. And I prune: if the agent already does something right without the line, the line goes.
What I left out on purpose
- A tour of the repo. The ETH Zurich paper found repo overviews did not help agents. The agent can read the tree.
- Tools I cannot name a task for. Tool definitions and MCP servers all live in the agent's context. In my last post I said driving the browser through MCP was slower than writing the test myself (and that I might be holding it wrong). fix is the one I would try: use the agent at dev time to generate the page map and selectors, then run plain deterministic Playwright in CI. went further and suggested having the AI write a deterministic test generator instead of the tests. Agent at dev time, boring code at run time.
-
Parallel agents on day one. One reliable agent beats three confused ones. When yours is reliable,
git worktreelets a second one work on its own branch without overwriting the first. - Trusting how fast it feels. More on that next. ## Does it work? I try not to trust my gut
Here is why I care about measuring. In 2025, METR ran a randomized trial and found experienced open-source developers took 19% longer with AI tools while believing they were faster (paper). In February 2026 METR said that finding is outdated: speedups now seem likely, but selection effects, with developers unwilling to work without AI, made its new data unreliable (METR's update). I read both as the same lesson. How fast it feels is not a measurement. Pick two or three numbers and check them again in a month. I would track the share of agent PRs with a green first CI run, review rounds per PR, and bugs found after merge.
Where it still breaks
- Happy path bias. The plan step works because I add the failure modes myself. The agent does not ask what happens when the webhook fires twice.
- Review load grows. More PRs, same number of reviewers. I have not solved this one.
- Mobile. On Flutter the agent still hands me web patterns wearing a Flutter costume. The structure of the suite stays my job. ## Now tell me what I am missing
Three things I want to hear from you in the comments. I will reply to every one.
One. What is the one line in your AGENTS.md (or CLAUDE.md, or rules file) you would never delete? And what is the line you deleted that made the agent better?
Two. When did you last sabotage your own verify step on purpose? If you have never seen it go red, how do you know it can?
Three. Where is your line for plan-first? Mine is "can I describe the diff in one sentence". I suspect it breaks somewhere and I want to know where.
Drop your take below. Especially if you disagree.
Previous in this series: I Let AI Write My Tests for 6 Months. Here Is What Actually Survived Production. I write about test automation and AI tooling at nileshblog.tech.
Sources and further reading
- Anthropic: Building effective agents
- Anthropic: Effective context engineering for AI agents
- Anthropic: Effective harnesses for long-running agents
- Claude Code docs: Best practices, Memory and CLAUDE.md, Hooks
- AGENTS.md and GitHub's lessons from 2,500+ agents.md files
- ETH Zurich: Evaluating AGENTS.md and Lulla et al.: Impact of AGENTS.md on agent efficiency
- METR: Changing our developer productivity experiment design
- Playwright: Test Agents
- Boris Tane: How I use Claude Code
- Git: git worktree
Top comments (0)