DEV Community

Jeff
Jeff

Posted on Originally published at powerduck.com

Every Agent Task Should End With a Command You Can Run

The pattern

You come back to an agent session. The summary says:

Done. Refactored the auth middleware, added rate limiting, updated the tests, all green.

You trust it. You merge. Three days later someone hits a 500 on a route that didn't change.

The pattern isn't that the agent lied. It's that the agent's definition of "done" is the thing that makes the loop end, and your definition of "done" — the thing that proves the system still behaves — is a different check that nobody ran.

What 'all green' actually means

When an agent says tests pass, it usually means one of these:

  • It ran the test suite and the exit code was 0.
  • It ran a test command and saw output it interpreted as success.
  • It wrote a test, ran it, saw it pass, and never re-ran the ones it didn't write.
  • It assumed the existing suite covers the path it just changed, because that's cheaper than checking.

None of these is the same as "the behavior the user relied on still works." They're all checks the agent designed, ran, and observed itself — which is the same structural problem as an AI review bot reviewing its own diff.

The cheap fix isn't a fancier evaluator. It's a discipline: every agent task ends with a one-line command you can run yourself.

The rule

At the end of any non-trivial agent task, ask for one thing:

Before you close this out, give me a single command I can paste into a fresh terminal that proves the thing you changed still works. No commentary, just the command and what output I should expect.

That's it. The agent will produce something like:

npm test -- --grep "auth middleware"
# expect: 12 passing, 0 failing, ~3s
Enter fullscreen mode Exit fullscreen mode

or

curl -s http://localhost:3000/health | jq .status
# expect: "ok"
Enter fullscreen mode Exit fullscreen mode

The interesting failure mode isn't the agent refusing. It's the agent discovering it can't produce one. When you ask for the command and it says "um, the test suite should cover it," you've just learned the task wasn't actually closed.

Why this surfaces the real gap

A command has to be runnable by someone other than the agent. That single constraint forces three things the summary doesn't:

  1. The command has to exist. The agent has to know the test command, the dev server URL, the seed script. If it doesn't, that's a real gap in your onboarding docs, not a model weakness.
  2. The expected output has to be specific. "Tests pass" is not specific. "12 passing, 0 failing, ~3s" is. Vague expected output is a tell that the agent hasn't actually looked.
  3. You can run it in a minute. The cost of skepticism drops from "open the agent's diff and read 200 lines" to "paste one line, wait 3 seconds." You start actually doing the check.

When the command is a contract check

For API work, the runnable command is usually not a unit test. It's something that hits the running service and compares the response against a contract the agent didn't write in this session.

That's the highest-leverage version of the rule: the command checks against an artifact that predates the task. An agent can rationalize a test it wrote five minutes ago. It can't rationalize a spec that's been in the repo for three months, when the live response returns a field the spec doesn't list.

We built Powerduck around this loop: the OpenAPI file lives locally, and the agent's "done" command is a contract check against the running endpoint — status codes, media types, response bodies, all compared to the spec, not to what the agent thinks it should return. If the command fails, the agent hasn't finished, no matter how confident the summary sounds.

The trap to avoid

Don't let the command become a ritual the agent memorizes and you stop reading. If every task ends with the same npm test and you always say "looks good," you've just moved the trust problem one layer down.

Rotate the command with the task. A refactor of the response shape needs a contract check, not a unit test. A race-condition fix needs a repeated-run command. A new endpoint needs a curl against the live route. The command should match the risk, not match the agent's muscle memory.

Try this on your next session

Pick the next agent task you were about to close without running anything. Before you merge, ask for the command. Most of the time you'll get one in ten seconds and run it. The one time you don't — the time the agent hesitates — is the time this rule just saved you from a merge you'd regret.

Top comments (0)