Every failure I have had with coding agents looks the same at the end. The agent writes a summary that reads like
success: the feature is implemented, the tests pass, the branch is ready. Then I open the thing, and half of it is not
there.
The interesting part is that the model is not lying. It has no way to check. It generated the most plausible
continuation of the conversation, and "done, with a tidy summary" is an extremely plausible continuation. Nothing in
the loop ever asked for evidence.
So I stopped trying to make the model more honest and changed the shape of the loop instead. In the agentic workforce I run now, no
model is allowed to say done. A command says done, and it has to exit 0.
The proof command gets written before the work exists
When I type an ask, the first thing that happens is not work. It is a contract.
/crew add a pricing page with a monthly / annual toggle
The card that opens carries six things:
- goal - what this is, in my words
- artifact - the file or surface that changes
- where it lands - the branch
- done when - the same thing, in plain language
- proof command - the exact command whose exit code decides
- budget - tokens, so a run cannot spin
The order is the point. The proof is frozen while the work does not exist yet. At that moment I am uninformed and
honest, so I write:
npm run test:e2e -- pricing.spec.ts
and not "looks right to me". Later, with work already on the branch and the temptation to call it finished, the
command is a fact from an earlier conversation. Neither I nor the agent can quietly move it. That is most of the value:
not detection, but removing the option.
Done is decided by a tool, not by a model
In the workforce, done is not a status an agent sets. It is the exit code of that command, run as a tool call, at the point
where the card would otherwise close.
- the worker does the work
- the proof runs, exactly as written when the card opened; a non-zero exit is the only failure signal that counts
- the coordinator re-runs the proof before the card can close, so one lucky green run does not close it
- a card that asks for a stronger guarantee gets one independent verifier run instead
The model writes the code and the summary. It never writes the verdict.
One owner per card, one writer per card
The second failure mode is drift: three agents on the same branch, each doing something locally reasonable, together
producing nothing. So a card has one writer, and a card has one owner - a coordinator that wakes on every card event
with the whole record in hand and takes exactly one decision: retry, rescope, split, or ask me.
The asks are bounded too. Before any work starts, the chat asks at most five questions in one batch, or none. After
that I am out of the loop until the card ends, and it ends either with a verified result in the thread I asked from, or
with one concrete question. Token budgets are per card, so a run that reaches its budget stops instead of burning
through the rest of the night, and a worker that fails five tool calls in a row stops itself.
What it looks like from the outside
Illustrative card, from the demo board:
16:12:04 /crew add a pricing page with a monthly / annual toggle
16:12:09 contract: artifact app/pricing/page.tsx
proof npm run test:e2e -- pricing.spec.ts
16:12:11 card opened - coordinator owns it, your turn ends
...
16:41:21 done - proof exit 0 - 2 runs, 1 fix
That last line is what reaches the chat: not a summary of feelings, the proof that passed and how many attempts it
took. If the proof had failed twice, I would have got one question instead, with the two failure lines attached.
Where this shape stops working
Three honest limits, because a post about verification that only lists wins is its own kind of bullshit.
-
A proof proves what you wrote, not what you meant.
test:e2e -- pricing.spec.tscan be green while the toggle does nothing, if the test asserts the wrong thing. The leverage is all in writing the command well, and that is human work. - It is bad for work where done is a judgement. Naming, positioning, "write the launch post" - no command decides those. The workforce falls back to a person, which is correct behaviour, but it means this shape only pays off on tasks with a checkable end.
- Proofs cost a minute each. On trivial work that is a bad trade. On the work you would otherwise re-check at 23:00 anyway, it is the cheapest minute in the task.
There is also a fourth thing I did not expect: because the proof fixes the artifact, "can you also just..." turns into
a second card instead of a scope creep inside the first one. Annoying, and correct.
The pattern is the point, not my plugin
The generalisable rule is not complicated:
- before the agent starts, write down an executable definition of done;
- let only that definition close the loop;
- let a tool run it, not the model that did the work.
Any agent framework can do this. Most do not, because "done" is much easier to print than to prove.
I built it as a plugin for Hermes Agent - Apache-2.0, self-hosted, runs on your own machine, no telemetry - and the
whole thing is github.com/macd2/Crew:
hermes plugins install macd2/Crew
python3 ~/.hermes/plugins/crew/install.py --profile NAME
hermes -p NAME plugins doctor crew
If you would rather watch it before installing anything - the 53 second launch film, one ask through to the verified
result:
crew.forgecoreai.com is the same loop as a page, if you would rather read than watch.
So the question I would actually like answered in the comments: what does done mean in your setup? Do you have a
command, or is it a feeling? I am collecting the ones people genuinely trust - a solo dev with pytest, a team with a
deploy smoke test, someone checking a number in a dashboard. A proof you cannot write is a task you cannot delegate,
and the interesting stories are the tasks where you had to admit that.
Drafted with an AI assistant and edited by me.
Top comments (0)