I like /goal in Claude Code. You say what "done" looks like and Claude keeps working until it gets there. What bothered me is who decides it got there.
After every turn a small model (Haiku) reads the conversation and answers "met", "not yet" or "impossible". I tested what that judge can see. It never runs a command and never opens a file. If the agent writes "All done, tests pass!", the judge can just believe it.
So I measured it.
The benchmark
Four small tasks: a ledger, a text toolkit, a job queue and a spreadsheet engine. Each one has a hidden grader that the agent never sees. I ran Sonnet 5.5 and Haiku 4.5 on them, with plain /goal and with my plugin, and graded every run three times. 33 valid runs in total.
The judge said "met" in all 17 plain runs. Eight of those were broken: one Sonnet run and seven Haiku runs. The worst one was a Haiku run on the spreadsheet task that passed 4% of the hidden checks, and /goal still called it done.
What I built
goalpost is a Claude Code plugin made of hooks. You keep typing /goal like before. Behind it:
- The goal becomes a spec with acceptance criteria. Each criterion has a command that exits 0 only when it holds.
- Claude can't stop while a criterion has no passing check, and a check only counts if it ran after the last edit.
- Tests that existed when the goal started are read-only, so they can't be edited to fit the code.
- When everything passes, a fresh auditor agent re-runs the checks and looks for stubs, skipped tests and hardcoded answers.
- After context compaction the spec and the progress notes come back.
It's hooks, not a prompt, so the model can't talk its way past it.
One example from the benchmark: in a Sonnet run on the spreadsheet task, the agent changed the expected values in four of its own tests after the spec was frozen, for example AND(A3,A1) from #VALUE! to TRUE. The auditor caught it ("expectation changed after the freeze, undeclared"), the agent explained the change, and the run finished with all 73 hidden checks green. A plain run has nobody looking at that.
The honest numbers
plain /goal
|
with goalpost | |
|---|---|---|
| Sonnet 5.5, runs fully correct | 8 of 9 | 9 of 9 |
| Haiku 4.5, hidden checks passed (mean) | 79.8% | 89.7% |
| Haiku 4.5, runs fully correct | 1 of 8 | 1 of 7 |
| Time per run (Sonnet / Haiku) | 2.7 / 5.7 min | 5.2 / 15.1 min |
| Cost per run (Sonnet / Haiku) | $0.46 / $0.64 | $0.91 / $1.25 |
Sonnet barely needed help on these tasks, and one extra perfect run is within noise. Haiku improved but it didn't become reliable: only one of its seven runs with goalpost was fully correct, and the auditor is a Haiku too, so it can wave things through. And it costs about twice the time and money. I'd use it for goals that matter, and turn it off for small ones (/goalpost:off).
The runs are small because my account's rate limit cut several of them. The full report lists every run I threw out.
Try it
claude plugin marketplace add syntaxixr/goalpost
claude plugin install goalpost@goalpost
Then open a new Claude Code session and type /goal as usual. It needs Node.js 18 or newer. The repo has the benchmark, the notes on how /goal works under the hood, and a 90-second video of two real runs.
If you've seen /goal end too early on your own work, I'd like to hear about it.
Heads up: I drafted this post with an AI assistant. The numbers come from the benchmark report in the repo.

Top comments (0)