DEV Community

Kaiji
Kaiji

Posted on

One Command, an Agent That Runs All Night: 6 Rules for Writing /goal Text

The agents that die overnight usually don't die because the model is weak. They die because the goal text has no finish line in it. The agent can't tell whether to stop or keep going — so it does both: it stops when it shouldn't, and it keeps going after everything is already broken.

That's not my theory. OpenAI published an official Cookbook post this year, Using Goals in Codex, about exactly this problem. I read it and condensed the parts you can actually use into six rules. At the end: an open-source tool I built that drafts goal text for you, with a real run as proof.

What /goal is

Codex shipped /goal in 0.128.0; since 0.133.0 it's been on by default. The usage is one sentence: give it a goal text, the agent hangs the goal on the thread, checks evidence after each round, keeps working when criteria are unmet, and stops when they're met. You don't stand behind it saying "continue".

/goal Reduce p95 latency below 120 ms without regressing correctness tests   # set
/goal          # view
/goal edit     # edit
/goal pause    # pause
/goal resume   # resume
/goal clear    # clear
Enter fullscreen mode Exit fullscreen mode

The goal lives on the thread in four states: running, paused, done, budget exhausted. Remember the last one — that's rule 5.

Other harnesses are building the same loop mechanism under different names. They all eat the same input: a plain goal text. So these six rules don't pick a tool.

The six rules

1. A goal must carry numbers.

The official contrast. Weak:

/goal Improve performance
Enter fullscreen mode Exit fullscreen mode

Strong:

/goal Reduce p95 latency below 120 ms on the checkout benchmark while keeping the correctness suite green
Enter fullscreen mode Exit fullscreen mode

The second one defines "done" as a measurable state. p95 went 180 → 135? Not done. Under 120 but the tests are red? Also not done. The Cookbook hammers one point: completion is decided by evidence. The model believing it's done doesn't count.

2. Copy the official six-element template.

The strongest goals state six things, and they come in three pairs. Result and verification surface answer what counts as done. Constraints and scope fence in what may be touched. Iteration policy and stop conditions govern everything in between. The Cookbook ships a template you just fill in — and what a filled-in one looks like, see the screenshot at the end of this post: a real feature discussion condensed into exactly one of these.

/goal <desired end state> verified by <specific evidence> while preserving <constraints>.
Use <allowed files>. Between iterations, <how to pick the next step>.
If blocked, <what to report>.
Enter fullscreen mode Exit fullscreen mode

3. Scope: neither too wide nor too narrow.

Ask it to fix one outlet when the disease lives in the breaker box, and it sits trapped in a single file. Ask it to "optimize the whole system" with no acceptance surface, and it can never declare victory. The official recommendation is the middle ground: all tests on the current branch pass, public interfaces untouched.

4. Details go in an attachment.

Goal text caps at 4000 characters. Put the overflow in a project file and add one line to the goal: read that file first, confirm, then execute. I keep mine in .goal/SPEC.md — the filename doesn't matter, the pattern does.

5. Budget and stop conditions are mandatory.

When the budget runs out, the agent parks in a budget exhausted state: summary only, no new work. The official post is explicit that this state is not "done". So don't sign the paper and go to sleep. A common stop condition: the same approach failing three times in a row.

6. Three kinds of work don't deserve a goal.

One-line fixes — just do them, a goal is ceremony. Work with no definable done-criteria — "make the code better" is a wish, not a criterion, so figure out what you actually want first. Anything touching production credentials or needing a human call — stay at the keyboard. Use plain conversation for those, and /plan first for big tasks.

What if you can't write one?

The Cookbook knows these goals are hard to write. Its suggestion is two steps: describe the task in plain words, let Codex draft the goal, then you tighten the acceptance criteria and the stop conditions.

I followed that idea and built an open-source skill, find-my-goal. You say a wish ("help me optimize my project"), it asks a few multiple-choice questions — max three rounds, and every question has a "you decide" option — then it actually reads your project, runs a baseline, and translates "optimize" into numbered criteria. What comes out is a six-element goal text you paste into /goal. It ships with an anti-cheating clause too: the agent may not modify tests or benchmarks to satisfy the criteria. And if a goal you already set is spinning in place, hand it over for an audit.

GitHub logo Kaiji-Z / find-my-goal

/goal drafter. Say a goal, answer a few questions, get a strong goal with executable criteria and brakes. 零门槛 /goal 起草器——说目标→答选择题→产出codex官方级别的强目标。 (EN/中文)

find-my-goal

If you can type, you can use /goal. · 中文说明

❌ /goal optimize my project
   → no criteria, no budget, no brake: the loop wanders, burns tokens, stalls halfway
✅ /goal Goal: cut full npm test time from baseline 84s to under 40s, all green
        Scope: only src/ and tests; no public API changes, no new deps
        Done when: npm test exit 0; 3 consecutive runs each ≤ 40s
        Stop if: needs new deps; same idea fails 3 times
        Budget: max 15 iterations
   → runs until the evidence says done

You don't learn to write the right-hand side — find-my-goal writes it for you. Say it in plain words, answer a few multiple-choice questions, paste the draft into /goal. Works in English or Chinese.

License: MIT Skill

TL;DR

# Install (one-time, pick one)
npx skills add Kaiji-Z/find-my-goal                                   # any agent (skills.sh CLI)
git clone https://github.com/Kaiji-Z/find-my-goal ~/.claude/skills/   # Claude Code
git
…
Enter fullscreen mode Exit fullscreen mode

Here's a real run from my machine. One feature discussion, condensed into a complete five-part goal draft — Goal / Scope / Done when / Stop if / Budget:

A real find-my-goal run: a feature discussion condensed into a complete goal draft

You probably have a flopped goal text sitting in your own history. Paste it in the comments — I'll break it down against these six rules.

Most people judge agents by the model: parameters, benchmarks. The real dividing line sits in daily operations. Same person, same model — a clear goal text runs all night and finishes clean; a fuzzy one falls apart in two rounds. The line is too ordinary, and that's exactly why nobody treats it as a capability.

Top comments (0)