<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Kaiji</title>
    <description>The latest articles on DEV Community by Kaiji (@kaiji1126).</description>
    <link>https://dev.to/kaiji1126</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4106482%2F787614ba-b364-4830-a1f5-88bbaeb61024.jpg</url>
      <title>DEV Community: Kaiji</title>
      <link>https://dev.to/kaiji1126</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/kaiji1126"/>
    <language>en</language>
    <item>
      <title>One Command, an Agent That Runs All Night: 6 Rules for Writing /goal Text</title>
      <dc:creator>Kaiji</dc:creator>
      <pubDate>Mon, 14 Sep 2026 09:48:44 +0000</pubDate>
      <link>https://dev.to/kaiji1126/one-command-an-agent-that-runs-all-night-6-rules-for-writing-goal-text-3ddm</link>
      <guid>https://dev.to/kaiji1126/one-command-an-agent-that-runs-all-night-6-rules-for-writing-goal-text-3ddm</guid>
      <description>&lt;p&gt;The agents that die overnight usually don't die because the model is weak. They die because the goal text has no finish line in it. The agent can't tell whether to stop or keep going — so it does both: it stops when it shouldn't, and it keeps going after everything is already broken.&lt;/p&gt;

&lt;p&gt;That's not my theory. OpenAI published an official Cookbook post this year, &lt;em&gt;Using Goals in Codex&lt;/em&gt;, about exactly this problem. I read it and condensed the parts you can actually use into six rules. At the end: an open-source tool I built that drafts goal text for you, with a real run as proof.&lt;/p&gt;

&lt;h2&gt;
  
  
  What /goal is
&lt;/h2&gt;

&lt;p&gt;Codex shipped &lt;code&gt;/goal&lt;/code&gt; in 0.128.0; since 0.133.0 it's been on by default. The usage is one sentence: give it a goal text, the agent hangs the goal on the thread, checks evidence after each round, keeps working when criteria are unmet, and stops when they're met. You don't stand behind it saying "continue".&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;/goal Reduce p95 latency below 120 ms without regressing correctness tests   &lt;span class="c"&gt;# set&lt;/span&gt;
/goal          &lt;span class="c"&gt;# view&lt;/span&gt;
/goal edit     &lt;span class="c"&gt;# edit&lt;/span&gt;
/goal pause    &lt;span class="c"&gt;# pause&lt;/span&gt;
/goal resume   &lt;span class="c"&gt;# resume&lt;/span&gt;
/goal clear    &lt;span class="c"&gt;# clear&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The goal lives on the thread in four states: running, paused, done, budget exhausted. Remember the last one — that's rule 5.&lt;/p&gt;

&lt;p&gt;Other harnesses are building the same loop mechanism under different names. They all eat the same input: a plain goal text. So these six rules don't pick a tool.&lt;/p&gt;
&lt;h2&gt;
  
  
  The six rules
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. A goal must carry numbers.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The official contrast. Weak:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;/goal Improve performance
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Strong:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;/goal Reduce p95 latency below 120 ms on the checkout benchmark &lt;span class="k"&gt;while &lt;/span&gt;keeping the correctness suite green
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The second one defines "done" as a measurable state. p95 went 180 → 135? Not done. Under 120 but the tests are red? Also not done. The Cookbook hammers one point: completion is decided by evidence. The model &lt;em&gt;believing&lt;/em&gt; it's done doesn't count.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Copy the official six-element template.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The strongest goals state six things, and they come in three pairs. Result and verification surface answer what counts as done. Constraints and scope fence in what may be touched. Iteration policy and stop conditions govern everything in between. The Cookbook ships a template you just fill in — and what a filled-in one looks like, see the screenshot at the end of this post: a real feature discussion condensed into exactly one of these.&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;/goal &amp;lt;desired end state&amp;gt; verified by &amp;lt;specific evidence&amp;gt; &lt;span class="k"&gt;while &lt;/span&gt;preserving &amp;lt;constraints&amp;gt;.
Use &amp;lt;allowed files&amp;gt;. Between iterations, &amp;lt;how to pick the next step&amp;gt;.
If blocked, &amp;lt;what to report&amp;gt;.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;&lt;strong&gt;3. Scope: neither too wide nor too narrow.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Ask it to fix one outlet when the disease lives in the breaker box, and it sits trapped in a single file. Ask it to "optimize the whole system" with no acceptance surface, and it can never declare victory. The official recommendation is the middle ground: all tests on the current branch pass, public interfaces untouched.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Details go in an attachment.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Goal text caps at 4000 characters. Put the overflow in a project file and add one line to the goal: read that file first, confirm, then execute. I keep mine in &lt;code&gt;.goal/SPEC.md&lt;/code&gt; — the filename doesn't matter, the pattern does.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5. Budget and stop conditions are mandatory.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;When the budget runs out, the agent parks in a &lt;em&gt;budget exhausted&lt;/em&gt; state: summary only, no new work. The official post is explicit that this state is not "done". So don't sign the paper and go to sleep. A common stop condition: the same approach failing three times in a row.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;6. Three kinds of work don't deserve a goal.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;One-line fixes — just do them, a goal is ceremony. Work with no definable done-criteria — "make the code better" is a wish, not a criterion, so figure out what you actually want first. Anything touching production credentials or needing a human call — stay at the keyboard. Use plain conversation for those, and &lt;code&gt;/plan&lt;/code&gt; first for big tasks.&lt;/p&gt;
&lt;h2&gt;
  
  
  What if you can't write one?
&lt;/h2&gt;

&lt;p&gt;The Cookbook knows these goals are hard to write. Its suggestion is two steps: describe the task in plain words, let Codex draft the goal, then you tighten the acceptance criteria and the stop conditions.&lt;/p&gt;

&lt;p&gt;I followed that idea and built an open-source skill, &lt;strong&gt;find-my-goal&lt;/strong&gt;. You say a wish ("help me optimize my project"), it asks a few multiple-choice questions — max three rounds, and every question has a "you decide" option — then it actually reads your project, runs a baseline, and translates "optimize" into numbered criteria. What comes out is a six-element goal text you paste into /goal. It ships with an anti-cheating clause too: the agent may not modify tests or benchmarks to satisfy the criteria. And if a goal you already set is spinning in place, hand it over for an audit.&lt;/p&gt;


&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/Kaiji-Z" rel="noopener noreferrer"&gt;
        Kaiji-Z
      &lt;/a&gt; / &lt;a href="https://github.com/Kaiji-Z/find-my-goal" rel="noopener noreferrer"&gt;
        find-my-goal
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      /goal drafter. Say a goal, answer a few questions, get a strong goal with executable criteria and brakes. 零门槛 /goal 起草器——说目标→答选择题→产出codex官方级别的强目标。 (EN/中文)
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;div class="markdown-heading"&gt;
&lt;h1 class="heading-element"&gt;find-my-goal&lt;/h1&gt;
&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;If you can type, you can use /goal.&lt;/strong&gt; · &lt;a href="https://github.com/Kaiji-Z/find-my-goal/README.zh-CN.md" rel="noopener noreferrer"&gt;中文说明&lt;/a&gt;&lt;/p&gt;

&lt;div class="snippet-clipboard-content notranslate position-relative overflow-auto"&gt;&lt;pre class="notranslate"&gt;&lt;code&gt;❌ /goal optimize my project
   → no criteria, no budget, no brake: the loop wanders, burns tokens, stalls halfway
✅ /goal Goal: cut full npm test time from baseline 84s to under 40s, all green
        Scope: only src/ and tests; no public API changes, no new deps
        Done when: npm test exit 0; 3 consecutive runs each ≤ 40s
        Stop if: needs new deps; same idea fails 3 times
        Budget: max 15 iterations
   → runs until the evidence says done
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;You don't learn to write the right-hand side — find-my-goal writes it for you. Say it in plain words, answer a few multiple-choice questions, paste the draft into &lt;code&gt;/goal&lt;/code&gt;. Works in English or Chinese.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://github.com/Kaiji-Z/find-my-goal/LICENSE" rel="noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/08cef40a9105b6526ca22088bc514fbfdbc9aac1ddbf8d4e6c750e3a88a44dca/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f4c6963656e73652d4d49542d626c75652e737667" alt="License: MIT"&gt;&lt;/a&gt;
&lt;a href="https://github.com/Kaiji-Z/find-my-goal/find-my-goal/SKILL.md" rel="noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/4a3a18f609c999db6982865fff0dee4b25a1173b72adb20fe061dfceab2927a9/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f747970652d6167656e742d2d736b696c6c2d677265656e" alt="Skill"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;TL;DR&lt;/h2&gt;

&lt;/div&gt;

&lt;div class="highlight highlight-source-shell notranslate position-relative overflow-auto js-code-highlight"&gt;
&lt;pre&gt;&lt;span class="pl-c"&gt;&lt;span class="pl-c"&gt;#&lt;/span&gt; Install (one-time, pick one)&lt;/span&gt;
npx skills add Kaiji-Z/find-my-goal                                   &lt;span class="pl-c"&gt;&lt;span class="pl-c"&gt;#&lt;/span&gt; any agent (skills.sh CLI)&lt;/span&gt;
git clone https://github.com/Kaiji-Z/find-my-goal &lt;span class="pl-k"&gt;~&lt;/span&gt;/.claude/skills/   &lt;span class="pl-c"&gt;&lt;span class="pl-c"&gt;#&lt;/span&gt; Claude Code&lt;/span&gt;
git&lt;/pre&gt;…
&lt;/div&gt;&lt;/div&gt;
  &lt;/div&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/Kaiji-Z/find-my-goal" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;/div&gt;



&lt;p&gt;Here's a real run from my machine. One feature discussion, condensed into a complete five-part goal draft — Goal / Scope / Done when / Stop if / Budget:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5zngbtdf8h208pkqgfj7.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5zngbtdf8h208pkqgfj7.png" alt="A real find-my-goal run: a feature discussion condensed into a complete goal draft" width="800" height="1405"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;You probably have a flopped goal text sitting in your own history. Paste it in the comments — I'll break it down against these six rules.&lt;/p&gt;

&lt;p&gt;Most people judge agents by the model: parameters, benchmarks. The real dividing line sits in daily operations. Same person, same model — a clear goal text runs all night and finishes clean; a fuzzy one falls apart in two rounds. The line is too ordinary, and that's exactly why nobody treats it as a capability.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>codex</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Stop Being Your Agent's QA: A Skill That Makes AI Coding Agents Prove Their Work</title>
      <dc:creator>Kaiji</dc:creator>
      <pubDate>Wed, 02 Sep 2026 16:11:56 +0000</pubDate>
      <link>https://dev.to/kaiji1126/stop-being-your-agents-qa-a-skill-that-makes-ai-coding-agents-prove-their-work-4gkn</link>
      <guid>https://dev.to/kaiji1126/stop-being-your-agents-qa-a-skill-that-makes-ai-coding-agents-prove-their-work-4gkn</guid>
      <description>&lt;p&gt;Every AI coding session ends the same way: the agent says "Done ✓" — and you spend the next 30 minutes manually clicking through everything to find out it isn't.&lt;/p&gt;

&lt;p&gt;I got tired of being my agent's test suite. So I built &lt;strong&gt;stop-manual-testing&lt;/strong&gt; — a skill that flips the loop: every task ends with machine-checkable criteria the agent must actually run and pass before it's allowed to say "done".&lt;/p&gt;

&lt;h2&gt;
  
  
  How it works
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Every task ends with a verification list the agent can execute itself: commands, checks, expected outputs&lt;/li&gt;
&lt;li&gt;If a check can't be automated, the skill tells you exactly what to verify by hand — and why&lt;/li&gt;
&lt;li&gt;The agent converges in a closed loop instead of handing you "Done ✓" and hoping for the best&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It's modular, checks can run in parallel, and it works across agent-driven workflows.&lt;/p&gt;


&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/Kaiji-Z" rel="noopener noreferrer"&gt;
        Kaiji-Z
      &lt;/a&gt; / &lt;a href="https://github.com/Kaiji-Z/stop-manual-testing" rel="noopener noreferrer"&gt;
        stop-manual-testing
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      Stop manually testing your AI agent. A skill that builds a machine-checkable verification system so the agent self-verifies and converges in a closed loop. | 让 AI agent 自己搭机器可判定的验证体系,在闭环里自己收敛,停止手动测试。
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;div class="markdown-heading"&gt;
&lt;h1 class="heading-element"&gt;stop-manual-testing&lt;/h1&gt;
&lt;/div&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;A skill that stops you from manually testing your AI agent. The agent reads it, builds a machine-checkable verification system, and self-converges in a closed loop — collapsing the ~90% of dev time you spend staring at runs and judging by gut feel.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/Kaiji-Z/stop-manual-testing#%E4%B8%AD%E6%96%87%E7%89%88" rel="noopener noreferrer"&gt;中文&lt;/a&gt; · English&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;div class="markdown-heading"&gt;
&lt;h1 class="heading-element"&gt;English&lt;/h1&gt;
&lt;/div&gt;
&lt;p&gt;A &lt;a href="https://platform.claude.com/docs/en/agents-and-tools/agent-skills/overview" rel="nofollow noopener noreferrer"&gt;skill&lt;/a&gt; for AI coding agents (Claude Code / Codex / ZCode / Cursor). Load it once, and instead of you manually clicking through the UI and eyeballing whether the agent "got better or worse this run," the agent builds itself a verification system where correctness is machine-checkable — then iterates inside a closed loop until it converges.&lt;/p&gt;
&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;The pain it solves&lt;/h2&gt;
&lt;/div&gt;
&lt;p&gt;If you develop AI agents, you are likely stuck here:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;~90% of your time goes to manual testing&lt;/strong&gt; — opening the web page, typing inputs, watching the agent step by step, judging with your own brain whether it got better…&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;
  &lt;/div&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/Kaiji-Z/stop-manual-testing" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;/div&gt;


&lt;p&gt;If you're spending more time reviewing your agent's work than actually creating, give it a try — feedback welcome.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>devtools</category>
      <category>productivity</category>
    </item>
  </channel>
</rss>
