DEV Community

San
San

Posted on

My specs never told me what to do next, so I built my own SDD harness

I do spec-driven development with an AI agent. I tried some well-known spec-driven harnesses, then built my own tool, and for months now I have used only that. This post explains why I needed my own tool and when it is useful.

Repo: https://github.com/opellen/opengoal

The question my specs did not answer

Like many developers who got hooked on AI coding, I was working on several projects at the same time, like crazy. At some point I felt the need for spec-driven development, so I used the well-known harness frameworks. But I ran into a bottleneck:

"What do I do next?"

The specs did not tell me what to do next.

As a workaround, I made a page for each project in Notion and managed things there. But it was a separate place from the project, so keeping it up to date was a hassle. So I wanted state management that lives inside the project.

Why I built my own

My SDD harness started in a hobby project: reverse engineering a game. Reverse engineering takes a lot of analysis and design over many sessions, and designs and plans get overturned all the time. Each time, brainstorming with the AI harness again and rewriting the spec was too much work. If that process runs through a CLI, it is even less agile. I thought about editing the prompts of the existing harnesses, but then I could not keep up with their updates. So I also wanted the prompts of my own harness to be extendable, in a way similar to Jekyll.

There were a few more reasons.

The AI sometimes ignored the later part of a slash command in a harness. I learned that the cause was long text written in a prose-like natural language style. So I started to optimize the prompts. I based this on research showing that pseudocode-style prompts are followed better than natural language and use fewer tokens.

  • EMNLP 2023: pseudocode-style prompts improved F1 scores by 7 to 16 points over natural language.
  • CodeAgents (2025): writing agent workflows in pseudocode reduced token use by 55 to 87% and improved performance by 3 to 36 points.

Forced strict TDD was another reason. Small tasks got the same strict TDD as big ones, and if errors still remained, I had to run the process again. The tokens burned for nothing that way were painful.

One question, one file

My harness is based on one rule: one question, one file. If you want to ask the project something, the answer is in exactly one file.

ROADMAP.md  ─────────────  where am I in the big picture?
     CONTEXT.md  ────────  what is this project?
            ● GOAL.md  ──  what do I do now?
Enter fullscreen mode Exit fullscreen mode

This structure is kept in the docs folder by default. You can also choose another location.

Some harness frameworks have a similar structure. But my harness, OpenGoal, does not use a CLI or hooks. It is plain markdown. Nothing extra is added at runtime, so it runs light. I mostly use it in Claude Code, but it can be installed into about 20 tools, including Cursor and Codex.

Spec-driven harness frameworks expect a carefully structured spec, but that can be too much for small tasks. So OpenGoal can create just a goal. You can split that goal into smaller pieces as far as you need, and for simple work you can skip the design document. I did not throw specs away. I write them to fit the size of the task, and GOAL.md is the spec that answers "what do I do next?"

How I use it

When OpenGoal is not set up in a project yet, or when I want to discuss something with the AI, I start with scout.

You: /opgl:scout
AI:  Found an existing codebase. No CONTEXT.md yet.
     Set up project context with `/opgl:context`?

You: /opgl:context init
AI:  ✓ docs/CONTEXT.md created
     To set a goal: `/opgl:goal init`
Enter fullscreen mode Exit fullscreen mode

When the discussion has given enough context about the goal, I ask it to write the goal like this.

You: /opgl:goal init Migrate auth module from JWT to session-based
AI:  ✓ docs/GOAL.md created (5 tasks)
Enter fullscreen mode Exit fullscreen mode

Then the AI may recommend breaking the goal down further, writing a design document, or, for simple work, starting right away without a design document.

AI:  Recommend `/opgl:goal breakdown` — the tasks contain several hidden steps.
AI:  Recommend `/opgl:design init` — implementation task with file-level decisions.
AI:  Recommend `/opgl:go` — this is a simple task.
Enter fullscreen mode Exit fullscreen mode

When the work needs a design, the design document comes first. In OpenGoal, this document works as the spec for a unit of work.

You: /opgl:design init
AI:  ✓ docs/DESIGN.md created
     Session store selection, migration strategy, rollback plan included
Enter fullscreen mode Exit fullscreen mode

During the work, when there is an important decision, it asks for my opinion.

You: /opgl:go
AI:  Task 1/5: Set up session store... ✓ done
     Task 2/5: Replace middleware... ✓ done
     Task 3/5 requires DB schema changes. Proceed?
Enter fullscreen mode Exit fullscreen mode

When you want to handle longer-term goals, you can write a roadmap.

You: /opgl:roadmap init
AI:  ✓ docs/ROADMAP.md created
     ├── M1: Auth migration
     ├── M2: Rate limiting
     └── M3: API versioning
Enter fullscreen mode Exit fullscreen mode

Finishing and pausing a goal

When all the tasks of a goal are done, I archive it. /opgl:goal archive moves GOAL.md and DESIGN.md, and PLAN.md and CHECKPOINT.md if they exist, into docs/archive/goals/. If there is a roadmap, its milestone is marked done too.

An archived goal is still used after that. When I later fix that part or work on the next goal, the AI finds the design, decisions, and checkpoints from that time in the archive and uses them as context. OpenGoal does not read the archive automatically. The archive stays in a fixed place, and CONTEXT.md or the next goal's documents often point to it, so the AI finds and reads it when needed.

You can save the state of a goal in progress as a checkpoint. I mostly use checkpoints to prepare for compaction, or to hand work over to another session or another AI tool. A checkpoint often keeps details that the compaction summary misses. Compaction shrinks the whole conversation at once, but a checkpoint writes down what the AI has found so far and the next steps separately.

You: /opgl:goal checkpoint
AI:  ✓ docs/CHECKPOINT.md saved
     (If the session drops, the next one picks up where you left off)
Enter fullscreen mode Exit fullscreen mode

Checkpoints are especially useful when you use several AI tools together. I subscribe to Claude Code, Codex, and Antigravity, and I use them together. When I write a checkpoint in Claude Code, move to Codex and just run /opgl:go, the work continues where it stopped.

Natural language and other features

You do not have to type the commands every time. When the context is clear, natural language works too. You can answer a recommended command with "ok" or "go", "proceed with task 6" runs /opgl:go, and "save progress" runs /opgl:goal checkpoint.

Beyond that, through my long dogfooding, OpenGoal came to handle suspending and resuming goals, sub-goals, a backlog, and so on. I squeezed as much as I could out of markdown alone.

Token use

Overall, OpenGoal was also designed to reduce the use of our precious tokens. CONTEXT.md plays an important part in that. The AI knows in advance where things are, so there are fewer tool calls, and it needs to re-explore the codebase less at the start of each session. Also, there are no hooks or MCP servers, so unused tool definitions do not take up tokens. The subagent skill that installs with it serves the same purpose. The main agent decides, by the difficulty of the task, which model to hand a subagent, and it makes very small fixes itself. Subagents gather the material for analysis, but the main agent does the analysis itself. I had the problem in other harnesses where the main agent lost the fine details of the context.

Why markdown alone

The biggest reason I built OpenGoal was this thought: all this machinery, built for a perfection we want to rely on, might become mostly unnecessary as AI improves quickly, and then markdown alone might be enough. And I think that has happened. I have finished various projects with only OpenGoal, from a small Obsidian plugin, to a native mobile editor with Notion-style block editing, to a state machine designer.

Top comments (0)