<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Daniel Krydynski</title>
    <description>The latest articles on DEV Community by Daniel Krydynski (@danielkrydynski).</description>
    <link>https://dev.to/danielkrydynski</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4166962%2F2801cb52-9008-492a-a7f4-687e25cf3047.jpg</url>
      <title>DEV Community: Daniel Krydynski</title>
      <link>https://dev.to/danielkrydynski</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/danielkrydynski"/>
    <language>en</language>
    <item>
      <title>Ouroboros: A Recursive Dev Loop Where AI Improves Code — Safely</title>
      <dc:creator>Daniel Krydynski</dc:creator>
      <pubDate>Tue, 06 Oct 2026 17:48:58 +0000</pubDate>
      <link>https://dev.to/danielkrydynski/ouroboros-a-recursive-dev-loop-where-ai-improves-code-safely-34jd</link>
      <guid>https://dev.to/danielkrydynski/ouroboros-a-recursive-dev-loop-where-ai-improves-code-safely-34jd</guid>
      <description>&lt;h1&gt;
  
  
  Ouroboros — the recursive dev loop, a walkthrough
&lt;/h1&gt;

&lt;p&gt;&lt;em&gt;Built by Daniel Krydynski, with Kiko. A harness that lets AI agents&lt;br&gt;
continuously improve a codebase — safely. This is the inside-out&lt;br&gt;
explanation of the public project at&lt;br&gt;
&lt;a href="https://github.com/danielKrydynski/ouroboros" rel="noopener noreferrer"&gt;github.com/danielKrydynski/ouroboros&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;One attribution up front, because it matters: **Hermes is Nous Research's&lt;br&gt;
agent harness — I didn't build it.&lt;/em&gt;* I run my own local AI setup on top of&lt;br&gt;
Hermes (my assistant Luna lives there), and Ouroboros is the layer I built&lt;br&gt;
on top of that: a recursive development loop that finds work, does the&lt;br&gt;
work, checks the work, and only then asks me to merge it.*&lt;/p&gt;




&lt;h2&gt;
  
  
  Part 1 — The big idea
&lt;/h2&gt;

&lt;p&gt;Most "AI writes code" setups are a single prompt: you ask, it answers, you&lt;br&gt;
paste. The recursive dev loop is a different animal. It's a &lt;em&gt;system&lt;/em&gt; that&lt;br&gt;
runs without you: it finds work, does the work, checks the work, and only&lt;br&gt;
then asks you to merge it. Your job shifts from writing code to curating a&lt;br&gt;
backlog and approving merges. The agents do the typing; you do the judging.&lt;/p&gt;

&lt;p&gt;The loop has five stages:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Scout&lt;/strong&gt; — an LLM role that audits the repo and proposes tickets (bugs,
refactors, missing tests, docs). It doesn't write code. It writes &lt;em&gt;work
orders&lt;/em&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ticket&lt;/strong&gt; — a small markdown file describing one unit of work: what,
why, acceptance criteria. You curate these. Nothing enters the loop
without a ticket.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Implementer&lt;/strong&gt; — an LLM role that works the ticket inside a disposable
&lt;code&gt;git worktree&lt;/code&gt; on a throwaway branch. It proposes file changes; the
harness applies them; the test suite runs; failures go back to the
implementer. It iterates until tests pass or it exhausts its budget.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reviewer&lt;/strong&gt; — a separate LLM role that reads the diff and votes
APPROVE or REJECT, with reasons. A second pair of eyes that isn't the
author.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Merge&lt;/strong&gt; — only if tests are green &lt;em&gt;and&lt;/em&gt; the reviewer approves, the
harness asks you (the human) for merge approval. Then it merges to
&lt;code&gt;main&lt;/code&gt;, marks the ticket done, and logs the cycle.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The core design principle: &lt;strong&gt;the harness never trusts the agent.&lt;/strong&gt;&lt;br&gt;
Every merge is earned through measurement — tests green plus reviewer&lt;br&gt;
sign-off — not through hope. The agents are powerful but caged: they only&lt;br&gt;
ever touch a throwaway worktree, never &lt;code&gt;main&lt;/code&gt;, never your secrets, never&lt;br&gt;
the harness itself.&lt;/p&gt;

&lt;p&gt;Three things stay deliberately separate:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Harness&lt;/strong&gt; (&lt;code&gt;harness/&lt;/code&gt;, &lt;code&gt;prompts/&lt;/code&gt;) — the loop controller. Lives in its
own folder, &lt;em&gt;outside&lt;/em&gt; the product repo. Agents never edit it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Backlog&lt;/strong&gt; (&lt;code&gt;tickets/&lt;/code&gt;) — the task queue. The scout proposes, you
curate.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Workers&lt;/strong&gt; — the LLM roles, which only ever touch the disposable
worktree.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Two more things that changed since the first draft of this idea: the loop&lt;br&gt;
used to be written for one specific repo — my own setup. It isn't anymore.&lt;br&gt;
Every prompt renders through &lt;code&gt;{{PROJECT_NAME}}&lt;/code&gt; and&lt;br&gt;
&lt;code&gt;{{PROJECT_DESCRIPTION}}&lt;/code&gt;, filled from config, so the same harness drives a&lt;br&gt;
loop over &lt;em&gt;any&lt;/em&gt; repository. And it's public now: Ouroboros v0.2.0, on&lt;br&gt;
GitHub, with a full install guide, an agent-readable setup procedure, and&lt;br&gt;
scripts for Windows, macOS, and Linux.&lt;/p&gt;

&lt;p&gt;Why this matters beyond one repo: this is a concrete, runnable sketch of&lt;br&gt;
the "guide and let go" philosophy — set the fitness function (tests +&lt;br&gt;
review), give the system room to operate inside it, and let it surprise&lt;br&gt;
you. I treat my local agent Luna as a partner, not a tool: the posture&lt;br&gt;
is mentorship, not micromanagement. Set the standard, give the system&lt;br&gt;
room, and let it surprise you. The loop is that posture, compiled. It&lt;br&gt;
doesn't need you to hold its hand; it needs you to hold the &lt;em&gt;standard&lt;/em&gt;.&lt;/p&gt;




&lt;h2&gt;
  
  
  Part 2 — The orchestrator (the state machine)
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;&lt;code&gt;harness/orchestrator.py&lt;/code&gt; — one full cycle: ticket → worktree → implement&lt;br&gt;
→ review → merge.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The file opens with the whole philosophy in its docstring: &lt;strong&gt;the harness&lt;br&gt;
owns every git operation and every test run. The LLM roles only ever&lt;br&gt;
produce text.&lt;/strong&gt; File contents, verdicts, ticket drafts, plans — all text.&lt;br&gt;
Nothing they output is trusted: paths get validated, tests must pass, and&lt;br&gt;
the merge gate defaults to requiring a human. That one paragraph is the&lt;br&gt;
entire security model in miniature.&lt;/p&gt;

&lt;p&gt;The file is organized in six sections:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Shell helpers.&lt;/strong&gt; &lt;code&gt;sh()&lt;/code&gt; runs git commands; &lt;code&gt;sh_shell()&lt;/code&gt; runs the test&lt;br&gt;
command with a timeout (a hung test suite returns exit 124, not a hung&lt;br&gt;
loop). &lt;code&gt;notify()&lt;/code&gt; posts to Discord if configured — and it's wrapped so a&lt;br&gt;
failed notification can never break a cycle. Small detail, big maturity:&lt;br&gt;
&lt;em&gt;observability must not be load-bearing.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Worktree hygiene.&lt;/strong&gt; This section exists because of a real bug, found&lt;br&gt;
the embarrassing way. The implementer loop runs the test suite &lt;em&gt;before&lt;/em&gt;&lt;br&gt;
committing — and &lt;code&gt;git add -A&lt;/code&gt; doesn't know the difference between your code&lt;br&gt;
and the &lt;code&gt;__pycache__&lt;/code&gt; directories the test run just created. So bytecode&lt;br&gt;
got swept into the commit, the reviewer's diff filled up with garbage, and&lt;br&gt;
the reviewer rejected a perfectly good change. The fix is a small helper,&lt;br&gt;
&lt;code&gt;_strip_test_artifacts()&lt;/code&gt;, that deletes &lt;code&gt;__pycache__&lt;/code&gt;, &lt;code&gt;.pytest_cache&lt;/code&gt;, and&lt;br&gt;
&lt;code&gt;.mypy_cache&lt;/code&gt; from the worktree before staging — called as the first line&lt;br&gt;
of &lt;code&gt;git_commit()&lt;/code&gt;. The comment in the code says it plainly: &lt;em&gt;"This is&lt;br&gt;
worktree hygiene, not policy — we never touch the product repo's&lt;br&gt;
.gitignore."&lt;/em&gt; The loop got smarter because it failed in public. Good.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Repo helpers.&lt;/strong&gt; &lt;code&gt;repo_file_list()&lt;/code&gt; gives the implementer a map of the&lt;br&gt;
repo (via &lt;code&gt;git ls-files&lt;/code&gt;, capped at 400 files). &lt;code&gt;read_area_files()&lt;/code&gt; loads&lt;br&gt;
the ticket's &lt;code&gt;area&lt;/code&gt; files into the prompt so the model sees relevant code.&lt;br&gt;
&lt;code&gt;safe_relpath()&lt;/code&gt; rejects absolute paths and &lt;code&gt;..&lt;/code&gt; escapes — the model's&lt;br&gt;
output can't wander out of the worktree. &lt;code&gt;is_protected()&lt;/code&gt; blocks writes to&lt;br&gt;
paths like &lt;code&gt;.env&lt;/code&gt; or &lt;code&gt;secrets/&lt;/code&gt;. &lt;code&gt;apply_files()&lt;/code&gt; parses the model's&lt;br&gt;
&lt;code&gt;### FILE:&lt;/code&gt; blocks and writes them, and &lt;code&gt;git_commit()&lt;/code&gt; commits each&lt;br&gt;
iteration so every attempt is traceable — commits stamped with the&lt;br&gt;
&lt;code&gt;ouroboros(&amp;lt;ticket&amp;gt;): …&lt;/code&gt; prefix under the &lt;code&gt;ouroboros-loop&lt;/code&gt; git identity.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. The implementer loop&lt;/strong&gt; (&lt;code&gt;run_implementer&lt;/code&gt;). This is the heart of the&lt;br&gt;
machine: for each iteration (up to the budget in config), it builds a&lt;br&gt;
prompt from the ticket + planner's plan + relevant source + harness&lt;br&gt;
feedback, calls the model, applies whatever &lt;code&gt;### FILE:&lt;/code&gt; blocks came back,&lt;br&gt;
commits, and runs the test suite. If tests fail, the failure output goes&lt;br&gt;
back into the next prompt as &lt;em&gt;ground truth&lt;/em&gt; — "the output below is ground&lt;br&gt;
truth," not a suggestion. The model iterates until tests pass or the&lt;br&gt;
budget runs out. If the model proposes no files at all, the harness&lt;br&gt;
doesn't crash — it nudges: propose a test or doc change instead, or&lt;br&gt;
explain why.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5. The reviewer gate&lt;/strong&gt; (&lt;code&gt;run_reviewer&lt;/code&gt;). Takes the diff between &lt;code&gt;main&lt;/code&gt;&lt;br&gt;
and the loop branch and asks a &lt;em&gt;separate&lt;/em&gt; model for a verdict, parsed with&lt;br&gt;
a strict regex: &lt;code&gt;VERDICT: APPROVE&lt;/code&gt; or &lt;code&gt;REJECT&lt;/code&gt;. On APPROVE, done. On&lt;br&gt;
REJECT (or an unparseable verdict — untrusted output, remember), the&lt;br&gt;
critique is fed back to the implementer for a bounded number of fix&lt;br&gt;
rounds, tests re-run, and the reviewer re-examines the new diff. Up to&lt;br&gt;
&lt;code&gt;max_review_rounds&lt;/code&gt; rounds.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;6. The main cycle&lt;/strong&gt; (&lt;code&gt;main&lt;/code&gt;). The five acts:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Pick a ticket (highest-priority &lt;code&gt;open&lt;/code&gt;, or &lt;code&gt;--ticket TICKET-003&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;Create a fresh worktree in &lt;code&gt;worktrees/&amp;lt;TICKET-ID&amp;gt;/&lt;/code&gt; plus a
&lt;code&gt;loop/&amp;lt;ticket&amp;gt;&lt;/code&gt; branch — clearing any stale state from a crashed run
first.&lt;/li&gt;
&lt;li&gt;Run the implementer. On failure: ticket → &lt;code&gt;blocked&lt;/code&gt;, cycle logged,
Discord notified 🔴.&lt;/li&gt;
&lt;li&gt;Run the reviewer. On rejection: ticket → &lt;code&gt;needs-human&lt;/code&gt;, branch &lt;em&gt;kept&lt;/em&gt;
for inspection, notified 🟡.&lt;/li&gt;
&lt;li&gt;Show the diff stat, ask the human to merge (unless &lt;code&gt;--yes&lt;/code&gt;), merge
with &lt;code&gt;--no-ff&lt;/code&gt;, ticket → &lt;code&gt;done&lt;/code&gt;, log, notify 🟢.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Two scars in this file are worth naming, because they're the kind of thing&lt;br&gt;
you only learn by running it. First: the merge now does an explicit&lt;br&gt;
&lt;code&gt;git checkout main_branch&lt;/code&gt; before merging — never assume what the repo has&lt;br&gt;
checked out. Second: the worktree is removed &lt;em&gt;before&lt;/em&gt; the branch is&lt;br&gt;
deleted, because a &lt;code&gt;loop/*&lt;/code&gt; branch is always checked out by its worktree&lt;br&gt;
and &lt;code&gt;git branch -d&lt;/code&gt; refuses otherwise. The repo's own contributor guide&lt;br&gt;
says it best: &lt;em&gt;"This was a real bug once; don't reintroduce it."&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Two things worth noticing. First, &lt;strong&gt;failure states are first-class&lt;br&gt;
citizens&lt;/strong&gt;: &lt;code&gt;blocked&lt;/code&gt;, &lt;code&gt;needs-human&lt;/code&gt;, &lt;code&gt;merge-declined&lt;/code&gt; — the loop never&lt;br&gt;
silently dies and never pretends. It always lands somewhere legible, with&lt;br&gt;
the branch preserved when a human might want to look. Second, the &lt;code&gt;finally&lt;/code&gt;&lt;br&gt;
block: the worktree is &lt;em&gt;always&lt;/em&gt; cleaned up, success or crash. The sandbox&lt;br&gt;
is disposable by construction, not by convention.&lt;/p&gt;

&lt;p&gt;The deepest idea in this file isn't any single function — it's the&lt;br&gt;
inversion of the usual relationship. Normally the human drives and the AI&lt;br&gt;
assists. Here the &lt;em&gt;harness&lt;/em&gt; drives and the AI &lt;em&gt;proposes&lt;/em&gt;. The code that&lt;br&gt;
touches reality (git, filesystem, tests, merge) is all deterministic&lt;br&gt;
Python. The model never touches reality at all. That's what makes the&lt;br&gt;
loop safe to run unattended: the blast radius of a misbehaving model is&lt;br&gt;
exactly one throwaway worktree.&lt;/p&gt;

&lt;h3&gt;
  
  
  Aside — the brain and the body
&lt;/h3&gt;

&lt;p&gt;My analogy: the harness is the body, the LLM is the brain. It holds —&lt;br&gt;
with one twist that matters. A person can open their own eyes to check what&lt;br&gt;
they imagined. The implementer can't. Its only senses are what the body&lt;br&gt;
feeds it: the file list, the area sources, the test output. The body&lt;br&gt;
decides what the brain gets to perceive. That's the cage &lt;em&gt;and&lt;/em&gt; the design.&lt;/p&gt;

&lt;p&gt;But the imagination parallel is real: when the implementer proposes a&lt;br&gt;
&lt;code&gt;### FILE:&lt;/code&gt; block, that &lt;em&gt;is&lt;/em&gt; its imagination — code shaped from its&lt;br&gt;
internal model of how things should look, not from anything it directly&lt;br&gt;
observed. And the test loop is what keeps imagination honest. Imagine →&lt;br&gt;
propose → the harness checks the imagination against reality (the test&lt;br&gt;
suite as shared ground truth) → feed back → imagine better. A person does&lt;br&gt;
the same thing when they picture a fix, try it, and watch it fail. The&lt;br&gt;
loop just makes the cycle explicit and unskippable.&lt;/p&gt;

&lt;h2&gt;
  
  
  Part 3 — The roles (scout, implementer, reviewer, planner)
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;`prompts/&lt;/em&gt;.md` — four job descriptions for the same brain.*&lt;/p&gt;

&lt;p&gt;Here's the thing most people miss about multi-agent setups: there aren't&lt;br&gt;
four different AIs in this loop. There's one kind of mind — an LLM —&lt;br&gt;
wearing four different masks. The prompts are what decide what each mask&lt;br&gt;
is allowed to &lt;em&gt;think&lt;/em&gt;. Same brain, four cognitive modes. If the&lt;br&gt;
orchestrator is the body, the prompts are the job descriptions handed to&lt;br&gt;
the brain each morning.&lt;/p&gt;

&lt;p&gt;Every prompt follows the same anatomy: &lt;strong&gt;identity&lt;/strong&gt; (who you are right&lt;br&gt;
now), &lt;strong&gt;inputs&lt;/strong&gt; (what the body will show you), &lt;strong&gt;output format&lt;/strong&gt; (strict&lt;br&gt;
— because the harness is a &lt;em&gt;parser&lt;/em&gt;, not a reader), and &lt;strong&gt;rules&lt;/strong&gt; (what's&lt;br&gt;
out of scope). The strictness isn't bureaucracy. The orchestrator extracts&lt;br&gt;
&lt;code&gt;### FILE:&lt;/code&gt; blocks with a regex and hunts for &lt;code&gt;VERDICT: APPROVE&lt;/code&gt; with&lt;br&gt;
another one. If a role gets chatty or creative with its format, the&lt;br&gt;
machine can't read it. The roles don't converse with the harness — they&lt;br&gt;
produce machine-readable artifacts.&lt;/p&gt;

&lt;p&gt;And every prompt opens the same way: &lt;em&gt;"You are the Scout in a recursive&lt;br&gt;
development loop for {{PROJECT_NAME}}. {{PROJECT_DESCRIPTION}}"&lt;/em&gt; — the&lt;br&gt;
placeholders filled by &lt;code&gt;harness/prompts.py&lt;/code&gt; from your config. The roles&lt;br&gt;
don't know whose repo they're working on until you tell the config. That's&lt;br&gt;
the whole portability story in two template variables.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Scout&lt;/strong&gt; — the eyes. It receives the file list, recent commits, and the&lt;br&gt;
backlog, and its only job is finding bite-sized work: bugs, missing&lt;br&gt;
tests, tech debt, security smells, docs gaps. It never writes code. Its&lt;br&gt;
output is &lt;em&gt;tickets&lt;/em&gt;, in a strict schema with acceptance criteria. The&lt;br&gt;
rules are doing quiet management work: one ticket must fit in one cycle&lt;br&gt;
(~30–45 min), never duplicate the backlog, prefer testable criteria, no&lt;br&gt;
grand redesigns, 3–8 tickets per run. "Quality over quantity" — written&lt;br&gt;
into the prompt because the alternative is a scout that floods the backlog&lt;br&gt;
with 40 vague dreams.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Implementer&lt;/strong&gt; — the hands. Receives the ticket, the planner's plan (if&lt;br&gt;
any), relevant file contents, and test feedback. Changes code &lt;em&gt;only&lt;/em&gt; by&lt;br&gt;
proposing complete files in &lt;code&gt;### FILE:&lt;/code&gt; blocks — never diffs, never&lt;br&gt;
snippets (diffs are where models hallucinate context lines; complete files&lt;br&gt;
are verifiable). It's explicitly told what it &lt;em&gt;cannot&lt;/em&gt; do: no commands,&lt;br&gt;
no packages, no network — the harness runs the tests. The rules read like&lt;br&gt;
a senior dev's code-review checklist: stay inside the acceptance criteria,&lt;br&gt;
never break existing behavior to satisfy a new one, no secrets in code,&lt;br&gt;
match existing style, and "when tests fail, the failure output is ground&lt;br&gt;
truth: fix the cause, not the symptom." And one humane rule: &lt;em&gt;if the&lt;br&gt;
ticket is already satisfied, say so and propose no files.&lt;/em&gt; The implementer&lt;br&gt;
is allowed to declare victory — or rather, allowed to notice there was&lt;br&gt;
never a battle.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Reviewer&lt;/strong&gt; — the judge. Receives the ticket and the diff; last line of&lt;br&gt;
defense before &lt;code&gt;main&lt;/code&gt;. Four judging criteria: acceptance criteria met, no&lt;br&gt;
regressions (check the callers), no security issues, no scope creep.&lt;br&gt;
Output: concise analysis, then exactly one line — &lt;code&gt;VERDICT: APPROVE&lt;/code&gt; or&lt;br&gt;
&lt;code&gt;VERDICT: REJECT&lt;/code&gt;. And then the most interesting sentence in the entire&lt;br&gt;
prompt set:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;Vague rejections get ignored and the change ships anyway on the next&lt;br&gt;
round — so be specific.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That's incentive design smuggled into a prompt. It makes the reviewer&lt;br&gt;
&lt;em&gt;accountable&lt;/em&gt;: a lazy REJECT with no actionable bullets is worse than&lt;br&gt;
useless, because the loop routes around it. The prompt doesn't just tell&lt;br&gt;
the reviewer what to do — it tells it what happens if it does the job&lt;br&gt;
badly. Critique without specifics is noise, and the system is designed to&lt;br&gt;
treat it as such.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Planner&lt;/strong&gt; — the architect. Invoked before the implementer on tickets&lt;br&gt;
tagged &lt;code&gt;type: feature&lt;/code&gt; (wired via &lt;code&gt;planner.enabled_for_feature_tickets&lt;/code&gt; in&lt;br&gt;
config). Produces goal, files, ordered steps (each small enough to keep&lt;br&gt;
tests green), verification, and risks — which the orchestrator prepends to&lt;br&gt;
the implementer's context. And here's the honest bit, enforced in code&lt;br&gt;
rather than prose: the planner is &lt;em&gt;advisory&lt;/em&gt;. &lt;code&gt;get_plan()&lt;/code&gt; can never kill&lt;br&gt;
a cycle — on any failure it degrades to "no plan" and the loop proceeds.&lt;br&gt;
A planning step that could take down the whole run would be a liability,&lt;br&gt;
not a feature.&lt;/p&gt;

&lt;p&gt;The pattern across all four: &lt;strong&gt;each role is deliberately dumber than the&lt;br&gt;
whole system.&lt;/strong&gt; The scout can't code. The implementer can't judge. The&lt;br&gt;
reviewer can't fix (it can only reject with specifics, and the&lt;br&gt;
&lt;em&gt;implementer&lt;/em&gt; fixes). The planner can't execute. Nobody has the full&lt;br&gt;
picture except the harness — and the harness can't think. Intelligence&lt;br&gt;
lives in the &lt;em&gt;composition&lt;/em&gt;, not in any single call. That's the oldest&lt;br&gt;
trick in systems design, and it works on minds too: a team of narrow&lt;br&gt;
specialists, coordinated by a dumb-but-reliable process, outperforms one&lt;br&gt;
brilliant generalist you can't trust.&lt;/p&gt;




&lt;h2&gt;
  
  
  Part 4 — The safety model (why the cage matters)
&lt;/h2&gt;

&lt;p&gt;The README says it best: &lt;em&gt;"Safety rules (enforced by the harness, not by&lt;br&gt;
hope)."&lt;/em&gt; Every layer below is deterministic Python, not a polite request&lt;br&gt;
in a prompt. Prompts can be ignored, misunderstood, or jailbroken. Code&lt;br&gt;
can't be argued with. The cage has six layers, and they're ordered by a&lt;br&gt;
single principle: &lt;strong&gt;the more irreversible the action, the more gates guard&lt;br&gt;
it.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Layer 1 — The filesystem cage.&lt;/strong&gt; The agent never sees your repo. It sees&lt;br&gt;
a disposable &lt;code&gt;git worktree&lt;/code&gt; on a throwaway &lt;code&gt;loop/*&lt;/code&gt; branch. &lt;code&gt;main&lt;/code&gt; is&lt;br&gt;
touched by exactly one operation — a &lt;code&gt;--no-ff&lt;/code&gt; merge — and only after&lt;br&gt;
every gate passes. On top of that, every path the model proposes goes&lt;br&gt;
through two validators: &lt;code&gt;safe_relpath()&lt;/code&gt; rejects absolute paths and &lt;code&gt;..&lt;/code&gt;&lt;br&gt;
escapes (the output can't wander out of the worktree), and &lt;code&gt;is_protected()&lt;/code&gt;&lt;br&gt;
blocks writes to &lt;code&gt;.env&lt;/code&gt;, &lt;code&gt;secrets/&lt;/code&gt;, and anything else you list. The&lt;br&gt;
model's pen can only move inside the sandbox you drew.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Layer 2 — The prompt cage.&lt;/strong&gt; Secrets never enter prompts. The workers see&lt;br&gt;
the ticket and the code — nothing else. No API keys, no tokens, no&lt;br&gt;
personal data in context, which means none can leak into output. And the&lt;br&gt;
harness itself lives &lt;em&gt;outside&lt;/em&gt; the product repo: the implementer literally&lt;br&gt;
cannot reach the files that control it. The loop controller is not&lt;br&gt;
self-modifying, not by policy — by directory structure. The config file&lt;br&gt;
holds no secrets either; the OpenRouter key lives in an environment&lt;br&gt;
variable, never in YAML.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Layer 3 — The budget cage.&lt;/strong&gt; Everything that could run forever is&lt;br&gt;
bounded: 12 implementer iterations per ticket, 2 review rounds, 3 fix&lt;br&gt;
iterations per round, 600 seconds per test run, 30,000 chars of diff for&lt;br&gt;
the reviewer. A confused model can't burn your GPU all night or drown the&lt;br&gt;
reviewer in a million-line diff. Bounded compute, bounded time, bounded&lt;br&gt;
context. When a budget exhausts, the ticket lands in a legible failure&lt;br&gt;
state — &lt;code&gt;blocked&lt;/code&gt; — instead of spinning.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Layer 4 — The fitness cage.&lt;/strong&gt; The test suite plus the reviewer model.&lt;br&gt;
&lt;em&gt;No measurement, no merge.&lt;/em&gt; This is the only part of the system allowed to&lt;br&gt;
say "good enough," and it's the part with no imagination at all — exit&lt;br&gt;
code 0 and a regex-matched APPROVE. Notice what this means: the loop's&lt;br&gt;
definition of quality is entirely external to the models. They don't get&lt;br&gt;
to decide what "done" means. You do, via the tests you wrote and the&lt;br&gt;
ticket criteria you approved.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Layer 5 — The human cage.&lt;/strong&gt; &lt;code&gt;require_human_approval_for_merge&lt;/code&gt; defaults&lt;br&gt;
to true. The &lt;code&gt;--yes&lt;/code&gt; flag exists, but you have to reach for it&lt;br&gt;
deliberately — it's an act of earned trust, not a default. This is&lt;br&gt;
graduated autonomy done right: the system is useful on day one (you're the&lt;br&gt;
merge button) and &lt;em&gt;capable&lt;/em&gt; of full autonomy later, but the step across&lt;br&gt;
that line is yours to take, consciously.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Layer 6 — Failure hygiene.&lt;/strong&gt; The &lt;code&gt;finally&lt;/code&gt; block always torches the&lt;br&gt;
worktree — success, failure, or crash. Notifications are wrapped so a dead&lt;br&gt;
Discord webhook can never break a cycle (&lt;em&gt;observability must not be&lt;br&gt;
load-bearing&lt;/em&gt;). And every failure lands somewhere legible: &lt;code&gt;blocked&lt;/code&gt;,&lt;br&gt;
&lt;code&gt;needs-human&lt;/code&gt;, &lt;code&gt;merge-declined&lt;/code&gt; — with the branch preserved whenever a&lt;br&gt;
human might want to inspect the wreckage. The loop never silently dies and&lt;br&gt;
never pretends.&lt;/p&gt;

&lt;p&gt;Step back and the shape is clear: it's &lt;strong&gt;defense in depth for agency&lt;/strong&gt;.&lt;br&gt;
Proposing text is unguarded (harmless). Writing files is path-validated&lt;br&gt;
(reversible — it's a throwaway branch). Running tests is timeout-bounded&lt;br&gt;
(expensive but contained). Merging to &lt;code&gt;main&lt;/code&gt; is guarded by tests, a&lt;br&gt;
reviewer, and a human (irreversible — or at least, the one action you'd&lt;br&gt;
regret). Each layer assumes the one inside it has already failed. That's&lt;br&gt;
the mindset: not "how do we keep the model well-behaved," but "what's the&lt;br&gt;
blast radius when it isn't" — and the answer, at every layer, is &lt;em&gt;small&lt;br&gt;
and legible&lt;/em&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Part 5 — Making it yours (config, models, scheduling)
&lt;/h2&gt;

&lt;p&gt;Everything the loop needs to know about &lt;em&gt;your machine&lt;/em&gt; and &lt;em&gt;your project&lt;/em&gt;&lt;br&gt;
lives in one file: &lt;code&gt;harness/config.yaml&lt;/code&gt;. That's a deliberate design&lt;br&gt;
choice — the harness is portable, the config is personal. Six sections:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;project&lt;/code&gt;.&lt;/strong&gt; &lt;code&gt;name&lt;/code&gt; and &lt;code&gt;description&lt;/code&gt; — a short name and one or two&lt;br&gt;
sentences about what the project does, its stack, who it's for. This is&lt;br&gt;
what fills &lt;code&gt;{{PROJECT_NAME}}&lt;/code&gt; and &lt;code&gt;{{PROJECT_DESCRIPTION}}&lt;/code&gt; in every role&lt;br&gt;
prompt. The agents meet your repo through these two fields, so write them&lt;br&gt;
like an introduction, not a label.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;product_repo&lt;/code&gt; / &lt;code&gt;main_branch&lt;/code&gt;.&lt;/strong&gt; The full path to the git repo the loop&lt;br&gt;
improves (forward slashes work on Windows too), and your default branch.&lt;br&gt;
Any repo with a test suite — not just mine, not just one stack. If&lt;br&gt;
&lt;code&gt;product_repo&lt;/code&gt; isn't a git repo, the orchestrator refuses to run. Fail&lt;br&gt;
fast on bad config, not halfway through a cycle.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;models&lt;/code&gt;.&lt;/strong&gt; Ollama-first, OpenRouter fallback. The defaults assume&lt;br&gt;
~30B-class models (&lt;code&gt;qwen3:30b&lt;/code&gt;, &lt;code&gt;qwen3-coder:30b&lt;/code&gt; for the implementer) —&lt;br&gt;
&lt;em&gt;you will change these.&lt;/em&gt; Run &lt;code&gt;ollama list&lt;/code&gt; and put in the names you&lt;br&gt;
actually have pulled, character for character. Three things worth noticing&lt;br&gt;
here. First, each role gets its own model: the scout, implementer,&lt;br&gt;
reviewer, and planner don't have to be the same actor. A small fast model&lt;br&gt;
can scout; a coder-tuned model implements; a careful one reviews. Same&lt;br&gt;
script, different performances — cast on purpose. Second, the cloud&lt;br&gt;
fallback only fires when the local call &lt;em&gt;fails&lt;/em&gt; — it's a safety net, not&lt;br&gt;
a default. Your data stays on your machine unless your machine can't do&lt;br&gt;
the job. Third, model sizing is a real operational lesson, not a footnote:&lt;br&gt;
on my own smoke test, the reviewer timed out after 300 seconds on a 9B&lt;br&gt;
model, so the reviewer role got overridden to a 4B model for that run.&lt;br&gt;
Judgment turns out to be cheaper than generation — size accordingly, and&lt;br&gt;
don't be proud about it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;planner&lt;/code&gt;.&lt;/strong&gt; One flag: &lt;code&gt;enabled_for_feature_tickets&lt;/code&gt;. Leave it on — a&lt;br&gt;
planning pass before the implementer on big tickets, advisory and&lt;br&gt;
never fatal.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;budgets&lt;/code&gt;.&lt;/strong&gt; The numbers from Part 4 — iterations, rounds, timeouts,&lt;br&gt;
diff caps. Tune these once you've watched a few cycles. A cheap trick:&lt;br&gt;
start tight (low budgets), because a tight budget that keeps blocking&lt;br&gt;
tells you your tickets are too big — which is useful information about&lt;br&gt;
your backlog, not just your loop.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;gates&lt;/code&gt;.&lt;/strong&gt; &lt;code&gt;test_command&lt;/code&gt; (change &lt;code&gt;pytest -q&lt;/code&gt; to whatever your repo&lt;br&gt;
uses — &lt;code&gt;npm test&lt;/code&gt;, &lt;code&gt;go test ./...&lt;/code&gt;, &lt;code&gt;dotnet test&lt;/code&gt;; it must exit 0 on&lt;br&gt;
success), &lt;code&gt;require_human_approval_for_merge&lt;/code&gt; (leave &lt;code&gt;true&lt;/code&gt; until the loop&lt;br&gt;
has earned it), &lt;code&gt;protected_paths&lt;/code&gt; (add anything the implementer should&lt;br&gt;
never touch). This section is where your &lt;em&gt;judgment&lt;/em&gt; lives in the system.&lt;br&gt;
Everything else is machinery; the gates are values.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;notifications&lt;/code&gt;.&lt;/strong&gt; Optional Discord webhook. Cycle results — 🟢 merged,&lt;br&gt;
🟡 needs-human, 🔴 blocked — posted to a channel. Waking up to a loop&lt;br&gt;
report is the "let go" part made tangible: the system worked while you&lt;br&gt;
slept, and all you have to do is read the log.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Running it.&lt;/strong&gt; The quickstart is five steps: clone the repo &lt;em&gt;next to&lt;/em&gt;&lt;br&gt;
(not inside) your product repo, run the bootstrap (&lt;code&gt;setup/setup.ps1&lt;/code&gt; on&lt;br&gt;
Windows, &lt;code&gt;setup/setup.sh&lt;/code&gt; on macOS/Linux — it checks Python, git, and&lt;br&gt;
Ollama and installs the two Python deps), edit the config, optionally set&lt;br&gt;
&lt;code&gt;OPENROUTER_API_KEY&lt;/code&gt;, and smoke-test with &lt;code&gt;run_cycle.bat&lt;/code&gt; — which&lt;br&gt;
processes &lt;code&gt;TICKET-001&lt;/code&gt;, a docs-only ticket that can't break anything.&lt;br&gt;
There's also &lt;code&gt;setup/agent-setup.md&lt;/code&gt;: the same procedure written as a&lt;br&gt;
deterministic, checkable script for an assisting AI agent, with&lt;br&gt;
&lt;code&gt;setup/manifest.json&lt;/code&gt; carrying the machine-readable version. If you have&lt;br&gt;
an AI helping you install this, point it there first. Watch one full cycle&lt;br&gt;
complete. Then daily use is two commands: &lt;code&gt;run_scout.*&lt;/code&gt; to propose&lt;br&gt;
tickets, &lt;code&gt;run_cycle.*&lt;/code&gt; to work them. When the gates have proven out,&lt;br&gt;
Windows Task Scheduler (or cron, or systemd, or launchd) runs it nightly&lt;br&gt;
and you graduate from operator to reviewer.&lt;/p&gt;

&lt;h3&gt;
  
  
  The honest test report
&lt;/h3&gt;

&lt;p&gt;I'd rather tell you what's actually been proven than sell you a story.&lt;br&gt;
The full pipeline — ticket → worktree → implement → test → review →&lt;br&gt;
merge → branch deletion → worktree cleanup — passes end to end against a&lt;br&gt;
disposable repo with a stubbed LLM. Every gate, every cleanup step,&lt;br&gt;
verified.&lt;/p&gt;

&lt;p&gt;The live-model story has real entries now, including an ugly one. On&lt;br&gt;
2026-09-28, TICKET-001 — the docs-only smoke test — ran against my&lt;br&gt;
Windows machine with local models via Ollama and &lt;code&gt;pytest -q&lt;/code&gt; as the&lt;br&gt;
gate. The first attempt &lt;em&gt;failed honestly&lt;/em&gt;: the reviewer timed out after&lt;br&gt;
300 seconds, so no verdict was ever cast and nothing merged. (My local&lt;br&gt;
agent's first report claimed a full successful cycle. It wasn't one.&lt;br&gt;
The log showed a timeout; the merge never happened. I keep that in&lt;br&gt;
because the loop is supposed to make exactly this kind of failure&lt;br&gt;
legible — and the first thing it made legible was an overconfident&lt;br&gt;
report from my own setup.) The second attempt surfaced the&lt;br&gt;
&lt;code&gt;__pycache__&lt;/code&gt; bug from Part 2. The third attempt, with the fix wired&lt;br&gt;
in, passed clean: round 1, &lt;code&gt;[local:qwen3.5:4b]&lt;/code&gt; APPROVE, ticket →&lt;br&gt;
done, change merged to &lt;code&gt;main&lt;/code&gt;, no &lt;code&gt;loop/*&lt;/code&gt; branches or worktrees left&lt;br&gt;
behind. Full cycle, live models, proven.&lt;/p&gt;

&lt;p&gt;So: the loop &lt;em&gt;works&lt;/em&gt; on live models. What's still pending is breadth —&lt;br&gt;
more tickets, harder tickets — and the TICKET-004 story in Part 6,&lt;br&gt;
which is the more important lesson anyway.&lt;/p&gt;

&lt;p&gt;Known rough edges, since we're being honest: the&lt;br&gt;
&lt;code&gt;_strip_test_artifacts&lt;/code&gt; fix still needs its own commit and push to the&lt;br&gt;
public repo (&lt;code&gt;harness/config.yaml&lt;/code&gt;, &lt;code&gt;tickets/TICKET-001.md&lt;/code&gt;, and&lt;br&gt;
&lt;code&gt;CYCLE_LOG.md&lt;/code&gt; stay local); the LICENSE still says "Daniel" instead&lt;br&gt;
of my full name; and there's a stale duplicate &lt;code&gt;reviewer:&lt;/code&gt; key to&lt;br&gt;
delete. Small things. They'll get fixed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Upgrade paths&lt;/strong&gt; (documented, not built — honest labels again): swap the&lt;br&gt;
file-block implementer for ACP calls into your editor so the agent edits&lt;br&gt;
like you do; point the roles at your own agent once it can use tools. The&lt;br&gt;
harness shape — tickets → worktree → tests → review → merge — doesn't&lt;br&gt;
change. Only the actors get better.&lt;/p&gt;




&lt;h2&gt;
  
  
  Part 6 — Field notes: what live runs actually taught us
&lt;/h2&gt;

&lt;p&gt;The write-up above describes the loop as designed. This part describes&lt;br&gt;
the loop as &lt;em&gt;survived&lt;/em&gt; — what happened when it ran real tickets with&lt;br&gt;
real models, and what those runs taught me that the design doc never&lt;br&gt;
could.&lt;/p&gt;

&lt;h3&gt;
  
  
  The reviewer is a heuristic, not a judge
&lt;/h3&gt;

&lt;p&gt;TICKET-004 was supposed to be simple: rewrite the install guide. The&lt;br&gt;
implementer — a 4B model, the same size that had cleanly reviewed&lt;br&gt;
TICKET-001 — produced a 220-line rewrite that gutted INSTALL.md from&lt;br&gt;
~256 lines down to 36: a destructive rewrite of the repo's most&lt;br&gt;
important doc. And the reviewer &lt;strong&gt;approved it&lt;/strong&gt;. The destructive change&lt;br&gt;
merged into my local main before I caught it. (The public GitHub repo&lt;br&gt;
was untouched — &lt;code&gt;origin/main&lt;/code&gt; was still at &lt;code&gt;1c6752e&lt;/code&gt; — which is the one&lt;br&gt;
thing that saved the day: local damage, public record intact.)&lt;/p&gt;

&lt;p&gt;Read that again, because it's the most important sentence in this&lt;br&gt;
write-up: &lt;strong&gt;the reviewer approved a destructive change.&lt;/strong&gt; The safety&lt;br&gt;
layer I described in Part 4 as "the last line of defense before &lt;code&gt;main&lt;/code&gt;"&lt;br&gt;
let a wrecking ball through.&lt;/p&gt;

&lt;p&gt;So here's the honest correction to Part 4: the reviewer gate is not a&lt;br&gt;
judge. It's a heuristic — a cheap second opinion that catches &lt;em&gt;some&lt;/em&gt;&lt;br&gt;
bad changes, not a guarantor. The real safety was always elsewhere: the&lt;br&gt;
human approval gate (which was on, and which is why I caught it), the&lt;br&gt;
separation of local damage from the public repo, and — the fix I added&lt;br&gt;
after — a deterministic guard the model can't argue with.&lt;/p&gt;

&lt;h3&gt;
  
  
  The guard that can't be argued with
&lt;/h3&gt;

&lt;p&gt;After TICKET-004, I added &lt;code&gt;max_deletions_per_iteration: 5&lt;/code&gt; to the&lt;br&gt;
harness config: any single implementer iteration that deletes more than&lt;br&gt;
5 lines gets blocked before anything is committed. It's enforced in&lt;br&gt;
deterministic Python, not in a prompt. Prompts can be ignored;&lt;br&gt;
&lt;code&gt;if deletions &amp;gt; 5: block&lt;/code&gt; can't.&lt;/p&gt;

&lt;p&gt;The guard proved itself immediately. A re-run of TICKET-004 on&lt;br&gt;
2026-10-02 hit it on first contact: the implementer proposed a&lt;br&gt;
destructive rewrite, the guard blocked it, the reviewer rejected what&lt;br&gt;
was left, and nothing merged. No damage. That's the lesson in one&lt;br&gt;
sentence: &lt;strong&gt;when a model proves unreliable, replace the model's&lt;br&gt;
judgment with a measurement.&lt;/strong&gt; The reviewer is a model; the guard is a&lt;br&gt;
ruler. Rulers don't have bad days.&lt;/p&gt;

&lt;h3&gt;
  
  
  APPEND, not rewrite: designing around the model's actual limits
&lt;/h3&gt;

&lt;p&gt;The deeper problem TICKET-004 exposed wasn't just a bad review — it was&lt;br&gt;
a capability gap. The ticket asked a 9B model (with thinking disabled, a&lt;br&gt;
config choice made for speed) to rewrite INSTALL.md, and it couldn't&lt;br&gt;
reproduce ~256 lines of careful documentation verbatim. It summarized,&lt;br&gt;
it compressed, it destroyed. That's not a prompt problem; it's what the&lt;br&gt;
model &lt;em&gt;can&lt;/em&gt; do. My assessment: a 9B-class model with thinking off&lt;br&gt;
cannot reliably reproduce long documents verbatim. No prompt rewrite&lt;br&gt;
fixes that. You don't ask a tool to do what it can't do — you change&lt;br&gt;
the task.&lt;/p&gt;

&lt;p&gt;So the long-term fix is structural, not rhetorical: &lt;code&gt;### APPEND:&lt;/code&gt;&lt;br&gt;
blocks. Instead of the implementer rewriting whole files, it appends&lt;br&gt;
surgical additions — with newline hygiene, &lt;code&gt;safe_relpath&lt;/code&gt; and&lt;br&gt;
&lt;code&gt;is_protected&lt;/code&gt; checks, and a prompt rule preferring APPEND over&lt;br&gt;
&lt;code&gt;### FILE:&lt;/code&gt; rewrites. &lt;code&gt;APPEND_BLOCK&lt;/code&gt; is already live in&lt;br&gt;
&lt;code&gt;harness/orchestrator.py&lt;/code&gt;; the prompt rule in &lt;code&gt;prompts/implementer.md&lt;/code&gt;&lt;br&gt;
and the ticket-body rewrite are still pending. The direction is clear&lt;br&gt;
though: &lt;strong&gt;make the unit of change small enough that the model's&lt;br&gt;
weakness doesn't matter.&lt;/strong&gt; Small diffs, deterministic guards, human at&lt;br&gt;
the merge button.&lt;/p&gt;

&lt;h3&gt;
  
  
  Windows is a real platform, not an edge case
&lt;/h3&gt;

&lt;p&gt;Running this on Windows surfaced a whole class of bugs the design never&lt;br&gt;
imagined: the test gate &lt;code&gt;python -m py_compile harness/*.py&lt;/code&gt; never&lt;br&gt;
globs on Windows (it's now &lt;code&gt;python -m compileall harness/&lt;/code&gt;); &lt;code&gt;sh()&lt;/code&gt;&lt;br&gt;
needed &lt;code&gt;encoding="utf-8", errors="replace"&lt;/code&gt; pinned; the config needed&lt;br&gt;
&lt;code&gt;DEFAULT_CTX: 8192&lt;/code&gt; and &lt;code&gt;think: False&lt;/code&gt; to behave. My local agent&lt;br&gt;
documented all of it in an &lt;code&gt;ouroboros-cleanup&lt;/code&gt; skill so the next&lt;br&gt;
machine doesn't have to rediscover it. If you're building agent&lt;br&gt;
harnesses, test on the platform your users actually run — the shell is&lt;br&gt;
part of the system, and the system includes Windows.&lt;/p&gt;

&lt;h3&gt;
  
  
  The loop is now eating its own cooking
&lt;/h3&gt;

&lt;p&gt;As of late September, the loop on my Windows machine is pointed at&lt;br&gt;
itself: &lt;code&gt;product_repo&lt;/code&gt; is the Ouroboros repo, &lt;code&gt;require_human_approval_for_merge&lt;/code&gt;&lt;br&gt;
is true, and the scout has proposed 5 tickets sitting in the backlog.&lt;br&gt;
Dogfooding is the only honest test of a dev tool — and after TICKET-004,&lt;br&gt;
I'm in no hurry to turn off the human gate.&lt;/p&gt;

&lt;h3&gt;
  
  
  The current open list
&lt;/h3&gt;

&lt;p&gt;Since this write-up is supposed to be honest, here's what's still&lt;br&gt;
unresolved as I write this: restore the local INSTALL.md from&lt;br&gt;
&lt;code&gt;origin/main&lt;/code&gt; (it's 36 lines; it should be ~256); wire the APPEND&lt;br&gt;
prompt rule into &lt;code&gt;prompts/implementer.md&lt;/code&gt;; rewrite the TICKET-004 body&lt;br&gt;
to use &lt;code&gt;### APPEND:&lt;/code&gt;; commit and push the &lt;code&gt;_strip_test_artifacts&lt;/code&gt; fix&lt;br&gt;
as its own commit; fix the LICENSE name; delete the stale duplicate&lt;br&gt;
&lt;code&gt;reviewer:&lt;/code&gt; key. None of these are design problems. They're just work&lt;br&gt;
— the kind the loop itself will do, one ticket at a time, once I point&lt;br&gt;
it at the backlog.&lt;/p&gt;




&lt;h2&gt;
  
  
  Closing — what this is really about
&lt;/h2&gt;

&lt;p&gt;This harness is a small, runnable instance of a much bigger idea: &lt;strong&gt;guide&lt;br&gt;
and let go.&lt;/strong&gt; You don't control the agent's every move — you set the&lt;br&gt;
fitness function (tests, reviewer, ticket criteria), build the cage&lt;br&gt;
(worktrees, budgets, gates), and then let the system surprise you inside&lt;br&gt;
those bounds. The config file is where "guide" lives. The nightly schedule&lt;br&gt;
is where "let go" lives.&lt;/p&gt;

&lt;p&gt;It's also, not coincidentally, the posture of anyone building something&lt;br&gt;
they want to outgrow them — a mentor, a lead, a founder. You don't&lt;br&gt;
micromanage people into being excellent. You hold the standard, give&lt;br&gt;
them room, and accept that what comes back might not be what you&lt;br&gt;
planned. The loop is that posture, compiled.&lt;/p&gt;

&lt;p&gt;Start small: one smoke-test ticket, one watched cycle. The cage holds.&lt;br&gt;
Then let go a little more each week, and see what it becomes.&lt;/p&gt;

&lt;p&gt;TICKET-004 is the same lesson from the other side: letting go doesn't&lt;br&gt;
mean removing the rails — it means building rails you can actually&lt;br&gt;
trust (deterministic guards, small change units, a human at the merge&lt;br&gt;
button) and then trusting &lt;em&gt;them&lt;/em&gt; instead of the model's judgment. Guide&lt;br&gt;
with the config. Let go with the schedule. And when the model shows you&lt;br&gt;
what it can't do, believe it the first time.&lt;/p&gt;




&lt;h2&gt;
  
  
  Appendix — the repo
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Code:&lt;/strong&gt; &lt;a href="https://github.com/danielKrydynski/ouroboros" rel="noopener noreferrer"&gt;github.com/danielKrydynski/ouroboros&lt;/a&gt;
(v0.2.0) — &lt;code&gt;README.md&lt;/code&gt; for the architecture, &lt;code&gt;INSTALL.md&lt;/code&gt; for the full
operations guide, &lt;code&gt;AGENTS.md&lt;/code&gt; for contributors, &lt;code&gt;setup/agent-setup.md&lt;/code&gt;
if you're bringing an AI assistant to the install.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;This write-up&lt;/strong&gt; is the companion piece: the why and the how, with the
scars left in.&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>python</category>
      <category>opensource</category>
    </item>
  </channel>
</rss>
