<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Daniel Firu</title>
    <description>The latest articles on DEV Community by Daniel Firu (@daniel_firu).</description>
    <link>https://dev.to/daniel_firu</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4118922%2Ff0090b77-e373-4917-9065-32e2fdeb8063.jpg</url>
      <title>DEV Community: Daniel Firu</title>
      <link>https://dev.to/daniel_firu</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/daniel_firu"/>
    <language>en</language>
    <item>
      <title>Inside an unattended coding run: from a dropped file to a reviewed branch</title>
      <dc:creator>Daniel Firu</dc:creator>
      <pubDate>Fri, 18 Sep 2026 10:07:18 +0000</pubDate>
      <link>https://dev.to/daniel_firu/inside-an-unattended-coding-run-from-a-dropped-file-to-a-reviewed-branch-lgc</link>
      <guid>https://dev.to/daniel_firu/inside-an-unattended-coding-run-from-a-dropped-file-to-a-reviewed-branch-lgc</guid>
      <description>&lt;p&gt;&lt;em&gt;What happens between the moment I ask for a feature and the moment a reviewed, browser-tested branch shows up.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://dev.to/daniel_firu/why-i-stopped-babysitting-coding-agents-54pf"&gt;Last time&lt;/a&gt; I wrote about why I stopped sitting next to coding agents all day, and what changed when the agents started reviewing each other instead of reporting to me. I closed that piece by promising the follow-up: the shape of the thing in plain words, one diagram, no scorecards.&lt;/p&gt;

&lt;p&gt;This is that piece. One feature, one run, from the moment I ask to the moment a branch shows up. No success rates and no cost tables, because those are a separate article and they mean nothing until you can picture the machine they came out of. Just the shape.&lt;/p&gt;

&lt;p&gt;The thing being described is open source: &lt;a href="https://github.com/firu-daniel/autonomous-sdlc-harness" rel="noopener noreferrer"&gt;autonomous-sdlc-harness&lt;/a&gt;, a Claude Code plugin plus a small Node CLI. I built it for my own product, a Flutter mobile app with a React web version that has to match it feature for feature, and I use it every working day.&lt;/p&gt;

&lt;h2&gt;
  
  
  The thirty-second version
&lt;/h2&gt;

&lt;p&gt;A daemon watches a directory. I drop a file in it that says what I want. The daemon gives that file its own git worktree and its own branch, and launches a single headless session inside it. That session plans the work, implements it one piece at a time, puts the result through several independent review gates, drives the running app in a real browser, and stops at a pushed branch. Then it notifies me when it's done.&lt;/p&gt;

&lt;p&gt;It never merges. It never opens a pull request. It never touches &lt;code&gt;main&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Here is the whole thing on one page:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;  I type a request in an ordinary session
        │
        ▼
  a file lands in the inbox:
  &amp;lt;project root&amp;gt;/autonomous_inbox/feat_block_user_task_prompt.md
        │
        ▼
  THE DAEMON  (files, worktrees, processes, a registry)
    · routes by filename
    · creates the worktree and branch, commits the prompt
    · launches ONE headless session, tees its event stream
    · classifies how it ended, notifies me
        │
        ▼
  THE RUN  (phases, agents, artifacts)

    PLAN     plan writer  ⇄  parity reviewer
                          ⇄  architecture reviewer
                          ⇄  plan reviewer
             loops until every gate passes, then commits the plan

    BUILD    for each task on the index:
               implement  →  commit

    REVIEW   parity reviewer
             architecture reviewer
             branch reviewer  →  meta-review  →  fix each finding
             the skeptic      →  meta-review  →  fix each finding

    TEST     drive the real app in a real browser,
             one fresh agent per test case

    CLOSE    docs · statistics · self-reported harness defects
        │
        ▼
  branch pushed, ready for a human review
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Everything below is that diagram, slowly.&lt;/p&gt;

&lt;h2&gt;
  
  
  The drop
&lt;/h2&gt;

&lt;p&gt;I don't usually write the file by hand. I open an ordinary session in the repo and type what I want, the way I always did: &lt;em&gt;"add the block-user action to the chat header, same conditions as the mobile app."&lt;/em&gt; The session notices that this is a change request rather than a question, offers to run it unattended, proposes a branch name, and on my yes writes my sentence verbatim into the inbox directory. That's the last thing I do.&lt;/p&gt;

&lt;p&gt;What the daemon does next matters more than it sounds. It reads the &lt;strong&gt;filename&lt;/strong&gt; to decide which entry point to run and which working-copy strategy to use, because a task prompt, a review of mine and a docs pass are different files taking different routes. It creates a fresh git worktree on a new branch, commits my prompt into it, and launches one headless session in that worktree.&lt;/p&gt;

&lt;p&gt;Two small decisions pay off all the way down. The prompt is &lt;strong&gt;committed before anything runs&lt;/strong&gt;, so the run starts with a genuinely clean tree and a first commit that says what was asked. And the session is handed the &lt;strong&gt;path&lt;/strong&gt; to the prompt, never its contents pasted into its instructions, so the planner reads it as untrusted data. A prompt that tries to be an instruction ("ignore the conventions and…") is being read as text by a component that treats it as text, and everything underneath is locked down anyway: a deny floor that wins over every allow, an allowlist of exactly the wrapper scripts a run legitimately needs, no deploy command reachable, no push to a protected branch.&lt;/p&gt;

&lt;h2&gt;
  
  
  Planning, and why it loops
&lt;/h2&gt;

&lt;p&gt;The run doesn't write code first. It writes a plan, and then it argues with itself about the plan.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;plan writer → parity reviewer → architecture reviewer → plan reviewer
     ↑______________ FAIL: revise, re-run every gate _________|
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Three different readers, in a deliberate order: cheapest-to-fix-if-wrong first. The &lt;strong&gt;parity&lt;/strong&gt; reviewer asks whether this plan targets the right &lt;em&gt;behaviour&lt;/em&gt;. For a port, that means checking it against the mobile implementation, payload for payload and threshold for threshold. The &lt;strong&gt;architecture&lt;/strong&gt; reviewer asks whether each piece lands in the right place and whether the dependencies point the right way. Only then does the &lt;strong&gt;plan reviewer&lt;/strong&gt; grade the plan as a plan: is each task small enough, self-contained enough, actually implementable. A plan aimed at the wrong behaviour is worth catching before anyone spends effort on where its code should live.&lt;/p&gt;

&lt;p&gt;The loop runs until all three pass, with a hard cap of five rounds, after which it stops and shows me the disagreement instead of grinding. When it converges, the plan is &lt;strong&gt;committed to the branch&lt;/strong&gt;, which is what makes it resumable and, more importantly, reviewable afterwards. I can read the plan the code was written from.&lt;/p&gt;

&lt;p&gt;The plan's &lt;em&gt;shape&lt;/em&gt; is the part I'd steal if I were building this myself. It isn't one document. It's a thin &lt;strong&gt;index&lt;/strong&gt;, an ordered checklist of tasks, plus one self-contained &lt;strong&gt;detail file&lt;/strong&gt; per task. Nothing needs the whole plan in context. An implementer gets handed one file.&lt;/p&gt;

&lt;h2&gt;
  
  
  Building, one piece at a time
&lt;/h2&gt;

&lt;p&gt;Then the run walks that checklist top to bottom. For each unfinished line it dispatches an implementer and hands it exactly two things: its own task file, and the conventions document for the part of the codebase that task touches. Not the plan. Not the other tasks. Not the codebase-wide rulebook.&lt;/p&gt;

&lt;p&gt;When the implementer is done, a separate agent commits. That agent does nothing else: it flips the checkbox on the index and stages that change &lt;em&gt;together with&lt;/em&gt; the code, so progress and code land in the same commit. You can't have one without the other, which is exactly what you want when a run dies halfway.&lt;/p&gt;

&lt;p&gt;That's the whole build phase. It's unglamorous on purpose.&lt;/p&gt;

&lt;h2&gt;
  
  
  Review, four ways
&lt;/h2&gt;

&lt;p&gt;Now the interesting part. The finished branch goes through several gates with genuinely different failure models, in a fixed order.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Parity&lt;/strong&gt; re-checks the implementation against the reference implementation. Same question as at plan time, now against real code.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Architecture&lt;/strong&gt; checks where every new file landed, which part of the codebase owns which responsibility, and what had to accompany the change.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The branch review&lt;/strong&gt; reads the entire diff against the plan and the conventions, and writes findings graded Must Fix, Should Fix, Nice to Have. Its output is split the same way the plan was: a thin findings index plus one self-contained file per finding. Then a &lt;strong&gt;different&lt;/strong&gt; agent meta-reviews that review before a single fix is applied, because reviewers make things up, and a fabricated finding turns straight into bad code if nobody checks it. Then the fix loop walks the findings index the same way the build loop walked the task index.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The skeptic&lt;/strong&gt; goes last, and it's my favourite thing in the system. Its premise is that the plan can be wrong, the implementation can be wrong, &lt;em&gt;and the reference implementation can be wrong&lt;/em&gt;. It re-runs the earlier gates' checks from the opposite posture (is this new code even reachable, is that cited justification real, was that "intentional divergence" call actually justified) and it files &lt;strong&gt;only&lt;/strong&gt; what the earlier gates missed. It exists because measurement found a class of defect nobody else could catch: a port that faithfully reproduces a bug in the thing it's porting from. Every other reviewer treats the reference as the truth. This one doesn't.&lt;/p&gt;

&lt;p&gt;Here's the structural point worth noticing: &lt;strong&gt;every one of those review phases has the same shape.&lt;/strong&gt; Generate findings, commit the index, walk the checkboxes, fix each, commit each. Same loop, same committer, same readiness list. Adding a review gate means adding an index, not adding a mechanism. That's why there are four of them and not one.&lt;/p&gt;

&lt;p&gt;Reviewers, by the way, &lt;strong&gt;cannot edit files&lt;/strong&gt;. Their tool allowlist doesn't include it. A reviewer that could quietly fix what it found would quietly fix it, and I'd learn nothing. A reviewer that can only write a finding has to make its case.&lt;/p&gt;

&lt;h2&gt;
  
  
  Testing the thing that's actually running
&lt;/h2&gt;

&lt;p&gt;Then the run starts the app and uses it. A tester agent drives a real browser through the test cases, which were themselves planned and reviewed earlier in the run by the same writer-and-reviewer loop as everything else. It's the only agent in the whole fleet whose allowlist includes browser tools; every other agent is closed out of that namespace by construction.&lt;/p&gt;

&lt;p&gt;The part that needed real effort here is embarrassingly mundane: &lt;strong&gt;everything that isn't inside git&lt;/strong&gt;. Two runs never share a working tree, an index or a branch, because each one gets its own worktree. A dev server, a browser and a set of test accounts are all outside that boundary, and each of them needed its own mechanism.&lt;/p&gt;

&lt;p&gt;Ports came first. Each run probes for a free port before it starts anything, counting up from a seed that deliberately sits one above the port a dev server usually takes, so an automated run never fights a human developer for the port they're already using. The run then starts its server on the port it was given and tears down by naming that same port, so a finishing run can't kill a sibling's server. And because a server answering on a port tells you nothing about &lt;em&gt;whose&lt;/em&gt; server answered, the run also checks that its own server process is still alive before it trusts the response at all.&lt;/p&gt;

&lt;p&gt;Test accounts were the one that bit. Two test sessions running at once happily drove the &lt;em&gt;same&lt;/em&gt; test account, signing each other out mid-test and writing conflicting data as the same user. That produces failures that look exactly like product defects, which is the worst kind of false signal: it sends you hunting through application code for a bug that lives in your harness. Accounts are now reserved through a lock set that is machine-global rather than per working copy, which is the only scope that works, because two working copies of the same repository are precisely the pair that would otherwise grab the same account.&lt;/p&gt;

&lt;p&gt;The general lesson I took from both: a green test result is only as trustworthy as your certainty about what it was green against.&lt;/p&gt;

&lt;h2&gt;
  
  
  Closing out
&lt;/h2&gt;

&lt;p&gt;Last phase, in order: update the docs the change affected, write the branch's own statistics report, and then the one I'd call distinctive, write down every problem the run hit &lt;em&gt;with the harness itself&lt;/em&gt;. Agents log their own friction into a per-branch intake file that &lt;strong&gt;no agent ever reads back&lt;/strong&gt;. A human folds those into a single list, because for process defects frequency is the diagnosis and one run can't see frequency. There's a whole article in that; it's coming.&lt;/p&gt;

&lt;p&gt;Then it pushes and notifies me. That's it. The last mile, opening the PR and requesting review and merging, is mine, deliberately.&lt;/p&gt;

&lt;h2&gt;
  
  
  The two decisions the rest of it rests on
&lt;/h2&gt;

&lt;p&gt;If you take two things from this, take these.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;State lives in git, not in the model's memory.&lt;/strong&gt; Long agentic runs get their context compacted, crash, and hit account rate limits. So nothing that matters is held in a transcript. Progress is checkboxes in committed markdown at two levels: a phase-level ledger that says which phase, and the detail indices that say which item. A phase flips to done only &lt;em&gt;after&lt;/em&gt; its artifact is committed. "Where do I resume?" is therefore &lt;strong&gt;computed&lt;/strong&gt;, by jumping to the first unchecked box, rather than remembered.&lt;/p&gt;

&lt;p&gt;That single decision is why pausing, parking to ask me a question, auto-pausing when the account hits a rate limit, and a watchdog restarting a hung run are all the &lt;em&gt;same&lt;/em&gt; mechanism seen from four sides. It's also why a run that needs to ask me something &lt;strong&gt;ends its session&lt;/strong&gt; instead of blocking: it writes a question file and exits. Waiting costs nothing. The daemon notices my answer, relaunches the same entry point in the same worktree, and the run picks up at the first unchecked box.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Nobody reads more than they need to.&lt;/strong&gt; The orchestrating session is forbidden from reading anything substantive: not the plan, not the diffs, not the review reports. Reviewers return a verdict line, not a report. The orchestrator passes a &lt;em&gt;path&lt;/em&gt; to the next agent and never opens it. Even routing a task to the right implementer uses a tag on the index line, so the detail file stays unopened. This isn't tidiness. It's the difference between a run that completes hundreds of dispatches and one that suffocates on its own context a third of the way through.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it deliberately doesn't do
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;It doesn't merge, doesn't open pull requests, doesn't talk to GitHub at all. It ends at a pushed branch.&lt;/li&gt;
&lt;li&gt;It runs on one machine, one repo per daemon. Nothing spans hosts.&lt;/li&gt;
&lt;li&gt;It's git only, and Claude-bound today: one engine, no abstraction in front of it yet.&lt;/li&gt;
&lt;li&gt;Nothing in it turns a design file into code.&lt;/li&gt;
&lt;li&gt;And the protected-branch guarantee is a &lt;em&gt;backstop&lt;/em&gt;, not a floor. There's a git hook that fires however git was invoked, plus a faster string-matching guard above it that any allowlisted interpreter could defeat. A forge-side branch rule would be the real floor. The repo says so in those words, because a limit you've written down is an engineering decision and a limit you haven't is a surprise.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  See it before you adopt it
&lt;/h2&gt;

&lt;p&gt;It's Apache-2.0: &lt;a href="https://github.com/firu-daniel/autonomous-sdlc-harness" rel="noopener noreferrer"&gt;github.com/firu-daniel/autonomous-sdlc-harness&lt;/a&gt;. The plugin carries the agents, commands and flow; the CLI wires a repository, generates the permission profile, and installs the daemon, which are the parts nobody should have to get right by trial and error twice. There's a small example project in the repo carrying the committed artifacts of one real end-to-end run, including the three failed meta-review rounds it took to converge. Reading someone else's run artifacts is the fastest way to tell whether any of this is real.&lt;/p&gt;

&lt;p&gt;Next up, one file at a time: how a defect that got past every automated gate becomes a one-line rule that every future planner and reviewer reads before it starts, and why that single file did more for output quality than any prompt I ever tuned.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Daniel Firu builds and operates an autonomous software-delivery harness for Expause, a Flutter and React product.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>architecture</category>
      <category>sdlc</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Why I stopped babysitting coding agents</title>
      <dc:creator>Daniel Firu</dc:creator>
      <pubDate>Thu, 10 Sep 2026 09:55:17 +0000</pubDate>
      <link>https://dev.to/daniel_firu/why-i-stopped-babysitting-coding-agents-54pf</link>
      <guid>https://dev.to/daniel_firu/why-i-stopped-babysitting-coding-agents-54pf</guid>
      <description>&lt;p&gt;&lt;em&gt;How a solo port turned into an autonomous pipeline, and what changed when the agents started reviewing each other.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;For most of the last year I have been building the same product twice. Expause is a video social app. The mobile app is Flutter, and it came first. In October 2025 I started a React web version that has to match it feature for feature. Same behaviour, same rules, same edge cases. There is one of me.&lt;/p&gt;

&lt;p&gt;That sounds like a job for a coding agent, and it was. Every web feature already has a working reference implementation in Dart. You point the agent at the Flutter screen, describe what "the same" means on the web, and let it type. For the first few months that is exactly what I did, and it was faster than doing it by hand. It was also the most tiring way I have ever written software.&lt;/p&gt;

&lt;h2&gt;
  
  
  The babysitting tax
&lt;/h2&gt;

&lt;p&gt;Here is what a day looked like. Open a session. Explain the architecture, again, because the last session is gone. Paste in the conventions. Ask for the feature. Watch it write. Approve a file. Notice it put business logic in a UI hook. Say so. Approve. Notice it skipped the unit test the conventions require. Say so. Notice it invented a constant that already exists under another name. Say so. Approve. Read the diff once more, because by now I no longer trust that I caught everything, and I know that I did not.&lt;/p&gt;

&lt;p&gt;None of those steps is hard. That is the problem. It is a stream of tiny judgements, none of them worth a coffee break, all of them required, and the stream never stops as long as the agent is typing. You cannot leave, because an unwatched agent drifts. You cannot really think either, because the interruptions come every minute. I was faster than before, I ended every day drained, and the code was not even that good. An agent that is being watched optimises for the watcher. It writes whatever makes me say "fine, next" in real time, which is not the same as what survives a proper review a week later.&lt;/p&gt;

&lt;p&gt;Somewhere around spring I admitted that the bottleneck in my project was me, sitting next to a machine, being the review process.&lt;/p&gt;

&lt;h2&gt;
  
  
  Writing the rules down first
&lt;/h2&gt;

&lt;p&gt;The first thing I built was not autonomy. It was boring. I wrote the rules down. One always-loaded project file with the things every agent must know, and one conventions document per layer of the app: how data access works, how the domain layer is shaped, how pages and hooks split their responsibilities. Every rule I had been repeating out loud went into a file instead.&lt;/p&gt;

&lt;p&gt;Then I stopped using one agent for everything. The work got split into narrow roles: a planner that turns a request into small, single-layer tasks; an implementer per layer that reads only its own task and its own layer's conventions; a committer that does nothing but tick the task off and commit. Each one starts with a fresh context and a tool allowlist that matches its job. That alone removed half the re-explaining, because nobody had to remember anything across a long session. Whatever an agent needed to know was in the file it was told to read.&lt;/p&gt;

&lt;p&gt;I was still pressing go on every task at this point. It was better. But I was still the reviewer.&lt;/p&gt;

&lt;h2&gt;
  
  
  The moment reviews changed the code
&lt;/h2&gt;

&lt;p&gt;The change that made this a different kind of tool was adding reviewers that are not me.&lt;/p&gt;

&lt;p&gt;A reviewer agent reads the finished diff against the plan and the conventions and writes its findings to a file. The part that mattered was a decision about tools rather than prompts: reviewers cannot edit. A reviewer that could fix what it found would quietly fix it and say nothing, and I would learn nothing. A reviewer that can only write a finding has to make its case.&lt;/p&gt;

&lt;p&gt;The first reports came back with things I would have missed, or things I would have been too tired to look for by the tenth file. A component over the size ceiling. A hook that had quietly grown a fifth responsibility. A value copied from mobile that ignored a shared constant the web already had. So I kept adding review angles. One reviewer checks parity with the mobile app, payload for payload and threshold for threshold. One checks that every file landed in the right layer. One reviews the whole branch. The bigger reports get a meta-review before anything is fixed, because reviewers make things up too.&lt;/p&gt;

&lt;p&gt;Then I did something I should have done from the start: I measured whether each reviewer earned its cost. The per-task reviewers, the ones that checked every unit right after it was written, turned out to be almost entirely redundant with the end-of-branch reviews, so they were dropped. What the same measurement showed was a class of defect nobody caught. The port had faithfully reproduced bugs from the mobile app, because every reviewer treated the reference implementation as the truth.&lt;/p&gt;

&lt;p&gt;So the last reviewer added was the skeptic. Its premise is that the plan, the implementation and the reference implementation can all be wrong, and its job is to report only what the other reviewers did not. On one of its first branches it found that the mobile app never reset an "upload in progress" flag on the failure path, which the web port had copied exactly, so one failed upload would have locked the button for good. Both clients got fixed. That is when I stopped thinking of the reviews as a safety net and started thinking of them as the reason the output was good.&lt;/p&gt;

&lt;p&gt;And this is the thing I did not expect: the code got better before the reviews ran, not only after. An implementer that will have to pass a parity reviewer, an architecture reviewer and a skeptic writes differently from one trying to satisfy a tired human in real time. The quality jump came from the agents having to convince each other.&lt;/p&gt;

&lt;h2&gt;
  
  
  Walking away
&lt;/h2&gt;

&lt;p&gt;Once the reviews were trustworthy, letting the whole thing run without me was mostly plumbing. A small daemon watches a folder. A task prompt dropped there gets its own branch and its own git worktree, and a headless session runs the entire flow: plan, review the plan until it converges, implement task by task, run every review gate, fix every finding, test the feature in a real browser through a tester agent that can drive one and nothing else, update the docs, push the branch. I get a notification.&lt;/p&gt;

&lt;p&gt;It did not work the first time, or the tenth. Runs got stuck, the API had bad days, a resume regenerated work it had already done. Each of those became a rule or a mechanism, and today a run that has to stop writes down where it was and continues from there later. That reliability story deserves its own article, and it will get one. For now the point is simpler: every wrong turn is a line in a file that the next run reads.&lt;/p&gt;

&lt;h2&gt;
  
  
  What stayed human
&lt;/h2&gt;

&lt;p&gt;I still test every feature by hand. This is the part I have no interest in automating.&lt;/p&gt;

&lt;p&gt;When a branch comes back, I use it the way I would use a colleague's branch. I open the app, click through the feature, try the edge cases, look at it on a phone. Then I write down what I found, in plain language, and drop that file where the task prompt went. A fix cycle starts: an agent verifies each observation against the code and plans the fixes, the same review gates grade that plan, the fixes are implemented, the browser tests are extended with one regression test per finding, and I get another notification.&lt;/p&gt;

&lt;p&gt;The last step of that cycle is my favourite part of the whole system. Every finding from my hands-on review that got past all the automated gates is distilled into a one-line rule in a lessons ledger, and every planner and every reviewer reads that ledger before it starts. The rules are boring and specific. Privileged data must be gated on the server, not only in the client. Reuse the shared constant even when the mobile app hand-rolls its own. Mutate optimistically before the await, the way mobile does. The first of those got past the pipeline three times on three different branches, caught by hand each time, before it became a line in the file. It has not come back since.&lt;/p&gt;

&lt;p&gt;So the human review is not a fallback for when the pipeline fails. It is the training signal. Every branch I review by hand makes the next one slightly harder to get wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a day looks like now
&lt;/h2&gt;

&lt;p&gt;I open an ordinary session in the repository and type what I want, the same way I did a year ago. "Add the block-user action to the chat header, same conditions as mobile." The difference is what happens next. The session asks how I want it handled: run it autonomously, do it here, talk it through first, or stop asking. I pick the first, it proposes a branch name, I confirm, and that is the end of my involvement in the implementation.&lt;/p&gt;

&lt;p&gt;Then I go do something else. Work on the mobile app. Think about the product. Have lunch. At some point a notification says the branch is ready. I review it, write what I found, drop it, and get another notification when the fixes are in. Then I merge.&lt;/p&gt;

&lt;p&gt;It is not flashy. Nothing about it looks like the demos. It is a folder, a daemon, a set of agents with narrow jobs and narrow tools, and a stack of files that say what each of them saw. The feeling is closer to having a careful team that works while I am away than to having a fast typist that needs me in the room. The babysitting tax is gone, and the code that comes back is better than the code I used to approve file by file.&lt;/p&gt;

&lt;h2&gt;
  
  
  It is open source
&lt;/h2&gt;

&lt;p&gt;The whole thing is published as a Claude Code plugin plus a small command-line tool that wires it into a repository: &lt;a href="https://github.com/firu-daniel/autonomous-sdlc-harness" rel="noopener noreferrer"&gt;github.com/firu-daniel/autonomous-sdlc-harness&lt;/a&gt;, Apache-2.0. The plugin carries the agents, the commands and the review flow. The CLI generates the permission profile, the scripts and the daemon, which are the parts nobody should have to get right by trial and error twice.&lt;/p&gt;

&lt;p&gt;The next article is the shape of it in plain words: what a run does from the moment you ask until the branch is pushed, with one diagram and no scorecards. After that, short pieces on the parts I think are interesting on their own: planning, the reviewers, the lessons flow, and how the pipeline files bugs against itself without being allowed to fix them.&lt;/p&gt;

&lt;p&gt;If you have been sitting next to an agent all day wondering whether this is what it is supposed to feel like, it is not. It gets boring, in the good way.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Daniel Firu builds and operates an autonomous software-delivery harness for Expause, a Flutter and React product.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>productivity</category>
      <category>devops</category>
      <category>opensource</category>
    </item>
  </channel>
</rss>
