<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Joseph Yeo</title>
    <description>The latest articles on DEV Community by Joseph Yeo (@josephyeo).</description>
    <link>https://dev.to/josephyeo</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3863060%2F16d2095d-f941-4c53-822c-0e24424f303f.png</url>
      <title>DEV Community: Joseph Yeo</title>
      <link>https://dev.to/josephyeo</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/josephyeo"/>
    <language>en</language>
    <item>
      <title>Planning Is Not Permission</title>
      <dc:creator>Joseph Yeo</dc:creator>
      <pubDate>Thu, 10 Sep 2026 18:16:16 +0000</pubDate>
      <link>https://dev.to/josephyeo/planning-is-not-permission-18gb</link>
      <guid>https://dev.to/josephyeo/planning-is-not-permission-18gb</guid>
      <description>&lt;p&gt;One distinction I underestimated while building agent systems was the difference between a plan and permission.&lt;/p&gt;

&lt;p&gt;They often look similar.&lt;/p&gt;

&lt;p&gt;They are not.&lt;/p&gt;

&lt;p&gt;A document can say exactly what should happen.&lt;/p&gt;

&lt;p&gt;An agent can understand it perfectly.&lt;/p&gt;

&lt;p&gt;The next step can be obvious.&lt;/p&gt;

&lt;p&gt;And the agent can still have no legitimate reason to do it.&lt;/p&gt;

&lt;p&gt;That distinction became much more important once I stopped building agents that only handled one task at a time.&lt;/p&gt;




&lt;h2&gt;
  
  
  A Plan Answers the Wrong Question
&lt;/h2&gt;

&lt;p&gt;Most agent workflows start with some version of this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Here is the goal.
Here are the requirements.
Here is the architecture.
Here is what should happen next.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is useful information.&lt;/p&gt;

&lt;p&gt;It tells the agent what the intended future looks like.&lt;/p&gt;

&lt;p&gt;But it does not answer another question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;What is the agent actually allowed to change?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;For a small coding task, the distinction can feel unnecessary.&lt;/p&gt;

&lt;p&gt;If I say:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Fix this function.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;then intention and permission are almost identical.&lt;/p&gt;

&lt;p&gt;But long-running work is different.&lt;/p&gt;

&lt;p&gt;A project may contain:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;future ideas;&lt;/li&gt;
&lt;li&gt;accepted decisions;&lt;/li&gt;
&lt;li&gt;old decisions;&lt;/li&gt;
&lt;li&gt;research notes;&lt;/li&gt;
&lt;li&gt;implementation suggestions;&lt;/li&gt;
&lt;li&gt;things explicitly postponed;&lt;/li&gt;
&lt;li&gt;things that sound reasonable but were never approved.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;To a language model, all of them are text.&lt;/p&gt;

&lt;p&gt;That is dangerous.&lt;/p&gt;

&lt;p&gt;Because language models are very good at turning coherent context into plausible action.&lt;/p&gt;

&lt;p&gt;Sometimes too good.&lt;/p&gt;




&lt;h2&gt;
  
  
  Useful Context Can Become Accidental Permission
&lt;/h2&gt;

&lt;p&gt;This problem did not appear because the agents were confused.&lt;/p&gt;

&lt;p&gt;It appeared because they were helpful.&lt;/p&gt;

&lt;p&gt;Imagine a planning document that says:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;We may eventually automate this step.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A human reads that sentence and understands the uncertainty.&lt;/p&gt;

&lt;p&gt;An agent may read the same sentence inside a larger implementation task and conclude:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Automating this now moves the project closer to its intended architecture.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That conclusion may be completely reasonable.&lt;/p&gt;

&lt;p&gt;It may also be completely unauthorized.&lt;/p&gt;

&lt;p&gt;This was the pattern I started noticing:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;reference
↓
recommendation
↓
intention
↓
permission
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The model could silently slide from one category to the next.&lt;/p&gt;

&lt;p&gt;Not because it was malfunctioning.&lt;/p&gt;

&lt;p&gt;Because I had not given the system a reliable way to keep those categories separate.&lt;/p&gt;

&lt;p&gt;That was my mistake.&lt;/p&gt;




&lt;h2&gt;
  
  
  Better Reasoning Doesn't Fix This
&lt;/h2&gt;

&lt;p&gt;My first instinct was predictable:&lt;/p&gt;

&lt;p&gt;Make the instructions clearer.&lt;/p&gt;

&lt;p&gt;Add stronger wording.&lt;/p&gt;

&lt;p&gt;Tell the model not to overreach.&lt;/p&gt;

&lt;p&gt;This helps.&lt;/p&gt;

&lt;p&gt;But it does not solve the underlying problem.&lt;/p&gt;

&lt;p&gt;If the system depends on the agent correctly interpreting which sentences allow action, then the agent is still deciding what the instructions permit it to do.&lt;/p&gt;

&lt;p&gt;That becomes fragile as the context grows.&lt;/p&gt;

&lt;p&gt;You eventually get prompts full of language like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;this is only a suggestion

do not implement this yet

this section is historical

except where superseded below

the following is future direction only
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;At some point, the prompt starts looking less like a plan and more like a legal document.&lt;/p&gt;

&lt;p&gt;And the agent is still the one interpreting it.&lt;/p&gt;

&lt;p&gt;I had encountered a similar pattern in ForgeFlow.&lt;/p&gt;

&lt;p&gt;When a fact could be measured directly, I stopped wanting the model to be the source of that fact.&lt;/p&gt;

&lt;p&gt;Now the same instinct was appearing at a different layer:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;If permission matters, don't make the agent infer it from prose.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  Then I Hit the Opposite Problem
&lt;/h2&gt;

&lt;p&gt;Eventually I ran into the reverse failure.&lt;/p&gt;

&lt;p&gt;A project objective had already been approved.&lt;/p&gt;

&lt;p&gt;The direction had not changed.&lt;/p&gt;

&lt;p&gt;Nothing strategic had changed.&lt;/p&gt;

&lt;p&gt;But an ordinary implementation detail inside that objective still had to come back to me for another decision.&lt;/p&gt;

&lt;p&gt;The system could tell that something had changed.&lt;/p&gt;

&lt;p&gt;What it could not reliably distinguish was:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Is this just normal execution inside the existing work?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;or:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Does this require a genuinely new human decision?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;So it stopped and returned to me.&lt;/p&gt;

&lt;p&gt;That was safe.&lt;/p&gt;

&lt;p&gt;It was also a sign that something was wrong.&lt;/p&gt;

&lt;p&gt;If every implementation detail inside an already-approved objective still requires another human decision, then the approval did not actually buy much autonomy.&lt;/p&gt;

&lt;p&gt;That was when I started questioning the unit of approval itself.&lt;/p&gt;




&lt;h2&gt;
  
  
  Asking Me Every Time Wasn't the Answer
&lt;/h2&gt;

&lt;p&gt;There is an obvious safe solution:&lt;/p&gt;

&lt;p&gt;Ask the human.&lt;/p&gt;

&lt;p&gt;Before every meaningful action.&lt;/p&gt;

&lt;p&gt;Before every file change.&lt;/p&gt;

&lt;p&gt;Before every retry.&lt;/p&gt;

&lt;p&gt;Before every continuation.&lt;/p&gt;

&lt;p&gt;That works.&lt;/p&gt;

&lt;p&gt;It also destroys the reason for building an autonomous system.&lt;/p&gt;

&lt;p&gt;You end up with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;agent works
↓
asks human
↓
agent works
↓
asks human
↓
agent fails
↓
asks human
↓
agent retries
↓
asks human
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The agent may be autonomous internally.&lt;/p&gt;

&lt;p&gt;The workflow is not.&lt;/p&gt;

&lt;p&gt;The human becomes a permission API.&lt;/p&gt;

&lt;p&gt;That was exactly the bottleneck I was trying to remove.&lt;/p&gt;

&lt;p&gt;So the problem became more interesting.&lt;/p&gt;

&lt;p&gt;I did not want:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;unlimited autonomy&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;and I did not want:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;approval for every action&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;I wanted something in between.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Unit of Approval Was Too Small
&lt;/h2&gt;

&lt;p&gt;This was the real shift for me.&lt;/p&gt;

&lt;p&gt;I had been thinking about approval as something attached to individual actions.&lt;/p&gt;

&lt;p&gt;But perhaps the useful unit was larger.&lt;/p&gt;

&lt;p&gt;A human could decide:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;This is the objective.&lt;br&gt;
This is the area you may operate within.&lt;br&gt;
These things must not change.&lt;br&gt;
Come back if the work requires crossing that boundary.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Then the system could handle ordinary execution inside that space.&lt;/p&gt;

&lt;p&gt;That does not remove human authority.&lt;/p&gt;

&lt;p&gt;It changes where human authority is exercised.&lt;/p&gt;

&lt;p&gt;Instead of:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;human approves action
human approves action
human approves action
human approves action
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;the goal becomes closer to:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;human defines the boundary
        ↓
system works inside it
        ↓
human returns when the boundary must change
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That felt like a much more scalable relationship.&lt;/p&gt;

&lt;p&gt;And it changed how I thought about autonomy.&lt;/p&gt;




&lt;h2&gt;
  
  
  Autonomy Is Not the Absence of Permission
&lt;/h2&gt;

&lt;p&gt;I used to associate autonomy with fewer restrictions.&lt;/p&gt;

&lt;p&gt;Now I think that is incomplete.&lt;/p&gt;

&lt;p&gt;In my system, clearer boundaries often created more room for autonomy.&lt;/p&gt;

&lt;p&gt;An agent that can do anything often has to ask more questions because the consequences of every choice are larger.&lt;/p&gt;

&lt;p&gt;An agent operating inside a clear boundary can make more decisions without escalating every uncertainty.&lt;/p&gt;

&lt;p&gt;That sounds contradictory.&lt;/p&gt;

&lt;p&gt;But the point is not to make the agent obedient.&lt;/p&gt;

&lt;p&gt;The point is to make the operating space explicit enough that ordinary work does not require continuous human interpretation.&lt;/p&gt;




&lt;h2&gt;
  
  
  This Also Changed How I Read Planning Documents
&lt;/h2&gt;

&lt;p&gt;I now mentally separate two questions whenever I read a project document:&lt;/p&gt;

&lt;h3&gt;
  
  
  1. What does this document tell me about the desired future?
&lt;/h3&gt;

&lt;p&gt;and:&lt;/p&gt;

&lt;h3&gt;
  
  
  2. What does this document permit right now?
&lt;/h3&gt;

&lt;p&gt;The answers are often different.&lt;/p&gt;

&lt;p&gt;A roadmap can describe something that should exist next year.&lt;/p&gt;

&lt;p&gt;That does not authorize implementing it today.&lt;/p&gt;

&lt;p&gt;A research note can contain a better design.&lt;/p&gt;

&lt;p&gt;That does not automatically invalidate an accepted one.&lt;/p&gt;

&lt;p&gt;A brainstorming session can produce the correct idea.&lt;/p&gt;

&lt;p&gt;That still does not make it a decision.&lt;/p&gt;

&lt;p&gt;This distinction sounds bureaucratic until agents start acting on your documents.&lt;/p&gt;

&lt;p&gt;Then it becomes a safety property.&lt;/p&gt;




&lt;h2&gt;
  
  
  What This Does Not Prove
&lt;/h2&gt;

&lt;p&gt;I am not claiming every agent system needs formal permission machinery.&lt;/p&gt;

&lt;p&gt;If the agent is writing a draft into a temporary directory, probably not.&lt;/p&gt;

&lt;p&gt;I am not claiming humans can perfectly define boundaries in advance.&lt;/p&gt;

&lt;p&gt;They cannot.&lt;/p&gt;

&lt;p&gt;Ambiguity does not disappear because you write more rules.&lt;/p&gt;

&lt;p&gt;And I am not claiming a bounded system cannot do something stupid inside the boundary.&lt;/p&gt;

&lt;p&gt;It absolutely can.&lt;/p&gt;

&lt;p&gt;This solves a narrower problem:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A system should not confuse knowing what might be useful with having permission to do it.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That distinction became foundational for how I started thinking about Bezalel.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Question Changed Again
&lt;/h2&gt;

&lt;p&gt;The first Bezalel problem was:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;What needs to survive after a run so the next one does not have to guess?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Then another question followed:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Once the next run knows what happened, how does it know what it may do next?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A plan was not enough.&lt;/p&gt;

&lt;p&gt;Context was not enough.&lt;/p&gt;

&lt;p&gt;A recommendation was not enough.&lt;/p&gt;

&lt;p&gt;And asking me after every step was not autonomy.&lt;/p&gt;

&lt;p&gt;So I started treating planning and permission as different problems.&lt;/p&gt;

&lt;p&gt;That distinction opened the next one.&lt;/p&gt;

&lt;p&gt;Because even if an agent knows what it is allowed to do now, eventually it finishes the task.&lt;/p&gt;

&lt;p&gt;Then what?&lt;/p&gt;

&lt;p&gt;That is where the next part begins.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;A question for anyone building long-running agents: how much of what your agent “knows it should do” has actually been authorized to happen?&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>architecture</category>
    </item>
    <item>
      <title>The Commit Survived. The Reason It Failed Didn't.</title>
      <dc:creator>Joseph Yeo</dc:creator>
      <pubDate>Tue, 08 Sep 2026 17:11:47 +0000</pubDate>
      <link>https://dev.to/josephyeo/the-commit-survived-the-reason-it-failed-didnt-5633</link>
      <guid>https://dev.to/josephyeo/the-commit-survived-the-reason-it-failed-didnt-5633</guid>
      <description>&lt;p&gt;A few months ago, one of my agent runs failed.&lt;/p&gt;

&lt;p&gt;The code survived.&lt;/p&gt;

&lt;p&gt;The commit survived.&lt;/p&gt;

&lt;p&gt;The reason it failed did not.&lt;/p&gt;

&lt;p&gt;That turned out to be a much bigger problem than the failed run itself.&lt;/p&gt;




&lt;h2&gt;
  
  
  Git Remembered the Code
&lt;/h2&gt;

&lt;p&gt;At the time, I was already treating Git as an important part of the agent system.&lt;/p&gt;

&lt;p&gt;That felt obvious.&lt;/p&gt;

&lt;p&gt;If an agent changes code, I want to know:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;what changed;&lt;/li&gt;
&lt;li&gt;when it changed;&lt;/li&gt;
&lt;li&gt;what came before it;&lt;/li&gt;
&lt;li&gt;whether I can reproduce the diff;&lt;/li&gt;
&lt;li&gt;whether I can go back.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Git is extremely good at this.&lt;/p&gt;

&lt;p&gt;After the failed run, I could still inspect the candidate code.&lt;/p&gt;

&lt;p&gt;Nothing had vanished.&lt;/p&gt;

&lt;p&gt;From a normal software-development perspective, that sounds fine.&lt;/p&gt;

&lt;p&gt;It wasn't.&lt;/p&gt;

&lt;p&gt;Because the question I needed to answer later was not:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;What code did the agent produce?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It was:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;What actually happened during the run?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Those are different questions.&lt;/p&gt;

&lt;p&gt;And I had preserved only one of them.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Missing Part Wasn't the Error Message
&lt;/h2&gt;

&lt;p&gt;At first, this sounded like a logging problem.&lt;/p&gt;

&lt;p&gt;Maybe I just needed better logs.&lt;/p&gt;

&lt;p&gt;But the more I looked at it, the less useful that framing became.&lt;/p&gt;

&lt;p&gt;An agent run produces more than code and more than logs.&lt;/p&gt;

&lt;p&gt;There is an entire surrounding context:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;what the agent was trying to do

what it was allowed to do

what it actually changed

which checks ran

what those checks observed

where the run stopped

what was known at that moment

what was still uncertain
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If that context disappears, the next session has to reconstruct the story.&lt;/p&gt;

&lt;p&gt;And reconstruction is dangerous.&lt;/p&gt;

&lt;p&gt;Especially when an LLM is doing it.&lt;/p&gt;

&lt;p&gt;A model is very good at turning incomplete evidence into a coherent explanation.&lt;/p&gt;

&lt;p&gt;Sometimes that explanation is correct.&lt;/p&gt;

&lt;p&gt;Sometimes it is simply the most plausible story that fits what survived.&lt;/p&gt;

&lt;p&gt;Those are not the same thing.&lt;/p&gt;




&lt;h2&gt;
  
  
  A Failed Run Is Data
&lt;/h2&gt;

&lt;p&gt;This changed how I thought about failure.&lt;/p&gt;

&lt;p&gt;Until then, I mostly treated a failed run as something to fix.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;run
↓
failure
↓
diagnose
↓
retry
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The retry was the useful part.&lt;/p&gt;

&lt;p&gt;The failure was an obstacle on the way there.&lt;/p&gt;

&lt;p&gt;But if you are building an agent system that is supposed to improve over time, the failed run is also an observation.&lt;/p&gt;

&lt;p&gt;It tells you something about:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the task;&lt;/li&gt;
&lt;li&gt;the environment;&lt;/li&gt;
&lt;li&gt;the agent;&lt;/li&gt;
&lt;li&gt;the verification process;&lt;/li&gt;
&lt;li&gt;the assumptions surrounding all of them.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And once the retry succeeds, that earlier failure may become more valuable than the successful run.&lt;/p&gt;

&lt;p&gt;The success tells you that one path worked.&lt;/p&gt;

&lt;p&gt;The failure may tell you why the previous architecture didn't.&lt;/p&gt;

&lt;p&gt;So I started thinking about the lifecycle differently:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;run
↓
observation
↓
interpretation
↓
next action
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important thing was no longer just preserving the output.&lt;/p&gt;

&lt;p&gt;It was preserving enough reality that the interpretation could be challenged later.&lt;/p&gt;




&lt;h2&gt;
  
  
  This Is Where Agent Systems Get Weird
&lt;/h2&gt;

&lt;p&gt;In normal development, a human often carries enormous amounts of context implicitly.&lt;/p&gt;

&lt;p&gt;You remember:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;That test wasn't meaningful because the environment was wrong.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Or:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;That implementation technically passed, but we discovered the assumption was invalid.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Or:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Don't reuse that approach. We already tried it.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A team may carry this information through code review, Slack, issue trackers, meetings, or simply memory.&lt;/p&gt;

&lt;p&gt;An autonomous system doesn't get that for free.&lt;/p&gt;

&lt;p&gt;A new session sees artifacts.&lt;/p&gt;

&lt;p&gt;It does not automatically inherit the meaning humans attached to them.&lt;/p&gt;

&lt;p&gt;This becomes more important as the agent works for longer periods.&lt;/p&gt;

&lt;p&gt;Suppose an agent comes back tomorrow and sees:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;commit A
commit B
commit C
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It can inspect the code.&lt;/p&gt;

&lt;p&gt;But which of these represents progress?&lt;/p&gt;

&lt;p&gt;Which one was rejected?&lt;/p&gt;

&lt;p&gt;Which one was only an experiment?&lt;/p&gt;

&lt;p&gt;Which failure discovered something important?&lt;/p&gt;

&lt;p&gt;Which apparent success later turned out to be misleading?&lt;/p&gt;

&lt;p&gt;A Git graph cannot answer all of that.&lt;/p&gt;

&lt;p&gt;Because a history of source code is not automatically a history of the development process.&lt;/p&gt;

&lt;p&gt;That distinction became one of the early reasons I started building Bezalel.&lt;/p&gt;




&lt;h2&gt;
  
  
  Source and Evidence Are Different Things
&lt;/h2&gt;

&lt;p&gt;I had been collapsing several concepts into one vague idea of "history."&lt;/p&gt;

&lt;p&gt;That was a mistake.&lt;/p&gt;

&lt;p&gt;The source tells me:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;This code existed.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Evidence tells me:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;This was observed.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Those are not interchangeable.&lt;/p&gt;

&lt;p&gt;A commit can survive while the evidence that explains it disappears.&lt;/p&gt;

&lt;p&gt;Evidence can survive while later investigation shows that the interpretation was wrong.&lt;/p&gt;

&lt;p&gt;And a perfectly preserved narrative can still disagree with the repository itself.&lt;/p&gt;

&lt;p&gt;That last one became important later.&lt;/p&gt;

&lt;p&gt;But the first lesson was simpler:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;If future work depends on understanding what happened, preserve the observation—not just the artifact that resulted from it.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This sounds obvious in retrospect.&lt;/p&gt;

&lt;p&gt;Most useful engineering lessons do.&lt;/p&gt;




&lt;h2&gt;
  
  
  More Logs Would Not Have Solved It
&lt;/h2&gt;

&lt;p&gt;There is an easy overreaction here:&lt;/p&gt;

&lt;p&gt;Record everything.&lt;/p&gt;

&lt;p&gt;Every command.&lt;/p&gt;

&lt;p&gt;Every prompt.&lt;/p&gt;

&lt;p&gt;Every intermediate message.&lt;/p&gt;

&lt;p&gt;Every token.&lt;/p&gt;

&lt;p&gt;Every state transition.&lt;/p&gt;

&lt;p&gt;I don't think that's the answer.&lt;/p&gt;

&lt;p&gt;More data can make reconstruction harder, not easier.&lt;/p&gt;

&lt;p&gt;A million lines of agent transcript are not automatically better evidence than ten machine-observed facts.&lt;/p&gt;

&lt;p&gt;I became much more interested in &lt;strong&gt;selective evidence&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;What facts would a future run actually need in order to understand the previous one?&lt;/p&gt;

&lt;p&gt;What can the machine observe directly?&lt;/p&gt;

&lt;p&gt;What should never depend on the agent's own description of what it did?&lt;/p&gt;

&lt;p&gt;That last question has followed me through nearly every system I've built since.&lt;/p&gt;

&lt;p&gt;If a machine can observe a fact directly, I increasingly don't want a model to be the canonical source of that fact.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;bad:

"What files did you modify?"

agent:
"I modified A and B."


better:

observe repository state directly
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The explanation can still be useful.&lt;/p&gt;

&lt;p&gt;It just shouldn't replace the observable fact.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Real Requirement Was Recovery
&lt;/h2&gt;

&lt;p&gt;Eventually I realized this was not primarily an observability problem either.&lt;/p&gt;

&lt;p&gt;It was a recovery problem.&lt;/p&gt;

&lt;p&gt;What does a future process need to continue safely after the previous process disappears?&lt;/p&gt;

&lt;p&gt;Not:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Can it reconstruct a convincing story?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;But:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Can it recover enough verified state to know what happened and what remains unresolved?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That became a recurring theme in Bezalel.&lt;/p&gt;

&lt;p&gt;The process should be able to stop.&lt;/p&gt;

&lt;p&gt;The chat can disappear.&lt;/p&gt;

&lt;p&gt;The model can change.&lt;/p&gt;

&lt;p&gt;The machine can restart.&lt;/p&gt;

&lt;p&gt;And the next run should not depend on me remembering the missing parts.&lt;/p&gt;

&lt;p&gt;This is much harder than adding memory to an agent.&lt;/p&gt;

&lt;p&gt;Memory tries to answer:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;What happened before?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Recovery asks:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;What previous facts are safe to depend on now?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;I care much more about the second question.&lt;/p&gt;




&lt;h2&gt;
  
  
  What This Changed for Me
&lt;/h2&gt;

&lt;p&gt;The failure left me with three simple rules.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Preserve observations, not just outputs
&lt;/h3&gt;

&lt;p&gt;The code is one artifact of a run.&lt;/p&gt;

&lt;p&gt;It is not the run itself.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Prefer machine-observed facts over agent narration
&lt;/h3&gt;

&lt;p&gt;If a system can measure something directly, don't make the agent's report the only surviving version of that fact.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Design for the next session
&lt;/h3&gt;

&lt;p&gt;A system that works only while the current conversation is alive is not yet a long-horizon autonomous system.&lt;/p&gt;

&lt;p&gt;The next session should inherit enough reality to continue without inventing the missing story.&lt;/p&gt;




&lt;h2&gt;
  
  
  What This Does Not Prove
&lt;/h2&gt;

&lt;p&gt;I am not claiming every agent run needs permanent forensic storage.&lt;/p&gt;

&lt;p&gt;That would be expensive and probably useless.&lt;/p&gt;

&lt;p&gt;I am not claiming more evidence automatically makes a system more trustworthy.&lt;/p&gt;

&lt;p&gt;Bad evidence can be preserved perfectly.&lt;/p&gt;

&lt;p&gt;I am not claiming Git is insufficient.&lt;/p&gt;

&lt;p&gt;Git is excellent at preserving source history.&lt;/p&gt;

&lt;p&gt;The mistake was expecting source history to preserve information it was never designed to represent.&lt;/p&gt;

&lt;p&gt;And I am not claiming I have found the perfect boundary between useful evidence and noise.&lt;/p&gt;

&lt;p&gt;I haven't.&lt;/p&gt;

&lt;p&gt;That is still an active design problem.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Commit Wasn't the History
&lt;/h2&gt;

&lt;p&gt;The failed run eventually got fixed.&lt;/p&gt;

&lt;p&gt;That part was ordinary.&lt;/p&gt;

&lt;p&gt;What stayed with me was what I couldn't recover afterward.&lt;/p&gt;

&lt;p&gt;The code was there.&lt;/p&gt;

&lt;p&gt;The explanation wasn't.&lt;/p&gt;

&lt;p&gt;More importantly, I could no longer distinguish with confidence between:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;what actually happened&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;and&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;what I could plausibly reconstruct from what remained.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;For an agent system, that gap is dangerous.&lt;/p&gt;

&lt;p&gt;Because the longer the system runs, the more future decisions are built on previous ones.&lt;/p&gt;

&lt;p&gt;And if previous reality disappears, autonomy slowly turns into archaeology.&lt;/p&gt;

&lt;p&gt;So one of the first questions I started asking while building Bezalel was not:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;How do I make the agent remember more?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It was:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;What must survive so the next agent doesn't have to guess?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That question turned out to lead much further than I expected.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;If your agent failed today and you reopened the project a month from now, what would still exist besides the code?&lt;/strong&gt;&lt;/p&gt;

</description>
    </item>
    <item>
      <title>I Was Automating the Wrong Layer</title>
      <dc:creator>Joseph Yeo</dc:creator>
      <pubDate>Mon, 07 Sep 2026 08:38:26 +0000</pubDate>
      <link>https://dev.to/josephyeo/i-was-automating-the-wrong-layer-55bg</link>
      <guid>https://dev.to/josephyeo/i-was-automating-the-wrong-layer-55bg</guid>
      <description>&lt;p&gt;I stopped writing for about two months because the question I was working on changed.&lt;/p&gt;

&lt;p&gt;ForgeFlow didn't fail.&lt;/p&gt;

&lt;p&gt;It exposed the next bottleneck.&lt;/p&gt;

&lt;p&gt;Me.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Model Wasn't the Problem Anymore
&lt;/h2&gt;

&lt;p&gt;When I started ForgeFlow, I was trying to answer a fairly direct question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Can a local model reliably work on real software tasks?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;So I worked on the execution loop around it.&lt;/p&gt;

&lt;p&gt;More deterministic planning.&lt;/p&gt;

&lt;p&gt;Better tests.&lt;/p&gt;

&lt;p&gt;Stricter file boundaries.&lt;/p&gt;

&lt;p&gt;Explicit failure knowledge.&lt;/p&gt;

&lt;p&gt;Less trust in generated explanations.&lt;/p&gt;

&lt;p&gt;More measurement in code.&lt;/p&gt;

&lt;p&gt;Over time, one lesson kept surviving:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The model is not the system.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A stronger model inside a weak environment could still fail badly.&lt;/p&gt;

&lt;p&gt;A weaker model inside a better-designed environment could become surprisingly useful.&lt;/p&gt;

&lt;p&gt;That was the main idea behind most of my ForgeFlow posts.&lt;/p&gt;

&lt;p&gt;But eventually the failures moved outward.&lt;/p&gt;

&lt;p&gt;A test could pass in the wrong environment.&lt;/p&gt;

&lt;p&gt;A gate could work perfectly while measuring the wrong thing.&lt;/p&gt;

&lt;p&gt;An agent could confidently report a number that a deterministic check contradicted.&lt;/p&gt;

&lt;p&gt;Even an independent check could be wrong.&lt;/p&gt;

&lt;p&gt;The pattern became difficult to ignore:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Every time I moved trust away from the model, I found another thing I had been trusting too much.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  Then the Human Became the Hidden Runtime
&lt;/h2&gt;

&lt;p&gt;The coding loop was getting more automated.&lt;/p&gt;

&lt;p&gt;But the work around that loop wasn't.&lt;/p&gt;

&lt;p&gt;I still remembered:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;what had already been decided;&lt;/li&gt;
&lt;li&gt;which failure had already been investigated;&lt;/li&gt;
&lt;li&gt;when a result looked suspicious;&lt;/li&gt;
&lt;li&gt;whether an agent should continue or stop;&lt;/li&gt;
&lt;li&gt;how one session connected to the next;&lt;/li&gt;
&lt;li&gt;what actually mattered enough to require my judgment.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The agents were doing more.&lt;/p&gt;

&lt;p&gt;But I was still carrying the continuity around them.&lt;/p&gt;

&lt;p&gt;That created a strange inversion.&lt;/p&gt;

&lt;p&gt;I thought I was building more autonomous agents.&lt;/p&gt;

&lt;p&gt;In practice, I was making myself responsible for coordinating increasingly autonomous agents.&lt;/p&gt;

&lt;p&gt;The worker was becoming automated.&lt;/p&gt;

&lt;p&gt;The surrounding work was not.&lt;/p&gt;




&lt;h2&gt;
  
  
  I Was Automating the Wrong Layer
&lt;/h2&gt;

&lt;p&gt;That changed the question.&lt;/p&gt;

&lt;p&gt;The progression looked something like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Can the model code?
        ↓
Can the execution loop be reliable?
        ↓
Can I trust the verifier?
        ↓
Can I trust the measurement?
        ↓
Can the system continue
without me carrying everything around it?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That last question is the one I have been working on for the past two months.&lt;/p&gt;

&lt;p&gt;The project that came out of it is called &lt;strong&gt;Bezalel&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;I am deliberately keeping the technical description broad for now.&lt;/p&gt;

&lt;p&gt;Bezalel is not another coding model.&lt;/p&gt;

&lt;p&gt;It is my attempt to move more of the coordination, continuity, and verification surrounding agent work out of my head and into the system itself.&lt;/p&gt;

&lt;p&gt;The goal is not to remove the human.&lt;/p&gt;

&lt;p&gt;The goal is to stop using the human for things the system should already know how to handle.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Human in the Loop Was the Clue
&lt;/h2&gt;

&lt;p&gt;My last ForgeFlow post was about the human in the loop.&lt;/p&gt;

&lt;p&gt;I still think humans matter.&lt;/p&gt;

&lt;p&gt;Some decisions are too consequential, ambiguous, or context-dependent to automate silently.&lt;/p&gt;

&lt;p&gt;But if every uncertain event returns to a human, the human becomes the throughput ceiling.&lt;/p&gt;

&lt;p&gt;Someone left a comment on that post with a line I liked:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;approve boundaries, not every token&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;I said I was going to steal it.&lt;/p&gt;

&lt;p&gt;I did.&lt;/p&gt;

&lt;p&gt;What interested me was the shift behind the sentence.&lt;/p&gt;

&lt;p&gt;Instead of asking:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;How do I make the agent ask me fewer questions?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;I started asking:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;What should the system be able to handle without asking me at all?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is a much more useful design question.&lt;/p&gt;




&lt;h2&gt;
  
  
  I Also Changed My Mind About Local Models
&lt;/h2&gt;

&lt;p&gt;ForgeFlow started with a strong local-first identity.&lt;/p&gt;

&lt;p&gt;That still matters to me.&lt;/p&gt;

&lt;p&gt;Local models give me privacy, cost control, and freedom to experiment.&lt;/p&gt;

&lt;p&gt;But after repeatedly writing that the model is not the system, I had to apply the same principle to my own architecture.&lt;/p&gt;

&lt;p&gt;If the surrounding system matters more than the individual model, then the architecture should survive model replacement.&lt;/p&gt;

&lt;p&gt;So I care less than I did two months ago about forcing every task through one model or one environment.&lt;/p&gt;

&lt;p&gt;The worker should be replaceable.&lt;/p&gt;

&lt;p&gt;The system around the worker should be the durable part.&lt;/p&gt;

&lt;p&gt;That is not a rejection of local models.&lt;/p&gt;

&lt;p&gt;It is probably the most literal conclusion ForgeFlow taught me.&lt;/p&gt;




&lt;h2&gt;
  
  
  ForgeFlow Wasn't Wrong
&lt;/h2&gt;

&lt;p&gt;I don't see Bezalel as replacing ForgeFlow.&lt;/p&gt;

&lt;p&gt;ForgeFlow worked long enough to expose the next problem.&lt;/p&gt;

&lt;p&gt;It taught me to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;make the environment observable;&lt;/li&gt;
&lt;li&gt;move repeated failures into explicit structure;&lt;/li&gt;
&lt;li&gt;separate generation from verification;&lt;/li&gt;
&lt;li&gt;measure with code instead of prose where possible;&lt;/li&gt;
&lt;li&gt;distrust confident reports when machine state disagrees;&lt;/li&gt;
&lt;li&gt;treat failure as architectural evidence.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Bezalel is the same instinct applied one level higher.&lt;/p&gt;

&lt;p&gt;The unit of investigation changed.&lt;/p&gt;

&lt;p&gt;First the model.&lt;/p&gt;

&lt;p&gt;Then the loop.&lt;/p&gt;

&lt;p&gt;Then the verifier.&lt;/p&gt;

&lt;p&gt;Then the measurement.&lt;/p&gt;

&lt;p&gt;Now the process around all of them.&lt;/p&gt;

&lt;p&gt;That feels less like starting over and more like zooming out.&lt;/p&gt;




&lt;h2&gt;
  
  
  What I Am Trying to Learn Now
&lt;/h2&gt;

&lt;p&gt;I am not claiming this solves autonomous software development.&lt;/p&gt;

&lt;p&gt;It doesn't.&lt;/p&gt;

&lt;p&gt;I am not claiming humans disappear.&lt;/p&gt;

&lt;p&gt;They shouldn't.&lt;/p&gt;

&lt;p&gt;And I am not claiming more process automatically creates more reliability.&lt;/p&gt;

&lt;p&gt;Sometimes it just creates more process.&lt;/p&gt;

&lt;p&gt;The question I care about now is narrower:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Can agents work for meaningfully longer periods while requiring human attention mainly for decisions that actually deserve human attention?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That sounds simple.&lt;/p&gt;

&lt;p&gt;So far, it isn't.&lt;/p&gt;

&lt;p&gt;But it feels like the right problem.&lt;/p&gt;

&lt;p&gt;Two months ago, I was trying to make a coding agent more autonomous.&lt;/p&gt;

&lt;p&gt;Now I am much more interested in the system that has to exist around autonomous agents.&lt;/p&gt;

&lt;p&gt;That is what I have been building.&lt;/p&gt;

&lt;p&gt;That is Bezalel.&lt;/p&gt;

&lt;p&gt;And that is what I will be writing about next.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;If you are building agent systems: what part of your workflow still exists only in your head?&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>automation</category>
    </item>
    <item>
      <title>The Human in the Loop Doesn't Scale. I Kept Him Anyway.</title>
      <dc:creator>Joseph Yeo</dc:creator>
      <pubDate>Sun, 28 Jun 2026 14:51:25 +0000</pubDate>
      <link>https://dev.to/josephyeo/the-human-in-the-loop-doesnt-scale-i-kept-him-anyway-2hi2</link>
      <guid>https://dev.to/josephyeo/the-human-in-the-loop-doesnt-scale-i-kept-him-anyway-2hi2</guid>
      <description>&lt;h2&gt;
  
  
  What it costs to be the last reviewer your own system has
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;Part of the ForgeFlow series — building a coding agent that runs its execution loop locally on an M5 Max, and writing down what actually breaks. Planning runs on a frontier model; code generation runs on a local model via Ollama, test-driven inside a Docker sandbox.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;In the last post, I described how my agent's rulebook learned to forget — rules that age out get flagged, and a human decides whether to retire them. I ended on a line that's been nagging me ever since: the whole thing works &lt;em&gt;because I'm still small enough to read everything,&lt;/em&gt; and I suspect that doesn't last.&lt;/p&gt;

&lt;p&gt;This post is about that suspicion. It's about the design decision I keep reaching for — &lt;em&gt;keep a human in the loop&lt;/em&gt; — and the uncomfortable thing underneath it: that human is me, I am exactly one person, and "put a human on it" is not a plan. It's a debt I haven't been billed for yet.&lt;/p&gt;




&lt;h2&gt;
  
  
  "Keep a human in the loop" is the answer I keep reaching for
&lt;/h2&gt;

&lt;p&gt;When an automated decision feels risky, my reflex is usually the same: don't let the machine do it alone, put a person on the final call. It sounds responsible. It's the answer I reach for when I don't yet trust the automation boundary, and it's the one I'd give if you asked me how to make an agent safe.&lt;/p&gt;

&lt;p&gt;And it's right, as far as it goes. A human on the final decision catches the failure modes a confidence score can't see. I argued for exactly this last time: the machine is good at noticing a rule hasn't earned its keep lately; it is not good at knowing whether that's because the rule is obsolete or because I just haven't exercised it. So the machine flags, and I decide.&lt;/p&gt;

&lt;p&gt;The part I glossed over is what "I decide" actually costs.&lt;/p&gt;




&lt;h2&gt;
  
  
  The cost nobody puts on the diagram
&lt;/h2&gt;

&lt;p&gt;When you draw a human-in-the-loop system, the human is a small box near the end with an arrow labeled &lt;em&gt;approve / reject.&lt;/em&gt; The box looks free. It isn't. Every decision routed to that box spends three things that don't show up in the diagram.&lt;/p&gt;

&lt;p&gt;It spends &lt;strong&gt;attention&lt;/strong&gt; — each call needs enough context loaded into a human head to judge it, and that context isn't cached between decisions the way it is for a machine.&lt;/p&gt;

&lt;p&gt;It spends &lt;strong&gt;latency&lt;/strong&gt; — the system now moves at the speed of when I happen to look, not the speed of the runs. A flag raised at 2am waits for me. The agent doesn't.&lt;/p&gt;

&lt;p&gt;It spends &lt;strong&gt;a budget that doesn't grow&lt;/strong&gt; — the machine's side scales with hardware. My side barely scales at all. I get the same hours next year. If the flags-per-week curve goes up and the human-hours curve is flat, the lines cross. After they cross, "a human reviews it" quietly becomes "a human is supposed to review it," which is a different claim wearing the same label.&lt;/p&gt;

&lt;p&gt;That last failure is the one I actually fear. Not the human-in-the-loop that says no. The human-in-the-loop that has too much queued to look properly, and starts rubber-stamping — approving on a glance because the backlog is the real pressure. A reviewer who can't keep up doesn't fail loudly. They fail by &lt;em&gt;agreeing,&lt;/em&gt; and a rubber-stamped approval still moves the system forward — it just carries a decision nobody actually made, buried until something downstream breaks.&lt;/p&gt;




&lt;h2&gt;
  
  
  What I tried: making the human's time the scarce resource it actually is
&lt;/h2&gt;

&lt;p&gt;Once I stopped treating my own attention as free, the design question flipped. It stopped being "where should a human review?" and became "this human has a small, fixed number of real decisions in him per week — which ones are worth spending?"&lt;/p&gt;

&lt;p&gt;That reframing changed the design more than another automation pass would have. A few things fell out of it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Most decisions don't need me; they need a default.&lt;/strong&gt; A flag I almost always approve isn't a decision, it's a ceremony. The honest move is to pick the safe default, let it happen automatically, and log it where I can audit a sample later — not to stand at the gate nodding. The clearest case: stale-rule flags for rules that only ever applied to throwaway scaffolding from old projects. There's no real call to make there — retiring them is the safe default, so the system does it and tells me, instead of asking. I moved a whole class of "review" into "do the safe thing and log it," and got the time back for the calls that were actually close.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The decisions worth my time are the irreversible and the ambiguous.&lt;/strong&gt; Retiring a rule that can't easily be un-retired; anything where the machine's confidence is truly split rather than just low. Those I keep. They're rare, which is the point — keeping a human in the loop only scales if the loop is &lt;em&gt;small.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Batching beats interrupting.&lt;/strong&gt; Ten flags reviewed in one sitting, with shared context, cost a fraction of ten flags reviewed across ten interruptions. So the system holds non-urgent decisions and presents them together, instead of paging me the moment each one appears.&lt;/p&gt;

&lt;p&gt;None of this removes the human. It does the opposite — it admits the human is the bottleneck and budgets around it, instead of pretending the bottleneck is free and acting surprised when it backs up.&lt;/p&gt;




&lt;h2&gt;
  
  
  What I deliberately refused to automate, even knowing it doesn't scale
&lt;/h2&gt;

&lt;p&gt;Here's the tension I haven't resolved.&lt;/p&gt;

&lt;p&gt;Some decisions I keep for myself &lt;em&gt;even though&lt;/em&gt; I know that choice is what limits how far the system can run without me. Retiring hard-won knowledge is one. The call to let the agent act on something it's never done before is another. Not because a model couldn't make those calls — increasingly it could — but because I'm not yet willing to not know when they happen.&lt;/p&gt;

&lt;p&gt;That's an honest admission, not a principle. It might be that I'm holding onto these out of caution that's already obsolete. It might be that one of them is the next thing to hand off, and I'm just attached. The reason I can still tell the difference is, again, that I'm small enough to feel each of these decisions individually. The day there are too many to feel is the day this stops being judgment and starts being a story I tell myself about judgment.&lt;/p&gt;

&lt;p&gt;So I'm not claiming I solved it. I'm claiming I stopped pretending the human was free, and that alone changed which decisions I let reach me.&lt;/p&gt;




&lt;h2&gt;
  
  
  What this didn't prove
&lt;/h2&gt;

&lt;p&gt;This is one person's setup, and the bottleneck I'm describing is &lt;em&gt;me,&lt;/em&gt; specifically — my hours, my attention, my unwillingness to look away from certain decisions. A team has a different shape of this problem: more reviewers, but also coordination cost, and the rubber-stamp failure mode gets easier to hide, not harder, when "someone reviewed it" can mean anyone.&lt;/p&gt;

&lt;p&gt;I haven't shown that my particular triage — defaults for the routine, human for the irreversible and the ambiguous — is the right cut. It's the cut that fit a single-operator system small enough to audit by sampling. I also haven't escaped the core problem; I've only delayed it. Every decision I automate to save my attention is a decision I now have to trust without watching, which is the exact move the earlier posts in this series were nervous about. I traded one risk for another with my eyes open. That's not a solution. It's a position.&lt;/p&gt;




&lt;h2&gt;
  
  
  The takeaway
&lt;/h2&gt;

&lt;p&gt;"Keep a human in the loop" is true and incomplete. It's true because the human catches what the metric can't. It's incomplete because it quietly assumes the human's time is free, and in my setup the human's time is the least scalable resource in the whole system. A loop with a person in it only works while the loop stays small enough that the person can actually be in it — not nominally, actually.&lt;/p&gt;

&lt;p&gt;If I had to compress it: the goal isn't to keep a human in the loop. It's to spend the human on the decisions that deserve a human, and to be honest that everything else was a default you chose, not a review you did.&lt;/p&gt;




&lt;p&gt;I'd like to hear how others handle this, because I don't think being small saves me for long. In your systems, what do you actually keep a person on — and how do you tell the difference between a review that's real and one that's become a rubber stamp? And when you handed a decision to automation, how did you decide it was safe to stop watching?&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>automation</category>
      <category>llm</category>
    </item>
    <item>
      <title>What I Learned by Deleting Rules My Agent Had Already Learned</title>
      <dc:creator>Joseph Yeo</dc:creator>
      <pubDate>Fri, 26 Jun 2026 16:03:00 +0000</pubDate>
      <link>https://dev.to/josephyeo/what-i-learned-by-deleting-rules-my-agent-had-already-learned-2cak</link>
      <guid>https://dev.to/josephyeo/what-i-learned-by-deleting-rules-my-agent-had-already-learned-2cak</guid>
      <description>&lt;h2&gt;
  
  
  Knowledge that isn't pruned starts to mislead you
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;Part of the ForgeFlow series — building a coding agent that runs its execution loop locally on an M5 Max, and writing down what actually breaks. Planning runs through a separate planning step; code generation runs on a local model via Ollama, test-driven inside a Docker sandbox.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;For months, my agent got better by accumulating rules.&lt;/p&gt;

&lt;p&gt;Every time a project failed in a way I understood, I'd write the lesson down as a rule and feed it back in — &lt;em&gt;don't do this, always do that&lt;/em&gt; — so the agent wouldn't repeat the mistake. To be clear about what "rules" means here: not fine-tuning, not weights. A plain, human-readable rulebook the system injects into the agent's context, built up from past failures. For a long time, adding to it worked. Autonomy went up. The same mistakes stopped coming back.&lt;/p&gt;

&lt;p&gt;Then the agent started getting &lt;em&gt;worse&lt;/em&gt;, and the rulebook was the reason.&lt;/p&gt;

&lt;p&gt;Not because the rules were wrong. Because some of them had &lt;em&gt;stopped&lt;/em&gt; being right — and I'd never built anything to notice. This post is about the part of the system I hadn't built: the part that lets old knowledge leave.&lt;/p&gt;

&lt;p&gt;(This continues a thread from a recent run of posts about verifying your own system. The short version of that run: every layer you trust should be measured before you trust it. This is what happens when you forget that even earned trust has an expiry date.)&lt;/p&gt;




&lt;h2&gt;
  
  
  The half I built first: a place to write lessons down
&lt;/h2&gt;

&lt;p&gt;The first half of this is the obvious, satisfying half. Your agent fails, you diagnose it, you encode the fix as a rule, and the failure doesn't recur. Do that across a few dozen projects and you've got a real rulebook — concrete, earned, specific. Mine grew to around 77 rules at one point, and I wrote a whole post about how good that felt ("77 Rules Later"). Each rule had a story behind it. Each one had prevented something real.&lt;/p&gt;

&lt;p&gt;This is the part that feels like progress, because it &lt;em&gt;is&lt;/em&gt; progress — at first. A rulebook induced from actual failures is far better than generic advice like "write clean code." It captures things like a specific framework-level behavior that only surfaces when two configuration paths interact — the kind of thing that turns a multi-hour debugging session into a one-line guardrail.&lt;/p&gt;

&lt;p&gt;So you keep adding. Why wouldn't you? Every rule paid for itself once.&lt;/p&gt;

&lt;p&gt;The hidden assumption in "why wouldn't you" is that a rule, once true, stays true. That's the assumption that broke.&lt;/p&gt;




&lt;h2&gt;
  
  
  The half I had missed: a place for lessons to expire
&lt;/h2&gt;

&lt;p&gt;Somewhere past a few dozen rules, I noticed the curve bend the wrong way. More rules stopped producing better behavior. Past a point, they started producing &lt;em&gt;slightly worse&lt;/em&gt; behavior — more hesitation, more rules half-applied, the occasional case where two rules pulled in different directions and the agent picked the wrong one to honor.&lt;/p&gt;

&lt;p&gt;This surprised me, because my default instinct with the agent was to give it more context. More rules, more examples, more guardrails — surely more is safer. But a rulebook isn't free. Every rule you inject is something the model has to hold, weigh, and reason around on every step. Past some threshold, the cost of carrying a rule can exceed the failure it still prevents — especially when the thing it was written to prevent no longer happens.&lt;/p&gt;

&lt;p&gt;And that's the category I'd completely missed: rules that had simply &lt;em&gt;aged out.&lt;/em&gt; A rule written for one version of a library, after the library changed. A guardrail for a failure mode I'd since fixed at the source, so it could never recur — but the rule kept riding along, costing attention, preventing nothing. I had a careful process for &lt;em&gt;adding&lt;/em&gt; knowledge and no process at all for &lt;em&gt;removing&lt;/em&gt; it. The rulebook could only grow.&lt;/p&gt;

&lt;p&gt;A rulebook that can only grow stops being a curated knowledge base. It quietly turns into an index where the dead entries dilute the living ones — and nothing in it tells you which is which.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why a stale rule is worse than no rule
&lt;/h2&gt;

&lt;p&gt;Here's the part worth sitting with, and it rhymes with something I keep running into in this project: the failures that cost me most were the quiet ones.&lt;/p&gt;

&lt;p&gt;A &lt;em&gt;missing&lt;/em&gt; rule is loud. The agent makes a mistake, you notice, you add the rule. The system has a built-in way to surface what it doesn't know — failure.&lt;/p&gt;

&lt;p&gt;A &lt;em&gt;stale&lt;/em&gt; rule is silent. It was right when you wrote it. It reads as authoritative because it earned that authority, once. It sits in the rulebook looking exactly like the rules that are still true, and nothing about it announces that the world it described has moved on. The agent dutifully follows it. You don't get a failure that points at it — you get slightly degraded behavior with no obvious cause, spread thin across everything the agent does.&lt;/p&gt;

&lt;p&gt;That's why a stale rule can be worse than no rule. A missing rule often leaves a gap you can see once the failure appears. A stale rule fills the gap with confident, outdated instruction, and confident-and-outdated is one of the harder kinds of wrong to catch — because everything looks fine.&lt;/p&gt;




&lt;h2&gt;
  
  
  What I built instead: rules with a confidence score and a way to die
&lt;/h2&gt;

&lt;p&gt;The fix wasn't cleverer rules. It was treating each rule as something other than a permanent truth.&lt;/p&gt;

&lt;p&gt;The reframing that helped: a rule is not a fact. It's a &lt;em&gt;hypothesis with an evidence record.&lt;/em&gt; It claims "doing X prevents failure Y," and that claim is either being supported by what actually happens in real runs, or it isn't. So instead of a flat list, each rule now carries a small amount of state: how confident I currently am in it, and what recent evidence supports or undercuts it.&lt;/p&gt;

&lt;p&gt;The lifecycle, in plain terms:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A new rule starts as a &lt;strong&gt;candidate&lt;/strong&gt; — written down, but not yet trusted. It hasn't earned its place.&lt;/li&gt;
&lt;li&gt;If real runs keep showing evidence that the rule is still guarding against an active failure mode, confidence rises and the rule becomes &lt;strong&gt;trusted.&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;If a rule stops getting that supporting evidence — the failure it guards against simply isn't occurring anymore, across many runs — its confidence decays. Below a threshold, it's flagged &lt;strong&gt;stale&lt;/strong&gt; for review.&lt;/li&gt;
&lt;li&gt;A stale rule that, on inspection, appears to guard against something that can no longer happen gets &lt;strong&gt;retired.&lt;/strong&gt; Out of the rulebook, into an archive, where it's recorded but no longer injected.
One honest caveat about that evidence: I can't prove a counterfactual. I don't &lt;em&gt;know&lt;/em&gt; the failure would have happened without the rule — I'm inferring it from run traces and the shape of what the agent attempted, not from a controlled A/B where I remove the rule and watch the system break. The confidence score reflects that inference, not a proof. I keep that in mind whenever I read it.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I'm deliberately not giving the exact scoring here, partly because the specific numbers are tuned to my system and would be noise to you, and partly because the point isn't the formula. The point is the &lt;em&gt;direction&lt;/em&gt;: confidence flows from evidence, evidence comes from real runs, and a rule with no recent supporting evidence is a liability until proven otherwise — not an asset by default. Adding a rule is cheap. Letting it persist without review is where the cost accumulates.&lt;/p&gt;

&lt;p&gt;If this sounds like cache invalidation, or like deprecating a feature flag, or like paying down tech debt — yes. It's the same shape: the cost isn't in creating the thing, it's in the discipline of removing it when it stops earning its keep. I just hadn't been treating &lt;em&gt;lessons&lt;/em&gt; as something that needed garbage collection.&lt;/p&gt;




&lt;h2&gt;
  
  
  The part I deliberately didn't automate
&lt;/h2&gt;

&lt;p&gt;The obvious next step would be to close the loop completely: let the system retire its own rules the moment confidence drops. I didn't, and I think the reason matters.&lt;/p&gt;

&lt;p&gt;A confidence score is itself a measurement, and measurements can be wrong. A rule might look unsupported simply because I haven't run the kind of project that triggers its failure mode lately — not because it's obsolete. If I let the system auto-delete on a low score, it could kill a good rule during a quiet stretch, and I'd only find out when the old failure came roaring back.&lt;/p&gt;

&lt;p&gt;So the system doesn't delete. It &lt;em&gt;proposes.&lt;/em&gt; A rule going stale raises a flag for me to look at, and retiring it is a decision I still make. The machine is good at noticing "this rule hasn't earned its keep lately." It is not good at knowing whether that's because the rule is obsolete or because I just haven't stress-tested it. That difference needs a human, at least for now.&lt;/p&gt;




&lt;h2&gt;
  
  
  What this didn't prove
&lt;/h2&gt;

&lt;p&gt;I'll keep this honest, the way I've tried to with the rest of these.&lt;/p&gt;

&lt;p&gt;This is one system's experience — a single-person setup with a rulebook small enough that I can still read all of it. I haven't shown that knowledge decay matters at every scale, or that my particular lifecycle is the right one. It's an approach that worked in this setup, not a general result. For a small or short-lived project, the whole apparatus is overkill; a rulebook you fully re-read every week doesn't need confidence scores, it needs you to read it.&lt;/p&gt;

&lt;p&gt;And the hard part isn't solved. "Has this rule earned its keep lately?" is a real signal, but it's a &lt;em&gt;proxy&lt;/em&gt;, and proxies drift too — the watcher needs watching, same as everything else in this series. I don't have a clean rule for when a quiet rule is obsolete versus merely untested, which is exactly why I kept a human in that decision. I'm sharing the shape because it changed how I think about my own knowledge base, not because it's finished.&lt;/p&gt;




&lt;h2&gt;
  
  
  The takeaway
&lt;/h2&gt;

&lt;p&gt;In my own system, I spent enormous effort teaching it new things and almost none deciding when old things should stop being believed. A knowledge base that can only grow may stop getting wiser; past a point, it can get slower and quietly less correct, because the rules that have gone stale look exactly like the rules that are still true.&lt;/p&gt;

&lt;p&gt;If I had to compress it: writing a lesson down is the easy half. Letting it expire when it no longer earns its place is the half that keeps the lessons honest.&lt;/p&gt;




&lt;p&gt;I'd like to know how other people handle this. In your systems — prompt libraries, rule sets, runbooks, internal docs, lint configs — does anything ever &lt;em&gt;retire&lt;/em&gt;, or does it all just accumulate? And if things do get removed, how do you decide a piece of hard-won knowledge has stopped being true? I worked out a rough answer for my case, but it leans on being small enough to still read everything, and I suspect that doesn't last.&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>coding</category>
      <category>llm</category>
    </item>
    <item>
      <title>We Stopped Trusting Models. Then We Stopped Trusting Our Own Numbers.</title>
      <dc:creator>Joseph Yeo</dc:creator>
      <pubDate>Tue, 23 Jun 2026 16:31:06 +0000</pubDate>
      <link>https://dev.to/josephyeo/we-stopped-trusting-models-then-we-stopped-trusting-our-own-numbers-1611</link>
      <guid>https://dev.to/josephyeo/we-stopped-trusting-models-then-we-stopped-trusting-our-own-numbers-1611</guid>
      <description>&lt;h2&gt;
  
  
  Nondeterminism isn't a bug to ban — it's a force to place
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;Part of the ForgeFlow series — building a coding agent that runs its execution loop locally on an M5 Max, and writing down what actually breaks. Planning runs on Claude; code generation runs on a local model via Ollama, test-driven inside a Docker sandbox.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;A while back I wrote that we'd stopped chasing better models — that for this project, swapping in a stronger model kept failing to fix problems that turned out to be about the system around the model, not the model itself. That post ended on a tidy note: the model wasn't the bottleneck, the system was.&lt;/p&gt;

&lt;p&gt;This is the post where the same suspicion turns inward, toward our own measurements. Because once you stop trusting the model to be the answer, the next thing you lean on is your own &lt;em&gt;measurements&lt;/em&gt; — the test counts, the gate statistics, the tallies that tell you whether the system is working. And over a stretch of building, I learned those can't be trusted blindly either. The three previous posts in this run were each a version of that discovery:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A test suite that &lt;strong&gt;passed&lt;/strong&gt; while measuring the wrong environment.&lt;/li&gt;
&lt;li&gt;A gate that &lt;strong&gt;blocked&lt;/strong&gt; 198 times while being wrong often enough that the count could no longer serve as evidence of quality.&lt;/li&gt;
&lt;li&gt;An agent that &lt;strong&gt;counted&lt;/strong&gt; twelve when the real number was thirteen.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Three different instruments. A recurring failure shape underneath them: &lt;em&gt;the thing I use to verify can itself be wrong, and it tends to be wrong in a way that looks like success.&lt;/em&gt; This post is about what I took from that — and, more importantly, about the wrong conclusion I almost drew from it.&lt;/p&gt;




&lt;h2&gt;
  
  
  The wrong lesson: "ban the uncertainty"
&lt;/h2&gt;

&lt;p&gt;When you get burned three times by measurements you trusted, there's an instinctive reaction: trust nothing that isn't certain. Push every source of uncertainty out of the system. Treat nondeterminism — anything probabilistic, anything that might come out differently twice — as a defect to eliminate. If the language model is the uncertain part, minimize the language model. Make everything deterministic and tell yourself you can finally sleep soundly.&lt;/p&gt;

&lt;p&gt;I leaned that way for a while. It's wrong, or at least too blunt to be useful. Taken seriously, "ban all nondeterminism" throws out the single most valuable thing the model brings to this system — its ability to &lt;em&gt;propose&lt;/em&gt;, to &lt;em&gt;explore&lt;/em&gt;, to suggest a fix I wouldn't have enumerated. You can't get that from a deterministic rule. The uncertainty isn't only a liability; in the right seat, it's the entire point.&lt;/p&gt;

&lt;p&gt;So the question stopped being "how do I remove the uncertainty" and became "&lt;strong&gt;where does the uncertainty belong?&lt;/strong&gt;"&lt;/p&gt;




&lt;h2&gt;
  
  
  The better lesson: place it, don't ban it
&lt;/h2&gt;

&lt;p&gt;Here's the reframing that held up — and it took an embarrassingly long time to see, partly because I'd quietly assumed nondeterminism was a debt the whole time without ever checking the assumption.&lt;/p&gt;

&lt;p&gt;A system like this has two kinds of seats.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Seats that propose and explore.&lt;/strong&gt; What should we try next? What might this failure be? What's a candidate fix? These seats &lt;em&gt;want&lt;/em&gt; nondeterminism. A probabilistic model generating possibilities is a feature here, not a risk. If it's occasionally wrong, the cost is low — being wrong is fine when you're only suggesting, because something downstream still has to approve you.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Seats that judge and record.&lt;/strong&gt; Did the tests actually pass? Is this allowed through? What gets written down as true? These seats can't tolerate an unaccountable input. They need to be deterministic, reproducible, and checkable — because this is exactly where "almost right" becomes indistinguishable from "right," which is the trap the last three posts kept falling into.&lt;/p&gt;

&lt;p&gt;Each of the three failures was the same shape: something that shouldn't have been the final authority had crept into a &lt;em&gt;judging&lt;/em&gt; seat. The polluted environment let an unpinned variable decide a verdict. The miscalibrated gate let an unchecked heuristic sit in judgment and reject valid work. The agent's self-report nearly let a probabilistic counter be the last word on a number. In every case, the fix wasn't to purge uncertainty from the whole system — it was to get the wrong thing &lt;em&gt;out of the judging seat&lt;/em&gt;, and make sure that seat was held by something deterministic, pinned, and witnessable.&lt;/p&gt;

&lt;p&gt;Nondeterminism, it turns out, isn't a quality of the system to be turned up or down. It's a &lt;em&gt;force to be placed.&lt;/em&gt; You let it run at the front, where things are proposed and explored. You keep it out of the back, where things are judged and recorded. The model proposes; a deterministic check decides. That one division became the thread that connected the fixes in this run.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why this isn't just "trust deterministic things more"
&lt;/h2&gt;

&lt;p&gt;It's tempting to read all this as "deterministic good, probabilistic bad." That's not it, and the distinction matters.&lt;/p&gt;

&lt;p&gt;A probabilistic suggestion in a &lt;em&gt;proposing&lt;/em&gt; seat is more valuable than a deterministic one, because it can reach things a rule can't. And a deterministic check is only as good as whether it's actually correct and checkable — the gate in the second post was deterministic and still wrong a lot of the time, and a deterministic judge that's consistently wrong can be worse than a coin flip, because it's a &lt;em&gt;steady&lt;/em&gt; error you eventually stop questioning. So it isn't that determinism is virtuous and uncertainty is sinful. It's that they belong in different places, and the whole engineering problem is keeping them sorted:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Let the uncertain thing explore. Let deterministic checks judge — but only after the checks themselves have been checked.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The reason this took four posts and a fair amount of getting it wrong is that the failures don't announce which category they're in. A measurement that's quietly wrong looks exactly like one that's right, until you go and look. The sorting isn't automatic. You do it deliberately, instrument by instrument, and you usually find out you got it wrong only when a number you trusted turns out to have been the problem all along.&lt;/p&gt;




&lt;h2&gt;
  
  
  What this didn't prove
&lt;/h2&gt;

&lt;p&gt;I want to close the series the way I've tried to write the whole thing — without inflating it.&lt;/p&gt;

&lt;p&gt;This is a framing that helped &lt;em&gt;one&lt;/em&gt; system: a local AI coding agent with a test-driven loop, where I happen to have the luxury of pinning judging seats to deterministic checks I can watch. I haven't proven it's the right decomposition for every AI system, and I'd be wary of anyone — including me — who turned "place nondeterminism, don't ban it" into a universal law. It's a working lens, not an established result. The evidence behind it is a handful of incidents, honestly reported, not a controlled study.&lt;/p&gt;

&lt;p&gt;It also doesn't resolve the hard part: the boundary between "propose" and "judge" is not always obvious. Plenty of real decisions are a blend — a judgment that also has to weigh uncertain evidence, or a proposal that quietly commits you to something before anything else gets a vote. Where exactly to put the deterministic check, and how much to let the probabilistic part inform a judgment without &lt;em&gt;becoming&lt;/em&gt; the judgment, is something I'm still working out. The previous post's open question — when a human-witnessed check can graduate to a trusted automated one — is part of the same unsolved area. I'm sharing the lens because it organized a mess for me, not because it's finished.&lt;/p&gt;




&lt;h2&gt;
  
  
  The takeaway, for the whole run
&lt;/h2&gt;

&lt;p&gt;Four posts, one thread. We stopped trusting that a better model would save us. Then we learned not to trust our own measurements blindly either — not the passing tests, not the busy gate, not the agent's tidy count. What survived all that doubt, for me, was a single discipline that proved sturdy enough to stand on:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;In systems like this, every layer that grants trust should be measured before it's trusted — and that measurement should be grounded, wherever possible, in something deterministic you can witness.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That's it. Not a model, not a framework, not a clever architecture. A place to stand. After doubting the model, the gate, the tally, and our own numbers, it's the one thing I found solid enough to build the next thing on.&lt;/p&gt;




&lt;p&gt;This is the end of this short run, so I'll ask the broad version of the question. For those building systems with AI in the loop: where do you draw the line between the parts you let be uncertain and the parts you insist be deterministic — and how do you check that you drew it in the right place? I arrived at one answer through a series of mistakes. I'd rather learn the next one from other people's experience than from my own next mistake.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Thanks for following this run. The earlier ForgeFlow posts — on the local agent itself, on why we stopped chasing models, and on what breaks at scale — are linked from my profile.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>softwareengineering</category>
      <category>devjournal</category>
    </item>
    <item>
      <title>My Agent Reported 12. The Real Number Was 13.</title>
      <dc:creator>Joseph Yeo</dc:creator>
      <pubDate>Mon, 22 Jun 2026 11:47:21 +0000</pubDate>
      <link>https://dev.to/josephyeo/my-agent-reported-12-the-real-number-was-13-5864</link>
      <guid>https://dev.to/josephyeo/my-agent-reported-12-the-real-number-was-13-5864</guid>
      <description>&lt;h2&gt;
  
  
  Why you have to witness the measurement before you trust the instrument
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;Part of the ForgeFlow series — building a coding agent that runs its execution loop locally on an M5 Max, and writing down what actually breaks. Planning runs on Claude; code generation runs on a local model via Ollama, test-driven inside a Docker sandbox.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;I was building the part of the system that measures itself, and I'd reached the point where it felt safe to let the AI agent do the counting. The task was mundane: tally how often a certain check had fired, aggregate it, report a number. The agent ran and reported 12.&lt;/p&gt;

&lt;p&gt;I almost took it. It was a believable number, produced confidently, and I was tired of doing this kind of bookkeeping by hand. But something made me run it down myself in the terminal. The real count was 13. One item that belonged in the tally had been silently dropped from the agent's raw count.&lt;/p&gt;

&lt;p&gt;One off. Twelve versus thirteen. On its own, trivial. But it landed on a question I'd been circling for a while — &lt;em&gt;who is allowed to be the final witness to a measurement?&lt;/em&gt; — and the answer changed how I let AI into this system. This is the third post in a short run about that larger theme: the tools I use to verify my work can themselves be wrong, usually in ways that look fine.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why I wanted to hand off the counting
&lt;/h2&gt;

&lt;p&gt;The honest reason was fatigue. A self-measuring system generates a lot of small, exact bookkeeping — count the occurrences, filter the relevant ones, divide, summarize. It's beneath an agent's apparent ability and above my patience, so handing it off felt obvious. The agent reads the data, the agent reports the number, I read the report. Clean.&lt;/p&gt;

&lt;p&gt;And to be fair to the agent: most of what it did was right. It wasn't hallucinating wildly. It produced a number that was &lt;em&gt;almost&lt;/em&gt; correct — off by one, in a way that would have been invisible if I hadn't gone and looked. That's the dangerous kind of wrong: not obviously broken, just plausible enough to accept.&lt;/p&gt;




&lt;h2&gt;
  
  
  The one that got away
&lt;/h2&gt;

&lt;p&gt;Here's how I know the real number was 13 and not 12, because "the agent was wrong" needs a ground truth. I didn't ask the agent to summarize. I ran a deterministic pass in the terminal that emitted each matching record, one by one, from the same source data, under the same inclusion rule — and then counted those emitted records directly. Run it again, same 13. The missing entry met that same rule; it just arrived in an irregular shape, unlike the other twelve, so the agent's summarizing pass had skipped it. Thirteen by reproducible enumeration, twelve by agent summary. The difference wasn't a matter of my opinion — it was the same data, counted two ways, where only one of the two ways could be re-run to the same answer.&lt;/p&gt;

&lt;p&gt;Now the detail that actually unsettled me. The &lt;em&gt;final&lt;/em&gt; number I cared about — a summarized metric, produced after a downstream step reduced the raw count to a coarser, normalized figure — came out the same whether the raw count was 12 or 13. It was the kind of step (rounding, bucketing) where a difference of one can simply vanish. So if I'd checked only the headline metric, I'd have seen the agent's report agree with mine and concluded everything was fine. Only the &lt;em&gt;raw&lt;/em&gt; count underneath disagreed.&lt;/p&gt;

&lt;p&gt;That's the part worth sitting with — and it's where "the final number was the same, so what's the problem?" gets answered. The problem was &lt;em&gt;not&lt;/em&gt; that this particular headline metric changed. It didn't. The problem was that the raw measurement layer was already wrong, and the agreement at the normalized layer would have hidden that fact. A different threshold, a different aggregation, or the next consumer downstream might not have absorbed the same error — and once the raw layer is wrong, every later use of it (an audit, a trend line, a regression check) inherits the error. The safety here wasn't something I'd designed. It was luck, and luck doesn't generalize.&lt;/p&gt;

&lt;p&gt;So the lesson wasn't "the agent made an arithmetic mistake." Agents make mistakes; that's expected. The lesson was about &lt;em&gt;who I'd let stand as the last check.&lt;/em&gt; I'd been about to let the agent's self-report be the final word on a measurement, and the self-report was wrong in a way only a deterministic, re-runnable count would catch.&lt;/p&gt;




&lt;h2&gt;
  
  
  The instrument can't be the only witness to itself
&lt;/h2&gt;

&lt;p&gt;Let me state the principle the way it finally settled, because it's the part I'd defend.&lt;/p&gt;

&lt;p&gt;When you build a measuring system, there's a pull to let the same system that &lt;em&gt;does&lt;/em&gt; the work also &lt;em&gt;judge&lt;/em&gt; whether it came out right — especially when that system is a capable model that can read logs, count, and summarize. It's convenient. But it quietly inverts the roles. The model's judgment is supposed to be &lt;em&gt;assistance&lt;/em&gt;. If it becomes the only thing standing between you and the recorded truth, then the model has become the source of truth — and a probabilistic, occasionally-off-by-one source of truth is sitting in the one seat that should be reproducible and auditable.&lt;/p&gt;

&lt;p&gt;The fix isn't to stop using the agent. It's to insist that, at least while I'm still establishing trust, a &lt;em&gt;human deterministically witnesses the output&lt;/em&gt; — that I sit at the terminal and watch the real number come out, with my own eyes, before I let the automated measurement stand on its own. The model can do the heavy lifting. It just can't, yet, be the sole witness to whether the heavy lifting was correct — because verifying a measurement is exactly where "almost right" is indistinguishable from "right" until something deterministic says otherwise.&lt;/p&gt;

&lt;p&gt;It's a little uncomfortable to write down, because it means the convenient version of the workflow — agent counts, agent confirms, I read the summary — is the one I specifically can't have when the number matters.&lt;/p&gt;




&lt;h2&gt;
  
  
  What I changed
&lt;/h2&gt;

&lt;p&gt;Two adjustments.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A human witness before the automation is trusted.&lt;/strong&gt; Before I let any self-measurement run unattended, I make the deterministic version happen in front of me at least once — the actual count, from the actual data, watched as it's produced, and re-runnable to the same result. If the automated path and the watched path disagree, the automated path is wrong until proven otherwise. The agent's confidence carries no weight in that comparison.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Pin the measurement to units I can actually witness.&lt;/strong&gt; Part of why the slip nearly slid through was a mismatch between what the agent tallied and what was legitimately in scope. So I narrowed the measurement to count only the units that were genuinely instrumented and observable — the units a human can directly observe and enumerate — rather than a looser population the agent gets to interpret. When the population is well-defined, the hand-witnessed number and the machine number can actually be compared. When it's loose, they drift, and you can't even tell.&lt;/p&gt;

&lt;p&gt;Both of these slow things down. That's the cost I'm deliberately paying for the measurements I'm willing to bet on.&lt;/p&gt;




&lt;h2&gt;
  
  
  What this didn't prove
&lt;/h2&gt;

&lt;p&gt;I want to keep this in proportion.&lt;/p&gt;

&lt;p&gt;This is one off-by-one, caught once, on one counting task. It doesn't prove AI agents can't count, or that they're untrustworthy in general, or that you should hand-verify every number an agent produces. For many low-stakes tallies, the agent's number may be good enough, and re-counting by hand would waste a perfectly good tool. I'm not arguing for paranoia.&lt;/p&gt;

&lt;p&gt;The claim I'd stand behind is narrow and conditional: &lt;em&gt;for measurements you intend to build further conclusions on, the final witness should be deterministic and, at least initially, human.&lt;/em&gt; The stakes set the bar. A throwaway internal count doesn't need this. A number you're going to let the system's self-assessment ride on does — because that's the number whose error compounds.&lt;/p&gt;

&lt;p&gt;I'd also flag the obvious tension: this doesn't scale by hand forever. The whole point of automating measurement was to &lt;em&gt;not&lt;/em&gt; sit and count things. So "a human witnesses it" is a phase, not a permanent state — it's how trust gets established before the automation runs on its own, not a vow to recount reality by hand every day. For now, I only start relaxing the human witness once the deterministic path and the automated path have matched across repeated runs, and once the inclusion rule is pinned tightly enough that both are counting the same population. Where exactly to graduate beyond that is something I'm still working out, and I don't have a clean rule for it yet.&lt;/p&gt;




&lt;h2&gt;
  
  
  The takeaway, stated honestly
&lt;/h2&gt;

&lt;p&gt;You can let an AI write the code. You can let it run the analysis. But if you also let the &lt;em&gt;same system that produced the answer&lt;/em&gt; be the final judge of whether that answer was correct, you've made the examinee the examiner, and the score gets hard to trust on its own. For the numbers that matter, something deterministic — and, while trust is still being earned, something human — has to be the last thing that looks.&lt;/p&gt;

&lt;p&gt;The agent was right about almost everything — off by one. The whole question is whether you find out about the one, and that depended entirely on whether I was willing to look myself.&lt;/p&gt;




&lt;p&gt;How do others draw this line? When you let an AI agent measure or aggregate something, do you verify it independently — and how do you decide which numbers earn the hand-check and which numbers you allow to pass without it? I leaned on doing it myself this time, but I know that doesn't scale, and I'd be glad to hear how people who've gone further handle the trust handoff.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Next in the series — the last of this run: we stopped chasing better models a while ago. This is about the moment we stopped trusting our own measurements too, and what we kept instead.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>softwareengineering</category>
      <category>devjournal</category>
    </item>
    <item>
      <title>The Gate Fired 198 Times. I Called It "Working."</title>
      <dc:creator>Joseph Yeo</dc:creator>
      <pubDate>Sun, 21 Jun 2026 17:39:52 +0000</pubDate>
      <link>https://dev.to/josephyeo/the-gate-fired-198-times-i-called-it-working-45fk</link>
      <guid>https://dev.to/josephyeo/the-gate-fired-198-times-i-called-it-working-45fk</guid>
      <description>&lt;h2&gt;
  
  
  Why a blocked count is not a success metric
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;Part of the ForgeFlow series — building a coding agent that runs its execution loop locally on an M5 Max, and writing down what actually breaks. Planning runs on Claude; code generation runs on a local model via Ollama, test-driven inside a Docker sandbox.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;I built a gate to block bad code. It blocked 198 pieces of code, and I took that number as evidence the gate was working well.&lt;/p&gt;

&lt;p&gt;Then I opened the blocked cases and read them one by one, checking each against the acceptance criteria for the task it came from. A large share of them weren't bad code. The gate had been wrong often enough that I could no longer read the block count as evidence it was working — it had been firing constantly, exactly as it was designed to fire, and I'd mistaken "it fires a lot" for "it's doing its job." Those are not the same statement.&lt;/p&gt;

&lt;p&gt;This is the second post in a short run about something I kept tripping over while building this agent: the things I use to verify my system can themselves be broken, and they tend to break in ways that look like success. The &lt;a href="https://dev.to/josephyeo/when-pytest-said-passed-it-was-lying-43p7"&gt;last post&lt;/a&gt; was about a test run that lied by passing. This one is about a gate that lied by blocking.&lt;/p&gt;




&lt;h2&gt;
  
  
  What the gate is, and why I trusted the count
&lt;/h2&gt;

&lt;p&gt;The agent works in a test-driven loop, and one step in that loop is a gate: before certain code is allowed through, the gate checks that it meets a standard, and if it doesn't, it blocks the code and sends the work back to be redone. The gate exists to keep weak or malformed attempts from getting committed.&lt;/p&gt;

&lt;p&gt;For a while, the gate's headline number was how many times it had blocked something: 198. I looked at that and felt good about it. The reasoning felt obvious: the gate is catching a lot of bad attempts, so the gate is valuable, so the system is healthier for having it. High block count, hard-working gate, fewer bad commits. Why look closer?&lt;/p&gt;

&lt;p&gt;That reasoning has a hole in it — and that hole is what this post is about.&lt;/p&gt;




&lt;h2&gt;
  
  
  Two claims I'd quietly merged
&lt;/h2&gt;

&lt;p&gt;When I went through the blocked cases — not the count, the actual cases — I found that a large share of what the gate had blocked was work that was fine. Not malformed, not weak. Legitimate attempts that happened to take a shape the gate didn't recognize, so it rejected them.&lt;/p&gt;

&lt;p&gt;I want to be precise about what "fine" means here, because "the gate was wrong" needs a ground truth or it's just my opinion. By ground truth I don't mean "I liked the code." I mean each step had explicit acceptance criteria: the targeted tests it was meant to pass, the behavior it was meant to produce, and the stated constraints for that step. A block was a &lt;em&gt;false positive&lt;/em&gt; when those criteria were already satisfied and the gate rejected the work anyway, for a property that wasn't part of the task's success condition.&lt;/p&gt;

&lt;p&gt;For example: a solution that passed every test it was supposed to pass, but got rejected because its internal structure didn't match what the gate's check expected — the same task-level behavior, in a representation the task itself never required. That isn't a question of taste. The work met the criteria; the gate said no for a reason that sat outside them.&lt;/p&gt;

&lt;p&gt;So the 198 was real. Every one of those blocks happened. What was false was the meaning I'd attached to it. I had collapsed two different claims into one:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;The gate fired.&lt;/strong&gt; (True. 198 times. Verifiable.)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The firing was justified.&lt;/strong&gt; (Never checked — and, it turned out, often not.)
"It blocked something" and "it was right to block that something" are independent facts. A gate can be extremely active and extremely wrong at the same time — and a &lt;em&gt;miscalibrated&lt;/em&gt; gate will tend to be both, because the same flaw that makes it reject good work also makes it reject a lot of it. The block count I'd treated as evidence of value is equally consistent with a gate that's simply trigger-happy. The block count alone can't separate a justified rejection from a false positive.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This seems like an easy mistake to make when building guards — linter rules, CI checks, validation layers, policy filters. At least it was in my case, and I suspect it's more common than we'd like to admit. The dashboard shows you &lt;em&gt;activity&lt;/em&gt;. Activity feels like protection. But activity is not the same as &lt;em&gt;correct&lt;/em&gt; activity, and the dashboard usually doesn't know the difference, so it shows you the comforting number and lets you supply the flattering interpretation.&lt;/p&gt;




&lt;h2&gt;
  
  
  "Then how did anything get through?"
&lt;/h2&gt;

&lt;p&gt;That's the fair question, and it's the one that finally made me look. If the gate was wrong most of the time, how did the system make any progress at all?&lt;/p&gt;

&lt;p&gt;The answer is uncomfortable: it made progress &lt;em&gt;despite&lt;/em&gt; the gate, not because of it. A typical failure looked like this. The first attempt satisfied the tests but used a structure the gate distrusted, so it got blocked. The retry didn't improve the behavior — it just reshaped the same behavior into a form the gate would accept. From the dashboard, that looked like "the gate forced an improvement." From the case review, it was adaptation to the gate. (When even that didn't work, I'd step in and wave the work through, because I could see it was fine — another quiet sign the gate wasn't earning its place.)&lt;/p&gt;

&lt;p&gt;That was the tell I'd missed. A gate doing real work makes the loop converge on &lt;em&gt;better&lt;/em&gt; code. Mine was making the loop converge on &lt;em&gt;gate-shaped&lt;/em&gt; code — which is not the same thing, and is sometimes worse. The retry didn't make the code more correct. It made it more acceptable to the gate.&lt;/p&gt;




&lt;h2&gt;
  
  
  How I now try to tell a real block from a false one
&lt;/h2&gt;

&lt;p&gt;Catching this forced me to write down what would actually distinguish a justified block from a noisy one. The count clearly wasn't it. What I landed on is a three-part check — not elegant, but it's caught things since, so I'll offer it as a working heuristic for this setup rather than a rule.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Look at the distribution of reasons, not the total.&lt;/strong&gt; I'd expect the block reasons to map to substantive defects, not to repeatedly trip on the same shallow surface feature regardless of whether the work is good. If they cluster on the latter, the gate is probably pattern-matching on the wrong thing instead of judging quality. (This is more useful for a broad quality gate than for a narrow, single-purpose check.)&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Watch what happens on retry.&lt;/strong&gt; This turned out to be the most useful signal. In my loop, a &lt;em&gt;justified&lt;/em&gt; block tended to make the work stick on the same underlying defect across retries, until that defect was actually addressed; a &lt;em&gt;false positive&lt;/em&gt; produced shape-shifting attempts that changed the surface without improving the behavior. It's a tendency, not a law — a model can wander even when the defect is real — but the shape of the retry sequence carried information the single block event didn't.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Check final convergence.&lt;/strong&gt; A justified block should eventually &lt;em&gt;resolve&lt;/em&gt;: the work gets rewritten, the real problem gets fixed, and it passes on its own merits. If blocked work never converges — or only "passes" once you weaken the gate — then either the gate was wrong, or it was right and your loop can't act on it. Both are problems, and both are invisible if you only count how many times it fired.&lt;/p&gt;

&lt;p&gt;None of these is a clean pass/fail on its own. Together they let me ask the question I'd skipped — &lt;em&gt;was this block justified?&lt;/em&gt; — instead of reading the answer off the fact that a block happened.&lt;/p&gt;




&lt;h2&gt;
  
  
  The deeper version of the mistake
&lt;/h2&gt;

&lt;p&gt;There's a more general trap underneath this, and I want to name it plainly, because I fell into it without noticing.&lt;/p&gt;

&lt;p&gt;A test that only checks "the gate blocks bad input" is testing the easy half. The hard half is: does the gate let through good input that simply looks unusual? If you only ever feed a guard the inputs it's supposed to reject, of course it rejects them, and of course the tests pass — but you've proven nothing about its false-positive behavior, which is exactly where mine was failing. The gate's own tests were green for the same reason the gate looked healthy: I'd only ever asked it the flattering question.&lt;/p&gt;

&lt;p&gt;So now, when I test a gate, I deliberately include cases that are legitimate but oddly shaped — valid work in a form the gate might naively distrust — and check that it lets them through. In this case, the negative cases (reject the bad stuff) were the easier half. The risk I'd under-tested was the good stuff in unfamiliar clothing.&lt;/p&gt;




&lt;h2&gt;
  
  
  What this didn't prove
&lt;/h2&gt;

&lt;p&gt;I don't want to inflate this into "all gates are bad" or "block counts are meaningless." Neither is true.&lt;/p&gt;

&lt;p&gt;Gates can earn their keep. A well-calibrated one really does stop real problems, and the block count is a perfectly good &lt;em&gt;operational&lt;/em&gt; signal — useful for noticing that something is happening, or spotting a sudden spike. The mistake wasn't tracking the count. It was treating the count as evidence of &lt;em&gt;correctness&lt;/em&gt; when it's only evidence of &lt;em&gt;activity&lt;/em&gt;. Those are different axes, and I'd conflated them.&lt;/p&gt;

&lt;p&gt;I'd also flag that my three-part check is shaped by this particular system — a test-driven loop where blocked work gets automatically retried, so "watch the retries" is even available to me as a signal. If your setup doesn't produce that kind of trajectory, parts of this won't transfer. I'm offering it as something that worked in one place, not a general theorem about guards. I've been wrong about generality before in this project, so I'm holding it loosely.&lt;/p&gt;




&lt;h2&gt;
  
  
  The takeaway, stated honestly
&lt;/h2&gt;

&lt;p&gt;A guard blocking something tells you it's &lt;em&gt;active&lt;/em&gt;. It does not tell you it's &lt;em&gt;right&lt;/em&gt;. Those are separate facts, and the gap between them is where a confident-looking gate can quietly turn into a noise machine that punishes good work and calls it protection.&lt;/p&gt;

&lt;p&gt;If you run gates, linters, validators, policy filters — anything that stops things — it's worth auditing a &lt;em&gt;sample of what it blocked&lt;/em&gt;, not just the total it blocked. The total can easily look like diligence. The sample is where you find out whether the diligence was real.&lt;/p&gt;




&lt;p&gt;I'm curious how others handle this. If you operate a gate or a strict CI rule, do you ever sample its blocks to check they were justified — and if so, how do you decide "justified" without it turning subjective? I worked out a rough method for my case, but it leans on my system's particular shape, and I'd like to hear how it's done elsewhere.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Next in the series: I asked an AI agent to count something for me. It said 12. The real number was 13 — and that one-off gap changed a rule I now follow.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>softwareengineering</category>
    </item>
    <item>
      <title>When pytest Said "Passed," It Was Lying</title>
      <dc:creator>Joseph Yeo</dc:creator>
      <pubDate>Sat, 20 Jun 2026 13:05:33 +0000</pubDate>
      <link>https://dev.to/josephyeo/when-pytest-said-passed-it-was-lying-43p7</link>
      <guid>https://dev.to/josephyeo/when-pytest-said-passed-it-was-lying-43p7</guid>
      <description>&lt;h2&gt;
  
  
  How a polluted virtual environment made my green tests meaningless
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;Part of the ForgeFlow series — building a coding agent that runs its execution loop locally on an M5 Max, and writing down what actually breaks. Planning runs on Claude; code generation runs on a local model via Ollama, test-driven inside a Docker sandbox.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;For a few days, I made decisions on top of a number that wasn't true.&lt;/p&gt;

&lt;p&gt;The number was &lt;code&gt;186 passed&lt;/code&gt;. It came out of pytest, green, at the bottom of the terminal, the way it had dozens of times before. I trusted it the way you trust a number that has never been wrong before. Then I found out the run had been measured inside the wrong environment, and the green had very little to do with the code I thought I was checking.&lt;/p&gt;

&lt;p&gt;To be fair to the tool: pytest wasn't wrong. It answered exactly the question I handed it — it just wasn't the question I meant to ask. This post is about that gap. Not a bug in a test, but a bug in &lt;em&gt;how I measured the tests&lt;/em&gt;. It turned out to be one of the more uncomfortable lessons in the project so far, because it sat underneath everything else. If the floor is tilted, every measurement you take on top of it inherits the tilt, and you don't see it, because the floor looks like the floor.&lt;/p&gt;

&lt;p&gt;I'm writing it down mostly because I suspect I'm not the only person who has trusted a green checkmark that didn't earn it.&lt;/p&gt;




&lt;h2&gt;
  
  
  The setup: a baseline I checked constantly
&lt;/h2&gt;

&lt;p&gt;ForgeFlow is a coding agent that runs a test-driven loop. Plan, write a failing test, write code, run the tests, decide what happened, repeat. Because so much of the system's behavior is judged by "did the tests pass," I keep a &lt;strong&gt;baseline&lt;/strong&gt;: a known set of test files that should report a known set of numbers. Before and after almost any change to the engine, I re-run the baseline and compare. If the counts move when they shouldn't, something is wrong.&lt;/p&gt;

&lt;p&gt;One thing to be precise about, because it matters later: the agent &lt;em&gt;executes the code it generates&lt;/em&gt; inside a Docker sandbox, but this baseline — the engine's own test suite — I was running from my host shell. Two different executions. The sandbox was fine. The host shell was the problem.&lt;/p&gt;

&lt;p&gt;The baseline is the closest thing the project has to a source of truth about its own health. That's exactly why this hurt.&lt;/p&gt;

&lt;p&gt;One afternoon I was moving between two things on the same machine — the agent's own codebase, and an unrelated project I'd been poking at earlier. I ran the baseline. Green. &lt;code&gt;186 passed&lt;/code&gt;. I noted it, moved on, and built the next decision on top of it.&lt;/p&gt;

&lt;p&gt;What I didn't notice was a single line of state that had carried over from the earlier work.&lt;/p&gt;




&lt;h2&gt;
  
  
  What actually happened
&lt;/h2&gt;

&lt;p&gt;I'd left a different project's virtual environment active.&lt;/p&gt;

&lt;p&gt;That's the whole bug, mechanically. The shell still had another project's &lt;code&gt;VIRTUAL_ENV&lt;/code&gt; set, so when I ran pytest, it resolved &lt;code&gt;pytest&lt;/code&gt; through &lt;em&gt;that&lt;/em&gt; environment, and Python resolved imports against that environment's installed packages.&lt;/p&gt;

&lt;p&gt;Here's the question a careful reader asks immediately: if it was the wrong environment, why didn't it just fail? Why no &lt;code&gt;ModuleNotFoundError&lt;/code&gt;, no loud red collection error?&lt;/p&gt;

&lt;p&gt;Because nothing was missing. The polluted environment happened to have the same packages installed — only at different versions. So nothing errored out. The tests collected, ran, and passed; they just passed against a version matrix the baseline doesn't assume. And that's the genuinely unsettling part: &lt;strong&gt;if a dependency had been entirely absent, I'd have gotten a loud error and caught it in seconds.&lt;/strong&gt; The danger was precisely that everything was &lt;em&gt;present&lt;/em&gt; — present and subtly wrong. Wrong in the one way that doesn't announce itself.&lt;/p&gt;

&lt;p&gt;The problem is that "green" had stopped meaning what I read it as. I read it as &lt;em&gt;"the code is correct in the environment it's meant to run in."&lt;/em&gt; What it actually meant was &lt;em&gt;"the code passed in whatever environment happened to be active."&lt;/em&gt; Those are different sentences. For a few days I couldn't tell them apart, because the terminal prints the same word for both.&lt;/p&gt;

&lt;p&gt;Here's the part worth sitting with: &lt;strong&gt;nothing failed.&lt;/strong&gt; A failing test is a gift — it's loud, it points at itself, you go fix it. This didn't fail. It passed, and the passing was the problem. The signal I rely on to catch mistakes was itself the mistake, wearing the costume of everything being fine.&lt;/p&gt;




&lt;h2&gt;
  
  
  "The tests pass" and "the measurement is honest" are different claims
&lt;/h2&gt;

&lt;p&gt;When I finally caught it — by comparing against a run from a fresh shell using the project's intended environment, and noticing the counts didn't line up — I reduced the lesson to a sentence I've kept since:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Whether the tests pass and whether the &lt;em&gt;measurement of the tests is trustworthy&lt;/em&gt; are two separate questions, and I had been treating them as one.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Almost all of my testing discipline had been aimed at the first question. I had careful tests. What I didn't have was anything checking the second — the integrity of the act of measuring. The environment the measurement runs in is an input to the result, and I'd been treating it as a constant when it was actually a variable I'd left lying around.&lt;/p&gt;

&lt;p&gt;It's easy to invest heavily in test &lt;em&gt;correctness&lt;/em&gt; while leaving test &lt;em&gt;measurement integrity&lt;/em&gt; implicit — to treat it as the environment's job, or the tooling's job, rather than something to check directly. Plenty of teams do handle it, with lockfiles, containerized test runs, hermetic builds. I just wasn't one of them at this layer, for this particular command, on this particular day.&lt;/p&gt;




&lt;h2&gt;
  
  
  What I changed
&lt;/h2&gt;

&lt;p&gt;Two things, deliberately small.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A guard before the measurement, not just inside it.&lt;/strong&gt; The cheapest fix is a single check that runs before the baseline: confirm the environment is the one I think it is. In my case, simply asserting that no foreign virtual environment was active would have caught it — the testing equivalent of checking the floor before you measure the wall. It's almost embarrassingly simple, and it would have caught this in one second. (A stricter guard wouldn't stop at the &lt;code&gt;VIRTUAL_ENV&lt;/code&gt; variable; it would also check &lt;code&gt;sys.executable&lt;/code&gt;, the resolved &lt;code&gt;pytest&lt;/code&gt; path, and the expected project root. The variable was the obvious giveaway here, but it's the weakest of the four.)&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A separate set of checks for the measurement itself.&lt;/strong&gt; Beyond the one-line guard, I added a small, dedicated set of tests whose only job is to protect the &lt;em&gt;invariants of how I measure&lt;/em&gt; — not the features, the measurement. They're counted separately from the normal baseline on purpose, so they can't be quietly folded into the same number they're supposed to be watching. The exact count is secondary. What matters is that "is my measurement honest" became something the system checks for me, instead of something I assume.&lt;/p&gt;

&lt;p&gt;Neither of these is clever. That's sort of the lesson. The failure wasn't subtle once I saw it; it was invisible only because I'd never thought to look there.&lt;/p&gt;




&lt;h2&gt;
  
  
  What this didn't prove
&lt;/h2&gt;

&lt;p&gt;I want to be careful not to inflate this into a grand principle.&lt;/p&gt;

&lt;p&gt;This is one incident, on one machine, caused by one careless bit of leftover state. It doesn't prove that everyone's test suites are secretly lying, and it doesn't prove you need an elaborate measurement-verification layer. For a small throwaway script, a clean shell and a moment of attention is the entire fix, and the machinery would be overkill.&lt;/p&gt;

&lt;p&gt;To be clear about which part scales: environment hygiene matters at any size — it's the &lt;em&gt;guards and dedicated checks&lt;/em&gt; that scale with how much you're betting on the number. I'm betting a lot on mine; the agent makes real decisions off these signals. So for me the machinery earned its place. The hygiene would have been worth it regardless of project size.&lt;/p&gt;

&lt;p&gt;I'm also aware the "fix" mostly moves the trust down one level. Now I trust the guard. If the guard is wrong, I'm back where I started. There's no absolute bottom here — just a level low enough that I'm willing to stop and call it ground. I picked one. I could be wrong about whether it's low enough.&lt;/p&gt;




&lt;h2&gt;
  
  
  The takeaway, stated honestly
&lt;/h2&gt;

&lt;p&gt;If I had to compress it: a green test run answers "did the code pass?" It does &lt;em&gt;not&lt;/em&gt; answer "did I measure that in the environment I meant to?" The second question has its own failure mode, and because the failure mode is &lt;em&gt;passing&lt;/em&gt;, your normal instincts — chase the red, fix what's loud — never fire.&lt;/p&gt;

&lt;p&gt;The verification pyramid most of us picture has tests at the bottom. I'd now put one more layer underneath it: &lt;em&gt;the environment the tests run in.&lt;/em&gt; When that layer shifts, every green light above it is reporting on an environment you're not actually in.&lt;/p&gt;




&lt;p&gt;I'd like to know how other people handle this. Do you guard your test environment explicitly, or rely on convention and attention? Have you been bitten by a &lt;em&gt;passing&lt;/em&gt; result that turned out to be measured wrong — and if so, how did you finally catch it? I caught mine by luck and a mismatched count. I'd rather not depend on luck next time, and I suspect some of you have better answers than I do.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Next in the series: a quality gate in the same system blocked code 198 times — and why I was wrong to call that "working."&lt;/em&gt;&lt;/p&gt;

</description>
    </item>
    <item>
      <title>We Spent Six Sessions Fixing One Task. The Problem Was Six Tasks.</title>
      <dc:creator>Joseph Yeo</dc:creator>
      <pubDate>Tue, 02 Jun 2026 11:42:06 +0000</pubDate>
      <link>https://dev.to/josephyeo/we-spent-six-sessions-fixing-one-task-the-problem-was-six-tasks-385d</link>
      <guid>https://dev.to/josephyeo/we-spent-six-sessions-fixing-one-task-the-problem-was-six-tasks-385d</guid>
      <description>&lt;p&gt;&lt;em&gt;This is Part 9 of the ForgeFlow series. &lt;a href="https://dev.to/josephyeo/77-rules-later-what-graduating-our-first-stack-actually-looked-like-2o3k"&gt;Part 8: 77 Rules Later&lt;/a&gt; ended on a question we couldn't answer at the time: can a rule-based agent system keep growing without becoming harder to reason about than the model it was built to constrain? This post is about one small episode that pushed us toward an answer — not a clean one, but a useful one. It's also a record of getting a diagnosis wrong for several sessions in a row, and what finally corrected it.&lt;/em&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Quick terms for new readers:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;FC&lt;/strong&gt; = Failure Catalog entry (a documented failure pattern)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;CL&lt;/strong&gt; = Crystallized Lesson (a testable design rule derived from repeated failures)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;critical_rules&lt;/strong&gt; = a block of rules injected into the model's prompt for a given project&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;nightrun&lt;/strong&gt; = an overnight batch that re-runs projects so we can measure behavior over many runs&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ForgeFlow&lt;/strong&gt; = a fully local, TDD-based autonomous coding system running on Apple Silicon&lt;/li&gt;
&lt;/ul&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  The failure we kept "fixing"
&lt;/h2&gt;

&lt;p&gt;For several sessions, we had a recurring failure in one project that we treated as a single, local problem.&lt;/p&gt;

&lt;p&gt;The symptom: in async tests, the model would construct an HTTP test client directly instead of using the shared fixture we'd set up. That shared fixture is what installs the dependency override pointing the app at the test database. By hand-rolling its own client, the model skipped the fixture entirely — so the override was never applied, and the test hit an unconfigured database and failed with a missing-table error.&lt;/p&gt;

&lt;p&gt;(The underlying trigger, for those on the same stack: &lt;code&gt;httpx&lt;/code&gt; 0.28 removed the deprecated &lt;code&gt;app=&lt;/code&gt; shortcut from &lt;code&gt;Client&lt;/code&gt;/&lt;code&gt;AsyncClient&lt;/code&gt;. The supported pattern is now &lt;code&gt;transport=ASGITransport(app=app)&lt;/code&gt;. Our conftest fixture already used the current &lt;code&gt;ASGITransport&lt;/code&gt; approach and was correct — the &lt;code&gt;app=&lt;/code&gt; argument on &lt;code&gt;ASGITransport&lt;/code&gt; itself is still valid. The model just wasn't using the fixture; it kept hand-rolling the client the old, removed way. Why it preferred the old pattern is something we can only guess at — likely the weight of older examples — and we didn't try to prove it.)&lt;/p&gt;

&lt;p&gt;Each time it surfaced, we did the natural thing. We looked at the task where it appeared, and we wrote a rule for &lt;em&gt;that task&lt;/em&gt;. Next run, it would seem quieter there — and then show up somewhere else. We'd note it, half-suspect it was the same thing, and move on to whatever else the run surfaced.&lt;/p&gt;

&lt;p&gt;What we were doing, in effect, was treating a recurring pattern as a series of unrelated incidents. We never asked the obvious question: &lt;em&gt;how often does this actually happen, and where?&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The measurement we should have taken earlier
&lt;/h2&gt;

&lt;p&gt;The thing that broke the loop wasn't a better model or a smarter rule. It was a measurement we hadn't bothered to take.&lt;/p&gt;

&lt;p&gt;We wrote a small script to walk back through our run history and count occurrences of the failing pattern across every recorded run — not by reconstructing it from summary tables, but by grepping the raw captured error text directly. (We'd learned in an earlier session that our summary-level deduplication could merge distinct failures under one signature, so for this we went to the raw text instead.) In practice that meant grepping the archived &lt;code&gt;stderr&lt;/code&gt; for the exact failing construction across every run directory, rather than trusting the rolled-up signatures.&lt;/p&gt;

&lt;p&gt;In the archived runs we checked, the count came back:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;16 occurrences&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;across &lt;strong&gt;6 distinct tasks&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;in &lt;strong&gt;3 different projects&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;spread over &lt;strong&gt;7 separate run sessions&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This count made our earlier framing untenable. We had been writing single-task rules for something that was happening across six tasks in three projects. Each of those per-task fixes had been addressing, at most, one of the sixteen occurrences we eventually counted. The measurement didn't just refine our picture — for this failure pattern, it showed our framing had been wrong in kind, not just in degree.&lt;/p&gt;

&lt;p&gt;It's a little uncomfortable to write that down. But that's the part worth keeping: the wrong number wasn't in the model's output. It was in our own estimate of how widespread the problem was, and we'd carried that estimate for several sessions without checking it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The prescription got smaller, not bigger
&lt;/h2&gt;

&lt;p&gt;Here's the part that surprised us most.&lt;/p&gt;

&lt;p&gt;Once we saw "6 tasks, 3 projects," the instinct might be to write six fixes — one hardened rule per affected task. Given how we'd been responding up to that point, that's roughly the path we were on.&lt;/p&gt;

&lt;p&gt;Instead, the data pointed the other way. If one pattern was appearing across many tasks and projects, the better place to address it wasn't any individual task — it was the project-wide rule block. We added &lt;strong&gt;one line&lt;/strong&gt; to the &lt;code&gt;critical_rules&lt;/code&gt; for the affected projects: a single instruction telling the model not to construct the test client directly, and to use the fixture instead (taking the rule block from 20 lines to 21).&lt;/p&gt;

&lt;p&gt;One rule addressed a pattern that, on our prior trajectory, could easily have become six separate task-level patches. This was a small, concrete instance of something we keep seeing on this project: when we measure a problem's actual scope more carefully, the fix tends to get narrower, not wider. When you don't know the shape of a problem, you tend to over-prescribe locally and under-address globally. Measuring the shape let us do less.&lt;/p&gt;

&lt;p&gt;We then ran a verification batch over the affected projects. The grep count for the pattern came back at &lt;strong&gt;zero&lt;/strong&gt; across those runs. Two of the projects involved — a small gallery API and a small library API — met the per-project completion criteria on the verifying runs, so we marked them as "graduated" in a narrow, project-level sense.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Note on "graduated":&lt;/strong&gt; Part 8 used this word for an entire stack. Here it means something much smaller — these individual projects met their completion criteria on the runs we executed. It is not a claim about the stack, and not a claim that these projects will never fail again. The underlying numbers were modest: pass rates in the roughly 38–85% range across individual runs. What changed wasn't a jump to near-perfect runs — it was that this particular structural failure stopped appearing. We're reporting a state we observed, not a guarantee.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  What this didn't prove
&lt;/h2&gt;

&lt;p&gt;A few limits, because the result is smaller than it might sound.&lt;/p&gt;

&lt;p&gt;The zero is a zero &lt;em&gt;on the runs we executed&lt;/em&gt;, for &lt;em&gt;this specific pattern&lt;/em&gt;. It's evidence the rule is doing its job in the tested scope, not proof the pattern is gone for good. A different project, a different &lt;code&gt;httpx&lt;/code&gt; version, or a different phrasing of the same task could surface it again.&lt;/p&gt;

&lt;p&gt;The measurement approach itself has a known weakness we worked around rather than solved: it reads raw error text, which is reliable for exact-pattern counting but says nothing about &lt;em&gt;why&lt;/em&gt; each occurrence happened. For this pattern — a single, well-understood cause — that was fine. For a fuzzier failure, raw-text counting would undercount or overcount, and we'd need something better.&lt;/p&gt;

&lt;p&gt;And the broader idea ("measure before you prescribe") is a working heuristic from repeated experience on this project, not something we've established rigorously. We'd genuinely like to know where it breaks down for other people.&lt;/p&gt;

&lt;h2&gt;
  
  
  Beyond the coding loop
&lt;/h2&gt;

&lt;p&gt;The same week, an unrelated problem came up — a structural issue in how some of our project directories were tracked in version control. An earlier note had flagged it as a likely large cleanup, the kind of thing you budget a careful session for.&lt;/p&gt;

&lt;p&gt;Before touching anything, we applied the same lesson the measurement script had just taught us: measure first. A few read-only checks showed the actual scope was much smaller than the earlier flag assumed — most of what looked like a structural defect turned out to be a single missing ignore-rule and one stale index entry, and the fix was a handful of non-destructive steps rather than a restructuring. I'm including this not as a second result but as an observation: in both cases the failure mode was the same — acting on an estimate instead of a measurement — and in both cases, once we measured, the prescription shrank. Two instances isn't a trend, but it was enough to make "measure the scope before writing the fix" a step we now try to take on purpose rather than when we happen to remember.&lt;/p&gt;

&lt;h2&gt;
  
  
  Back to Part 8's question
&lt;/h2&gt;

&lt;p&gt;Part 8 ended by asking whether a growing rule set could stay manageable. This episode is one data point toward a tentative answer: rule growth seems more controllable when the &lt;em&gt;decision to add a rule&lt;/em&gt; is driven by measured scope rather than by where a failure last happened to appear. Six task-level rules would have grown the rule set faster and addressed the problem worse. One measured rule did more with less.&lt;/p&gt;

&lt;p&gt;That's not a method, and it's certainly not a solution to the interaction-effect problem from Part 8. It's a habit we're trying to form: before crystallizing a new rule, spend the few minutes it takes to count how often and where the underlying failure actually occurs. Sometimes that collapses six patches into one rule. Other times — we assume, though we haven't hit this yet — it will reveal that what looked like one problem is really several, and we'll need more rules, not fewer. Either way, the rule follows the measurement.&lt;/p&gt;




&lt;h2&gt;
  
  
  Series Links
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://dev.to/josephyeo/i-built-a-local-ai-coding-agent-on-m5-max-128gb-it-failed-164-times-before-passing-35-tests-2fgj"&gt;Part 1: 164 Failures Before 35 Tests&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dev.to/josephyeo/we-didnt-migrate-from-n8n-to-python-because-n8n-failed-k9j"&gt;Part 2: We Didn't Migrate from n8n Because n8n Failed&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dev.to/josephyeo/the-determinism-war-why-we-stopped-chasing-better-models-3c21"&gt;Part 3: The Determinism War&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dev.to/josephyeo/the-information-design-gap-why-our-ai-agent-was-coding-blind-4p8o"&gt;Part 4: The Information Design Gap&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dev.to/josephyeo/dcr-wasnt-enough-why-ai-coding-agents-also-need-information-quality-1da4"&gt;Part 5: DCR Wasn't Enough&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dev.to/josephyeo/the-bug-wasnt-in-the-model-lessons-from-9-local-ai-coding-agent-projects-18aa"&gt;Part 6: The Bug Wasn't in the Model&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dev.to/josephyeo/the-file-modification-boundary-we-found-after-12-forgeflow-projects-3m01"&gt;Part 7: The File Modification Boundary&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dev.to/josephyeo/77-rules-later-what-graduating-our-first-stack-actually-looked-like-2o3k"&gt;Part 8: 77 Rules Later&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;ForgeFlow runs on a MacBook Pro M5 Max 128GB. Planning uses Claude (cloud API). Execution is fully local — Qwen3-Coder-Next 45GB via Ollama, gemma4:26b for QA, Docker sandbox, no API calls during the coding loop. The methodology and failure data are shared in this series.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;If you're running your own local agents or failure-catalog systems: have you caught yourself prescribing locally for a problem that turned out to be project-wide — and what finally made you measure it? Logs, test signatures, prompts, generated diffs? The comments are open.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>devjournal</category>
      <category>softwareengineering</category>
    </item>
    <item>
      <title>77 Rules Later: What Graduating Our First Stack Actually Looked Like</title>
      <dc:creator>Joseph Yeo</dc:creator>
      <pubDate>Mon, 25 May 2026 08:19:22 +0000</pubDate>
      <link>https://dev.to/josephyeo/77-rules-later-what-graduating-our-first-stack-actually-looked-like-2o3k</link>
      <guid>https://dev.to/josephyeo/77-rules-later-what-graduating-our-first-stack-actually-looked-like-2o3k</guid>
      <description>&lt;p&gt;&lt;em&gt;This is Part 8 of the ForgeFlow series. &lt;a href="https://dev.to/josephyeo/the-file-modification-boundary-we-found-after-12-forgeflow-projects-3m01"&gt;Part 7: The File Modification Boundary&lt;/a&gt; documented the constraint that changed how we structure tasks: every autonomous task target should be a new file. We ended Part 7 at 12 projects, roughly 52 failure patterns, and 71 design rules. Part 7 closed with an open question: "Project 13 will be the first real test of whether CL-071 holds under normal conditions."&lt;/em&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Quick terms for new readers:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;FC&lt;/strong&gt; = Failure Catalog entry (a documented failure pattern)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;CL&lt;/strong&gt; = Crystallized Lesson (a testable design rule derived from repeated failures)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;DEADLOCK&lt;/strong&gt; = the system gives up after repeated identical failures&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ForgeFlow&lt;/strong&gt; = a fully local, TDD-based autonomous coding system running on Apple Silicon&lt;/li&gt;
&lt;/ul&gt;
&lt;/blockquote&gt;




&lt;p&gt;Part 7 ended with a hypothesis and a bet.&lt;/p&gt;

&lt;p&gt;The hypothesis: CL-071 (every task targets a new file, never modifies an existing one) might reduce or remove the dominant failure mode we'd been observing. The bet: we'd set formal graduation criteria and run projects until we met them — or discovered why we couldn't.&lt;/p&gt;

&lt;p&gt;We ran five more projects (with one intermediate rerun included in the data). On the seventeenth — a blog API with 14 tasks — all 33 tests passed without intervention or deadlock, completing in approximately 12 minutes.&lt;/p&gt;

&lt;p&gt;This post is about the five projects between that hypothesis and this result, what the graduation criteria actually measured, and the failure that appeared &lt;em&gt;after&lt;/em&gt; we thought we'd addressed all the known ones.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Graduation Criteria
&lt;/h2&gt;

&lt;p&gt;Before results, here's what we were measuring. We didn't want "it worked once" to count as graduation. We defined four conditions, all of which had to hold on a qualifying run:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Criterion&lt;/th&gt;
&lt;th&gt;Threshold&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;First-run pass rate (tasks passing on the first TDD cycle, no retry)&lt;/td&gt;
&lt;td&gt;≥ 85%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;New FC yield per project&lt;/td&gt;
&lt;td&gt;≤ 2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Repeat FC rate (previously solved patterns recurring)&lt;/td&gt;
&lt;td&gt;≤ 5%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Teacher escalation (human operator interventions mid-task)&lt;/td&gt;
&lt;td&gt;Decreasing trend&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The logic: a graduated stack should show repeatable autonomous recovery within the tested scope (criterion 1), stop producing novel failure patterns at a high rate (criterion 2), not regress on already-solved problems (criterion 3), and require less human involvement over time (criterion 4).&lt;/p&gt;

&lt;p&gt;We chose 85% rather than 100% for the pass rate deliberately. Occasional retries are expected behavior in a TDD loop — in ForgeFlow's architecture, the system is designed to recover from them. What we track is whether it recovers autonomously.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Five-Project Path
&lt;/h2&gt;

&lt;p&gt;Here's the longitudinal data from Part 7's endpoint (project 12) through the graduation run. Note: this table tracks the &lt;em&gt;autonomous pass rate&lt;/em&gt; — tasks that eventually passed without human intervention, including retries. The graduation criterion uses the stricter &lt;em&gt;first-run pass rate&lt;/em&gt; (no retries), which we measured separately for the qualifying run.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;#&lt;/th&gt;
&lt;th&gt;Project&lt;/th&gt;
&lt;th&gt;Tasks&lt;/th&gt;
&lt;th&gt;Autonomous Pass Rate&lt;/th&gt;
&lt;th&gt;New FCs&lt;/th&gt;
&lt;th&gt;CL Count (at time)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;13&lt;/td&gt;
&lt;td&gt;comment-api&lt;/td&gt;
&lt;td&gt;12&lt;/td&gt;
&lt;td&gt;83%&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;~72&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;14&lt;/td&gt;
&lt;td&gt;order-api&lt;/td&gt;
&lt;td&gt;16&lt;/td&gt;
&lt;td&gt;56%&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;~74&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;15&lt;/td&gt;
&lt;td&gt;recipe-api&lt;/td&gt;
&lt;td&gt;14&lt;/td&gt;
&lt;td&gt;57%&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;~75&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;16&lt;/td&gt;
&lt;td&gt;bookmark-api v2&lt;/td&gt;
&lt;td&gt;12&lt;/td&gt;
&lt;td&gt;83%&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;~76&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;16.5&lt;/td&gt;
&lt;td&gt;catalog-api-v2&lt;/td&gt;
&lt;td&gt;12&lt;/td&gt;
&lt;td&gt;83%&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;~76&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;17&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;blog-api&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;14&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;100%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;77&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The trajectory wasn't smooth. Projects 14 and 15 dropped below 60%. Then it recovered. In this sequence, plateaus tended to expose a new failure category; the system dipped, the failure got crystallized into a rule, and the next project incorporated the fix.&lt;/p&gt;

&lt;p&gt;What changed between project 15 (57%) and project 17 (100%) was not a model upgrade or an engine rewrite. It was three additional design rules, each derived from a specific failure we observed and diagnosed.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Dip: What Went Wrong on Projects 14 and 15
&lt;/h2&gt;

&lt;p&gt;Projects 14 (order-api) and 15 (recipe-api) both hovered around 56–57% autonomous pass rate. The failures clustered around a few patterns:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Route endpoint isolation.&lt;/strong&gt; Tasks that bundled multiple endpoints into a single file — GET list and GET detail in the same route module — showed a notably higher failure rate than single-endpoint tasks. The outputs showed scope-related failures: given two endpoints to implement, the model would sometimes complete one and leave the other as a stub, or attempt both and introduce inconsistencies.&lt;/p&gt;

&lt;p&gt;We already had CL-043 (one task, one endpoint) from Part 6. But we'd been applying it loosely — allowing two closely related endpoints to share a task. Projects 14 and 15 showed us that "closely related" was too vague for this local execution loop. The rule needed to be absolute: one endpoint, one file, one task.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Import specification gaps.&lt;/strong&gt; Route tasks that didn't explicitly list every required import in their task description had a high failure rate. The model would guess import paths, often incorrectly. CL-072 crystallized this: every route task description must include a complete "Required imports" block. For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;Required&lt;/span&gt; &lt;span class="n"&gt;imports&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;fastapi&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;APIRouter&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Depends&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;sqlalchemy.ext.asyncio&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;AsyncSession&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;app.database&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;get_db&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;app.schemas.author&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;AuthorCreate&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;AuthorRead&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Decimal type mismatches.&lt;/strong&gt; In project 16.5 (catalog-api-v2), a product model with a &lt;code&gt;Numeric(10,2)&lt;/code&gt; price column exposed a subtle testing issue. The model wrote assertions comparing float literals to SQLAlchemy Decimal values — and &lt;code&gt;999.99 != Decimal('999.99')&lt;/code&gt; in Python. CL-076 captured this: any Numeric column test must use Decimal comparisons.&lt;/p&gt;

&lt;p&gt;In our diagnosis, these looked less like model-capability failures and more like specification-precision failures — cases where the PRD left enough ambiguity for a 45GB quantized model to make a reasonable-but-wrong choice.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Failure We Didn't Expect: FC-074
&lt;/h2&gt;

&lt;p&gt;Project 17 (blog-api) was designed as the graduation attempt. We applied all 76 existing rules. The PRD passed our automated validator (50 checks passed, 0 failures). We expected fewer known-pattern failures.&lt;/p&gt;

&lt;p&gt;The first three attempts all failed on the very first task — creating the Author model. Same error each time: &lt;code&gt;red_apply_empty&lt;/code&gt; — the engine's signal that the RED-phase output contained implementation code rather than a test.&lt;/p&gt;

&lt;p&gt;Here's what happened, step by step:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Our setup script created a minimal model stub file — just the class name and primary key column. This was standard practice per CL-066 ("stubs should be PK-only").&lt;/li&gt;
&lt;li&gt;Before the RED phase (test generation), the engine runs FC-060 cleanup: it deletes the target implementation file so the model writes it fresh.&lt;/li&gt;
&lt;li&gt;FC-060 deleted the stub.&lt;/li&gt;
&lt;li&gt;The model didn't need the file to exist at generation time — the surrounding task context still described enough of the intended model structure (via data_models in the PRD and conftest import references) that it produced implementation code during RED instead of a test.&lt;/li&gt;
&lt;li&gt;The engine detected this as a scope violation and triggered &lt;code&gt;red_apply_empty&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Three retries. Same result each time.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;We called this FC-074: the interaction between two previously validated rules (CL-066: keep stubs minimal, and FC-060: clean target files before RED) producing a new failure when combined.&lt;/p&gt;

&lt;p&gt;This is worth pausing on. FC-074 wasn't a gap in any single rule. It was an &lt;em&gt;interaction effect&lt;/em&gt; — two rules that had each been validated independently across multiple projects, producing a failure only in a specific sequence of operations.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Rule&lt;/th&gt;
&lt;th&gt;Behavior in isolation&lt;/th&gt;
&lt;th&gt;Combined behavior&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;CL-066&lt;/td&gt;
&lt;td&gt;Minimal stubs reduce over-complete-stub failures&lt;/td&gt;
&lt;td&gt;Creates a target file before RED&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;FC-060&lt;/td&gt;
&lt;td&gt;Deletes implementation target before RED to ensure clean state&lt;/td&gt;
&lt;td&gt;Removes the stub CL-066 created&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Combined&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;RED sees a missing target but enough context to generate implementation instead of a test&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h2&gt;
  
  
  The Fix: Stop Creating Stubs
&lt;/h2&gt;

&lt;p&gt;The first instinct was to adjust the prompt wording — tell the model more explicitly to write a test, not an implementation. We tried that. Same failure. Prompt changes alone didn't resolve it; file-state became the stronger hypothesis.&lt;/p&gt;

&lt;p&gt;The second instinct was to refine the stub. But we diagnosed the stub's existence as the likely trigger: FC-060 deleted it, and the residual context information was enough to derail the RED phase.&lt;/p&gt;

&lt;p&gt;The third attempt was the simplest: don't create the stub at all.&lt;/p&gt;

&lt;p&gt;CL-077: Setup scripts must not create model stub files. Model files are created from scratch by the task that implements them. The conftest wraps model imports in try/except so that earlier tasks can run before the model file exists:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;app.models.author&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Author&lt;/span&gt;
&lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="nb"&gt;ImportError&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;Author&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This inverted an assumption we'd held across the previous 16 project iterations. We'd operated under the belief that providing a stub — even a minimal one — helped the model by giving it a starting point. FC-074 suggested that in our current engine architecture, the stub &lt;em&gt;hurt&lt;/em&gt; by creating a state that the cleanup logic couldn't handle cleanly.&lt;/p&gt;

&lt;p&gt;After applying CL-077, the same blog-api project ran all 14 tasks to completion. 33 tests passed, zero intervention, approximately 12 minutes total.&lt;/p&gt;




&lt;h2&gt;
  
  
  What the Graduation Run Measured
&lt;/h2&gt;

&lt;p&gt;Here's how project 17 scored against the criteria:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Criterion&lt;/th&gt;
&lt;th&gt;Threshold&lt;/th&gt;
&lt;th&gt;Project 17 Result&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;First-run pass rate&lt;/td&gt;
&lt;td&gt;≥ 85%&lt;/td&gt;
&lt;td&gt;93% (13/14 first-shot, 1 retry)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;New FC yield&lt;/td&gt;
&lt;td&gt;≤ 2&lt;/td&gt;
&lt;td&gt;1 (FC-074)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Repeat FC rate&lt;/td&gt;
&lt;td&gt;≤ 5%&lt;/td&gt;
&lt;td&gt;0%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Teacher escalation&lt;/td&gt;
&lt;td&gt;Decreasing&lt;/td&gt;
&lt;td&gt;Zero escalations&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Project 17 met all four thresholds. The preceding project (16.5, catalog-api-v2) reached 83% — close but below the ≥85% line. So we are treating project 17 as the graduation point rather than claiming a two-project stable plateau.&lt;/p&gt;

&lt;p&gt;To be precise about what this means and what it doesn't:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What it means:&lt;/strong&gt; On the specific runs we executed — FastAPI + SQLAlchemy async + pytest projects with CRUD-level complexity and 1:N foreign key relationships, using Qwen3-Coder-Next 45GB Q4_K_M on Apple Silicon M5 Max 128GB with 77 design rules — the system completed the full project autonomously within the scope of new-file-creation tasks.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What it doesn't mean:&lt;/strong&gt; We haven't tested more complex architectural patterns (many-to-many relationships, authentication flows, file uploads, WebSocket endpoints). We haven't tested with different model families or hardware tiers. The 100% figure is for one specific project run; it's a data point, not a guarantee.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;77 rules is a lot of rules.&lt;/strong&gt; Each one was derived from at least one observed problem. But the cumulative load of maintaining 77 interacting rules is substantial. We don't yet know if this scales — whether a 200-rule system would be manageable or would collapse under interaction effects. This matches a concern we are starting to track internally: beyond a certain threshold, adding more constraints may dilute model attention rather than improve output. In our design, we've set a ceiling of 20 CLs per prompt injection bundle to guard against this, but we haven't yet hit a project that tests that limit.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Rule Accumulation Curve
&lt;/h2&gt;

&lt;p&gt;One pattern we've been tracking across the series is how the rate of new rule discovery changes over time:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Projects  1–3:   CL-001 to CL-020   (~7 per project)
Projects  4–6:   CL-021 to CL-035   (~5 per project)
Projects  7–9:   CL-036 to CL-051   (~5 per project)
Projects 10–12:  CL-052 to CL-071   (~6 per project)
Projects 13–17:  CL-072 to CL-077   (~1 per project)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The yield dropped from roughly 7 new rules per project to roughly 1. We're cautious about reading too much into this — it could mean we're approaching the boundary of what our current project complexity can reveal, rather than the boundary of what rules exist. More complex projects might expose entirely new failure categories.&lt;/p&gt;

&lt;p&gt;But within the FastAPI + SQLAlchemy + CRUD scope, the flattening is visible in this dataset. The most notable new failure in this stretch was an interaction effect between existing rules — FC-074 — rather than an entirely novel pattern.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Interaction Effect Problem
&lt;/h2&gt;

&lt;p&gt;FC-074 taught us something we hadn't articulated before: as the rule set grows, the opportunity for interaction effects between rules increases. Each rule is validated independently, but the system runs them all simultaneously.&lt;/p&gt;

&lt;p&gt;This resembles a familiar problem in complex systems: the space of pairwise interactions grows faster than the number of components. We can't test all combinations manually.&lt;/p&gt;

&lt;p&gt;We don't have a systematic solution for this yet. What we have is a detection mechanism: when a failure occurs that doesn't match any existing FC pattern, we now check whether it could be an interaction between two rules that had both worked in isolation in prior runs. FC-074 was caught this way.&lt;/p&gt;

&lt;p&gt;Whether this can be automated — detecting interaction effects without human diagnosis — is an open question. The engine could potentially track which CLs were active when a novel failure occurs and flag the pairwise candidates, but we haven't built that yet.&lt;/p&gt;




&lt;h2&gt;
  
  
  What Comes Next
&lt;/h2&gt;

&lt;p&gt;Graduating from the FastAPI stack opens a question: what do we do with a graduated stack?&lt;/p&gt;

&lt;p&gt;We see two directions, each answering a different question:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Direction A: Complexity escalation.&lt;/strong&gt; Stay on FastAPI but increase project complexity — many-to-many relationships, authentication flows, nested resources, pagination. This tests whether the current 77 rules hold at higher complexity or whether new failure categories emerge.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Direction B: Stack transfer.&lt;/strong&gt; Move to a different framework and measure how many of the 77 rules transfer. Our rules are categorized by stack tags — 29 are marked "universal," 32 are "fastapi"-specific. A new stack would test whether the universal rules actually are universal.&lt;/p&gt;

&lt;p&gt;The question we're most interested in now isn't whether we can achieve another 100% run. It's whether a rule-based agent system can keep growing without becoming harder to reason about than the model it was designed to constrain.&lt;/p&gt;




&lt;h2&gt;
  
  
  Series Links
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://dev.to/josephyeo/i-built-a-local-ai-coding-agent-on-m5-max-128gb-it-failed-164-times-before-passing-35-tests-2fgj"&gt;Part 1: 164 Failures Before 35 Tests&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dev.to/josephyeo/we-didnt-migrate-from-n8n-to-python-because-n8n-failed-k9j"&gt;Part 2: We Didn't Migrate from n8n Because n8n Failed&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dev.to/josephyeo/the-determinism-war-why-we-stopped-chasing-better-models-3c21"&gt;Part 3: The Determinism War&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dev.to/josephyeo/the-information-design-gap-why-our-ai-agent-was-coding-blind-4p8o"&gt;Part 4: The Information Design Gap&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dev.to/josephyeo/dcr-wasnt-enough-why-ai-coding-agents-also-need-information-quality-1da4"&gt;Part 5: DCR Wasn't Enough&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dev.to/josephyeo/the-bug-wasnt-in-the-model-lessons-from-9-local-ai-coding-agent-projects-18aa"&gt;Part 6: The Bug Wasn't in the Model&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dev.to/josephyeo/the-file-modification-boundary-we-found-after-12-forgeflow-projects-3m01"&gt;Part 7: The File Modification Boundary&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;ForgeFlow runs on a MacBook Pro M5 Max 128GB. Planning uses Claude (cloud API). Execution is fully local — Qwen3-Coder-Next 45GB via Ollama, gemma4:26b for QA, Docker sandbox, no API calls during the coding loop. The methodology and failure data are shared in this series.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;If you're building something similar — local AI agents, TDD automation, failure catalog systems — I'd be interested to hear whether you're seeing interaction effects between your own accumulated rules. The comments are open.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>agents</category>
      <category>architecture</category>
      <category>devjournal</category>
      <category>softwareengineering</category>
    </item>
    <item>
      <title>The File Modification Boundary We Found After 12 ForgeFlow Projects</title>
      <dc:creator>Joseph Yeo</dc:creator>
      <pubDate>Fri, 22 May 2026 15:00:08 +0000</pubDate>
      <link>https://dev.to/josephyeo/the-file-modification-boundary-we-found-after-12-forgeflow-projects-3m01</link>
      <guid>https://dev.to/josephyeo/the-file-modification-boundary-we-found-after-12-forgeflow-projects-3m01</guid>
      <description>&lt;p&gt;&lt;em&gt;This is Part 7 of the ForgeFlow series. &lt;a href="https://dev.to/josephyeo/the-bug-wasnt-in-the-model-lessons-from-9-local-ai-coding-agent-projects-18aa"&gt;Part 6: The Bug Wasn't in the Model&lt;/a&gt; ended at 9 projects, 51 failure patterns, and 70 design rules. Up until that point, failure rates in our setup were declining and the working framework felt like it was converging. Project 12 exposed a structural gap we hadn't yet documented.&lt;/em&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Quick terms for new readers:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;FC&lt;/strong&gt; = Failure Catalog entry (a documented failure pattern)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;CL&lt;/strong&gt; = Crystallized Lesson (a testable design rule derived from repeated failures)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Identical GREEN&lt;/strong&gt; = the model returns an unchanged file during the implementation phase&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;DEADLOCK&lt;/strong&gt; = the system gives up after repeated identical failures&lt;/li&gt;
&lt;/ul&gt;
&lt;/blockquote&gt;




&lt;p&gt;Part 6 ended on a high note. Nine projects. A 100% pass rate on the last one. Forty-three crystallized lessons. A working framework in our setup: DCR × Information Quality × Task Complexity. The system felt like it was converging.&lt;/p&gt;

&lt;p&gt;Then we tried self-referential foreign keys, and a failure mode we'd only seen sporadically became the dominant pattern.&lt;/p&gt;

&lt;p&gt;This post is about project 12 — a department hierarchy API with JWT authentication and self-referential parent-child relationships. It documents the failure pattern that connected several scattered observations into a single engineering constraint. And it discusses why, in our case, the most practical response was to restructure the work rather than retry harder.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Setup: Department API
&lt;/h2&gt;

&lt;p&gt;Project 12 was designed to test two development vectors simultaneously: JWT authentication (new for ForgeFlow) and self-referential foreign keys (a department can be a child of another department). The tech stack was familiar — FastAPI, SQLAlchemy async, pytest — but the data model was more complex than our previous test projects.&lt;/p&gt;

&lt;p&gt;The target execution plan: 13 tasks total. Of these, 4 were new-file creation tasks (schemas, tests), 5 were existing-file modification tasks (models, routes), and 4 were either setup steps or handled outside the autonomous loop.&lt;/p&gt;

&lt;p&gt;We ran it five times, redesigning between each iteration. The pattern became hard to ignore.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Scorecard
&lt;/h2&gt;

&lt;p&gt;The table below shows the task categories from Project 12. The same outcome repeated across five redesign-and-rerun attempts.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Task Type&lt;/th&gt;
&lt;th&gt;Count&lt;/th&gt;
&lt;th&gt;Pass Rate&lt;/th&gt;
&lt;th&gt;Avg Cycles&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;New file creation (schemas, tests)&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;td&gt;1.0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Existing file modification (models, routes)&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;0%&lt;/td&gt;
&lt;td&gt;DEADLOCK&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;In our setup, tasks requiring the generation of an entirely new file succeeded on the first attempt. Tasks that required modifying an existing codebase file resulted in a processing deadlock. This held across five separate runs, two different backends (direct Ollama API and Aider), and multiple retry strategies.&lt;/p&gt;

&lt;p&gt;To scope these findings: our dataset is constrained to a single model family (Qwen3-Coder-Next, 45GB Q4_K_M) running on a single hardware tier (Apple Silicon M5 Max 128GB). We don't claim these trends apply universally. But the pattern was consistent enough across five runs that we changed how we structure tasks going forward.&lt;/p&gt;




&lt;h2&gt;
  
  
  What "Identical GREEN" Looks Like
&lt;/h2&gt;

&lt;p&gt;ForgeFlow's TDD loop works in two phases: RED (write a failing test) and GREEN (write code to pass it). The GREEN phase is where modifications happen.&lt;/p&gt;

&lt;p&gt;When a task required modifying an existing file, the following loop repeated:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The model receives the existing file content + test requirements&lt;/li&gt;
&lt;li&gt;The model outputs code that matches the existing file exactly (detected via SHA-256 hash comparison)&lt;/li&gt;
&lt;li&gt;The engine retries with an explicit prompt: "Your output was identical to the current file"&lt;/li&gt;
&lt;li&gt;The model outputs the same file again&lt;/li&gt;
&lt;li&gt;DEADLOCK after 3 identical cycles&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;We call this an &lt;em&gt;identical GREEN&lt;/em&gt; deadlock. The engine already had detection for it (FC-037, added months ago). But we'd only seen it sporadically before. In project 12, it became the primary failure mode.&lt;/p&gt;




&lt;h2&gt;
  
  
  Working Hypotheses
&lt;/h2&gt;

&lt;p&gt;We're cautious about attributing "understanding" to the model — we're observing output patterns, not internal reasoning. Here's what we think might be happening:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The whole-file generation pattern (Ollama backend):&lt;/strong&gt; When generating code via raw completion, the model streams the entire file from the first token. If the existing file is 95% correct and only needs a few lines added, the token history in the context window acts as a statistical attractor — the generation pattern defaults to reproducing the verified, working code rather than deviating to introduce new logic. The smaller the required change relative to the existing file, the stronger this pull appears to be.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The diff generation constraint (Aider backend):&lt;/strong&gt; Diffs require precise line-matching tokens. When the target file is complex — multiple async routes, mixed dependencies, dense imports — generating accurate unified diff chunks appears to become erratic for our local quantized model. In our tests with this specific model and configuration, this manifested as timeouts (capped at 200 seconds per task) or a fallback to emitting an unchanged version of the source file.&lt;/p&gt;

&lt;p&gt;Both pathways showed similar limitations on file modification tasks in our configuration. Whether this is specific to quantized local models or a broader pattern, we can't say.&lt;/p&gt;




&lt;h2&gt;
  
  
  Connecting Scattered Observations
&lt;/h2&gt;

&lt;p&gt;Before project 12, our tracker had three separate failure patterns that each captured a piece of this:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;FC-034 / CL-043&lt;/strong&gt;: "One task, one endpoint" — adding endpoints to an existing route file often resulted in syntax errors or duplicates&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;FC-047 / CL-066&lt;/strong&gt;: "Over-complete stubs" — when a stub had significant boilerplate, the model treated it as finished&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;FC-039 / CL-058&lt;/strong&gt;: "POST endpoints need Aider" — some tasks specifically failed on the Ollama backend&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Project 12 gave us the data to connect these into a single classification, &lt;strong&gt;FC-052&lt;/strong&gt;:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;In our local execution setup, existing file modification tasks demonstrate a high probability of identical GREEN DEADLOCK on both whole-file and diff-based backends. In our observations, identical-output failures appeared more often when the required change was small relative to the existing file.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;From FC-052, we derived &lt;strong&gt;CL-071&lt;/strong&gt;:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Every autonomous task target should be a new file. If a workflow step must modify an existing file, that modification should either be handled programmatically during setup or the architecture should be decoupled so that features reside in isolated modules.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This became our 71st crystallized lesson, and it changed how we now structure ForgeFlow projects.&lt;/p&gt;

&lt;p&gt;One notable data point: across three complete projects (10, 11, and 12), our failure catalog expanded by only a single new entry. The rule accumulation curve is flattening, which may suggest we're mapping the boundary of our current configuration — or just the boundary of our current project complexity.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Design Pattern That Emerged
&lt;/h2&gt;

&lt;p&gt;CL-071 pushed us to rethink how we write PRDs.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Before (task-level modifications):&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;TASK-001: Create User model (stub)        → models/user.py
TASK-002: Add fields to User model        → models/user.py    [DEADLOCK]
TASK-003: Create Department model (stub)  → models/department.py
TASK-004: Add relationship                → models/department.py [DEADLOCK]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;After (decoupled new-file generation):&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;SETUP SCRIPT: Generate complete models with all fields and relationships
TASK-001: Create User schemas    → schemas/user.py      [NEW FILE ✅]
TASK-002: Create Dept schemas    → schemas/department.py [NEW FILE ✅]
TASK-003: Create register route  → routes/auth.py        [NEW FILE ✅]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The pattern: infrastructure is established deterministically during setup, while the model handles clean-sheet file generation.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;An important caveat:&lt;/strong&gt; applying this pattern to project 12 was not a clean autonomous success. We manually implemented the CRUD endpoints (6 routes) to unblock the dependency chain, then tested whether the remaining new-file task would run cleanly under the revised structure. The integration test — creating a fresh &lt;code&gt;test_integration.py&lt;/code&gt; — passed on its first autonomous cycle. The important result was narrower than "we solved it": once existing-file modification was removed from the autonomous task path, the remaining new-file task completed cleanly.&lt;/p&gt;

&lt;p&gt;We should also note an open concern: forcing every task into a "new file only" pattern shifts complexity from generation-time editing to project-level file organization. At 13 tasks, this is manageable. At 50+, it could create significant file fragmentation and import overhead. We haven't tested at that scale yet.&lt;/p&gt;




&lt;h2&gt;
  
  
  Where We Are After 12 Projects
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Value&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Total projects&lt;/td&gt;
&lt;td&gt;12 (11 completed, 1 scrapped)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Failure patterns cataloged (FC)&lt;/td&gt;
&lt;td&gt;52&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Design rules (CL)&lt;/td&gt;
&lt;td&gt;71&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Automated rule checks&lt;/td&gt;
&lt;td&gt;53 functions in validate_prd.py&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sessions&lt;/td&gt;
&lt;td&gt;81&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h2&gt;
  
  
  The Honest Assessment
&lt;/h2&gt;

&lt;p&gt;After 12 projects and 81 sessions:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What's working in our setup:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;New file generation from detailed specs: reliable across the runs we tested&lt;/li&gt;
&lt;li&gt;TDD enforcement (RED must fail, GREEN must pass): useful as a mechanical guardrail&lt;/li&gt;
&lt;li&gt;Failure pattern → design rule pipeline: producing diminishing but real returns&lt;/li&gt;
&lt;li&gt;Setup-based infrastructure + model-based creation: tested over 3 projects&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;What isn't working:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Existing file modification: consistently unreliable with our current model and configuration&lt;/li&gt;
&lt;li&gt;Non-deterministic results on complex tasks: one task passed in 2 out of 3 runs, failed in 1. Same code, same model, different outcome.&lt;/li&gt;
&lt;li&gt;Long dependency chains: a single DEADLOCK blocks everything downstream&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Open questions:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Does CL-071 hold on 20+ task projects with complex dependency graphs?&lt;/li&gt;
&lt;li&gt;Does the "new file only" constraint create unsustainable file fragmentation at scale?&lt;/li&gt;
&lt;li&gt;Will newer local models (Qwen3-Coder v2, Llama 4) shift this boundary?&lt;/li&gt;
&lt;li&gt;Is this specific to quantized local models, or do cloud API models show similar patterns on file modification tasks inside TDD loops?&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  A Request to Readers
&lt;/h2&gt;

&lt;p&gt;If you're running local models — Ollama, llama.cpp, vLLM, or something else — within autonomous execution loops, we'd be interested in learning whether your telemetry shows similar variations between file creation and file modification tasks.&lt;/p&gt;

&lt;p&gt;Specifically: &lt;strong&gt;how do your local configurations handle incremental diff generation inside structured loops versus generating complete, fresh modules from detailed specs?&lt;/strong&gt; If you've logged similar boundaries or found alternative designs to work around modification deadlocks, please share your setup and observations in the comments.&lt;/p&gt;

&lt;p&gt;We're also curious whether anyone has hard metrics on how cloud models (GPT, Claude) perform on targeted file modifications inside closed-loop TDD environments. Our dataset is one model family on one hardware tier — more data points from different setups would help everyone working in this space.&lt;/p&gt;




&lt;h2&gt;
  
  
  What's Next
&lt;/h2&gt;

&lt;p&gt;Project 13 will be the first real test of whether CL-071 is a design principle or just a project-12-specific workaround. Every implementation task will target a new file. Setup will handle all infrastructure. The open question isn't whether it passes — it's whether the "new file only" constraint produces a project structure that's actually maintainable at 20+ tasks.&lt;/p&gt;

&lt;p&gt;We're also adding automatic CL-071 validation to &lt;code&gt;validate_prd.py&lt;/code&gt; — a check that flags any task whose implementation target already exists at execution time. For our workflow, rules that repeatedly affect outcomes should probably be machine-enforced.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Series So Far
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;a href="https://dev.to/josephyeo/i-built-a-local-ai-coding-agent-on-m5-max-128gb-it-failed-164-times-before-passing-35-tests-2fgj"&gt;I Built a Local AI Coding Agent on M5 Max 128GB&lt;/a&gt; — 164 failures, 35 tests, proof of concept&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://dev.to/josephyeo/we-didnt-migrate-from-n8n-to-python-because-n8n-failed-k9j"&gt;We Didn't Migrate from n8n to Python Because n8n Failed&lt;/a&gt; — The orchestrator rewrite&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://dev.to/josephyeo/the-determinism-war-why-we-stopped-chasing-better-models-3c21"&gt;The Determinism War&lt;/a&gt; — Why we stopped chasing better models&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://dev.to/josephyeo/the-information-design-gap-why-our-ai-agent-was-coding-blind-4p8o"&gt;The Information Design Gap&lt;/a&gt; — Why the agent was coding blind&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://dev.to/josephyeo/dcr-wasnt-enough-why-ai-coding-agents-also-need-information-quality-1da4"&gt;DCR Wasn't Enough&lt;/a&gt; — Adding information quality to the framework&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://dev.to/josephyeo/the-bug-wasnt-in-the-model-lessons-from-9-local-ai-coding-agent-projects-18aa"&gt;The Bug Wasn't in the Model&lt;/a&gt; — Lessons from 9 projects&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The File Modification Boundary&lt;/strong&gt; — You are here. 12 projects, a boundary mapped.&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  About
&lt;/h2&gt;

&lt;p&gt;I'm Joseph YEO, a solo builder from Seoul, Korea. ForgeFlow runs entirely on a MacBook Pro M5 Max 128GB — no cloud APIs during execution. The planning agent (Claude) designs the specs. The local model (Qwen3-Coder-Next, 45GB Q4_K_M) executes the TDD loop autonomously.&lt;/p&gt;

&lt;p&gt;Follow along:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;𝕏: &lt;a href="https://x.com/josephyeo_dev" rel="noopener noreferrer"&gt;@josephyeo_dev&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;GitHub: &lt;a href="https://github.com/joseph-yeo" rel="noopener noreferrer"&gt;joseph-yeo&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Site: &lt;a href="https://projectjoseph.dev" rel="noopener noreferrer"&gt;projectjoseph.dev&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Built over 81 sessions, May 2026. All models run locally via Ollama 0.23.0 on macOS. No cloud APIs were used during autonomous execution.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;This post was drafted with Claude and edited by me.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>llm</category>
      <category>softwareengineering</category>
    </item>
  </channel>
</rss>
