<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Dexterlung</title>
    <description>The latest articles on DEV Community by Dexterlung (@dexterlung).</description>
    <link>https://dev.to/dexterlung</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4029426%2F0d3f03e5-03d2-46ff-975f-c565c23e82ce.jpg</url>
      <title>DEV Community: Dexterlung</title>
      <link>https://dev.to/dexterlung</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/dexterlung"/>
    <language>en</language>
    <item>
      <title>I Kept Feeling Like I Was Wasting My Time — While Designing a Cross-Model Controlled Experiment</title>
      <dc:creator>Dexterlung</dc:creator>
      <pubDate>Mon, 24 Aug 2026 13:05:21 +0000</pubDate>
      <link>https://dev.to/dexterlung/i-kept-feeling-like-i-was-wasting-my-time-while-designing-a-cross-model-controlled-experiment-1l8f</link>
      <guid>https://dev.to/dexterlung/i-kept-feeling-like-i-was-wasting-my-time-while-designing-a-cross-model-controlled-experiment-1l8f</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;I pulled 1,427 of my own prompts from six weeks with AI.&lt;br&gt;
I meant to see "how did my way of asking change."&lt;br&gt;
The sharpest thing wasn't the capability curve — it was that the curve and how I saw myself were a full tier apart.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Six months ago I wrote a piece called "Three Months, 1,604 Prompts: What Did AI Trade With Me?" That time I scanned "what I handle most." This time I wanted to look at something uglier: &lt;strong&gt;how do I ask — and did it change over six months?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I pulled every prompt I typed myself between June 15 and July 23 — 1,427 of them (stripping tool outputs, system messages, the fat-fingered interrupts) — cut them into six time windows, and measured them window by window.&lt;/p&gt;

&lt;p&gt;The numbers were clear. What actually stopped me was the person standing next to the numbers.&lt;/p&gt;




&lt;h2&gt;
  
  
  First, the numbers: my way of asking really did shift gears
&lt;/h2&gt;

&lt;p&gt;I tagged each prompt with a few categories of vocabulary, sorted by time:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Period&lt;/th&gt;
&lt;th&gt;Adversarial / verify / root-cause&lt;/th&gt;
&lt;th&gt;Meta / governance / method&lt;/th&gt;
&lt;th&gt;Delegation / automation / batch&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Mid-to-late June&lt;/td&gt;
&lt;td&gt;4%&lt;/td&gt;
&lt;td&gt;13%&lt;/td&gt;
&lt;td&gt;14%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Late June&lt;/td&gt;
&lt;td&gt;9%&lt;/td&gt;
&lt;td&gt;28%&lt;/td&gt;
&lt;td&gt;34%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Early July&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;21%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;36%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;34%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;In under six weeks, "make the AI push back, verify this, find the root cause" grew fivefold; "talk about method, governance, systematizing" nearly tripled. My average prompt length also jumped from just over 200 characters to around 1,000 — I'd started writing the kind of long, context-first "strategy prompt."&lt;/p&gt;

&lt;p&gt;If you only look at that table, it's an inspiring story: someone with no engineering background, in six months, going from "fix this bug for me" to "work backwards from my git scars to the pain most worth preventing."&lt;/p&gt;

&lt;p&gt;But I'm not here to write an inspirational post.&lt;/p&gt;




&lt;h2&gt;
  
  
  The same week, this is how I talked to the AI at 2am
&lt;/h2&gt;

&lt;p&gt;Early morning, July 18, I typed this:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Having the ability doesn't mean I've actually productized it… I don't have an SOP or a cold-start flow that can cold-start in one day and ship an MVP in three… nothing is pushing me forward. Facing it head-on relies entirely on my anxiety, so I keep opening new sessions and asking, over and over."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A little earlier, July 17:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"It's hard not to feel like no one would pay for my service, that other people's is better."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Earlier still, June 25:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Now that AI is this powerful and everyone can do things easily on their own — what am I even doing? Am I just wasting my time?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;These aren't cherry-picked extremes. Sentences like this show up again and again across those 1,427 prompts — &lt;strong&gt;late at night, fishing for reassurance, feeling like I go deep on single points but can't connect them into a loop, feeling like others do effortlessly what I have to grind for.&lt;/strong&gt; I'd often, in the same prompt, pour out a stack of self-doubt and then ask a genuinely hard technical question.&lt;/p&gt;

&lt;p&gt;I always thought I knew what I was doing. Laid open, it turned out my assessment of myself was frozen six months in the past.&lt;/p&gt;




&lt;h2&gt;
  
  
  But the same week, here's what I was actually doing
&lt;/h2&gt;

&lt;p&gt;This is where the gap is sharpest. Right around the days of "am I just wasting my time," my prompt log has these:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;I designed a cross-model controlled experiment with my own hands.&lt;/strong&gt; On July 10, to verify whether a methodology "skill" actually made the model smarter, I asked:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Can you open a sandbox or subagent right here and have haiku run it? And sonnet? If they'd be contaminated by this project's claude.md, tell me and I'll paste it manually."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;I didn't even notice — someone with no statistics training, no engineering background, instinctively knew to &lt;strong&gt;isolate the variable&lt;/strong&gt; (worried the project config would contaminate the experiment), to &lt;strong&gt;run a control group&lt;/strong&gt; (with skill vs. without), to &lt;strong&gt;cross-check across different models&lt;/strong&gt;. That's experimental design. And I felt like I was wasting my time.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;I started giving the AI red-team orders.&lt;/strong&gt; On July 21, I told it to attack a defense I'd just built myself:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Using a non-homologous model, ask 'what does this lens itself miss? Under what conditions would it give false reassurance?' — I want it to attack, not endorse."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;I settled on a "trust the scar, not my account" extraction method.&lt;/strong&gt; In early July, I wanted to capture a frontier model's judgment into a reusable skill. I didn't ask it "how do you think" — I knew that would get a beautiful but empty answer. I told it to work from my git history:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Git is the crystallization of scars: a repeated fix = a pain that was never prevented, that keeps recurring."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;I even started using the AI to recalibrate my own perception.&lt;/strong&gt; By July 23, I wasn't asking "how" anymore, I was asking:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"What's the real value I provide? _____? Please recalibrate me."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Giving red-team orders, designing controlled experiments, telling the difference between "self-report" and "behavioral evidence," turning around to have the AI correct my own positioning — &lt;strong&gt;these are not a beginner's questions.&lt;/strong&gt; This is someone who knows what he wants and knows how to force the AI to give up the real answer.&lt;/p&gt;

&lt;p&gt;That person and the one fishing for reassurance at 2am were the same me, the same week.&lt;/p&gt;




&lt;h2&gt;
  
  
  The gap itself is the point
&lt;/h2&gt;

&lt;p&gt;I set out to write a nice growth curve. By the time I got here, I'd changed my mind.&lt;/p&gt;

&lt;p&gt;What's actually worth writing down is &lt;strong&gt;the seam between self-assessment and actual judgment.&lt;/strong&gt; Because I'm almost certain that if you're also a solo operator building things with AI, you have this seam too. At night you feel like you're faking it, chasing someone else's taillights; by day you're doing things you don't even realize are hard.&lt;/p&gt;

&lt;p&gt;Three things I learned:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;One: your gut feeling about yourself is the least reliable instrument.&lt;/strong&gt; My read on "what I'm doing" lagged my actual ability by six months. If I'd gone and looked at the record earlier instead of going by feel, I'd have saved myself a lot of anxious nights. So now I periodically pull my own conversations and look — not out of vanity, but to calibrate. Feelings lie; the record doesn't.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Two: "I go deep on single points but can't connect them," I said this to myself so many times it became an excuse instead of a diagnosis.&lt;/strong&gt; The record shows that by July I was already running my first real client case, doing end-to-end dry runs, wiring scattered things into a flywheel. It's not that I can't connect — it's that I kept using "I can't connect" to block myself from seeing how much I already had.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Three: professional capability isn't shown by bragging, it's shown by laying it open.&lt;/strong&gt; In this whole piece I never said "I'm good." I just put my own prompts side by side — the ones fishing for reassurance, next to the ones giving red-team orders. The gap speaks for itself. That's more convincing than any "I'm a senior AI collaborator," because it's real, and it doesn't even hide my own awkwardness.&lt;/p&gt;




&lt;h2&gt;
  
  
  One small thing for you
&lt;/h2&gt;

&lt;p&gt;If you use Claude Code or a similar tool, your conversation history is sitting in jsonl files on your machine (Claude Code keeps them under &lt;code&gt;~/.claude/projects/&lt;/code&gt;). Spend half an hour writing a script to pull your own prompts from the past few months, and look at two things month over month:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Did your prompts get longer or shorter?&lt;/strong&gt; Longer usually means you started giving context and direction, not just orders.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Did words like "verify this / are you sure / is there a better way" rise as a share?&lt;/strong&gt; That's the signal of going from "commanding a tool" to "working with a collaborator you can challenge."&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Then — and this is the most important step — &lt;strong&gt;put those numbers next to the assessment of yourself in your head.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If they match, congratulations, you know yourself well.&lt;br&gt;
If they don't, if they're a full tier apart like mine, then you owe yourself an apology. You've come further than you think.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;The data for this comes from a script I ran over my own 171 sessions and 1,427 prompts from June–July. The method is the same as that piece six months ago, "Three Months, 1,604 Prompts" — except this time I didn't stop at the numbers; I also looked at the self-doubting me standing right next to them.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;本文原載於我的部落格：&lt;a href="https://coffeeshooters.com/content/impostor-gap-self-doubt-vs-demonstrated-judgment-en?utm_source=devto&amp;amp;utm_medium=social&amp;amp;utm_campaign=blog-impostor-gap-self-doubt-vs-demonstrated-judgment-en" rel="noopener noreferrer"&gt;I Kept Feeling Like I Was Wasting My Time — While Designing a Cross-Model Controlled Experiment&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>solodev</category>
    </item>
    <item>
      <title>How I Ask AI Changed: From "Fix This" to "Recalibrate Me"</title>
      <dc:creator>Dexterlung</dc:creator>
      <pubDate>Sat, 22 Aug 2026 13:05:11 +0000</pubDate>
      <link>https://dev.to/dexterlung/how-i-ask-ai-changed-from-fix-this-to-recalibrate-me-52f7</link>
      <guid>https://dev.to/dexterlung/how-i-ask-ai-changed-from-fix-this-to-recalibrate-me-52f7</guid>
      <description>&lt;p&gt;Two questions. Both are things I typed to an AI myself. Six weeks apart.&lt;/p&gt;

&lt;p&gt;June 24th, I asked:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"That checkout button does nothing. No console error either. Help me find the cause."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;July 23rd, I asked:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"I need you to correct my thinking. I'm not offering the kind of automation-plumbing service that everyone already does — what's the real value I provide? Please recalibrate me."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Same person. Same tool. Same AI. But those two questions are asking from &lt;strong&gt;two completely different floors of a building.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The first one: I know where the problem is (the button's broken), I want you to find the cause.&lt;br&gt;
The second one: I don't know who I am, I want you to correct how I see myself.&lt;/p&gt;

&lt;p&gt;I pulled all 1,427 prompts I typed to an AI between mid-June and late July. I set out to see "have I changed." What I found wasn't about how much I know technically — it's that the &lt;strong&gt;abstraction level of my questions&lt;/strong&gt; climbed, one rung at a time. And this ladder is something nobody ever taught me. I only noticed it existed by looking back at the record.&lt;/p&gt;

&lt;p&gt;So this is me pulling that invisible ladder apart. Because I'm increasingly convinced: &lt;strong&gt;real skill with AI isn't whether you can write a clever prompt — it's which rung your question is standing on.&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Rung 1: "Fix this" — the instruction level
&lt;/h2&gt;

&lt;p&gt;The bottom rung. You already know the shape of the answer; the AI is just your hands.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;"Help me find the cause."&lt;/li&gt;
&lt;li&gt;"Darken the whole thing — 100% opacity isn't dark enough."&lt;/li&gt;
&lt;li&gt;"Paste this to wake-and-send.py: --x 1966 --y 1303."&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This rung is useful and necessary — I still ask this way every single day. But if you &lt;strong&gt;only&lt;/strong&gt; stay here, the AI is just a version of you that types faster. As specific as you are, that's as specific as it gets; whatever you can't think of, it won't think of for you either.&lt;/p&gt;

&lt;p&gt;Mid-June me lived mostly on this rung. My prompts back then averaged barely 200 characters — short, direct, just giving orders.&lt;/p&gt;




&lt;h2&gt;
  
  
  Rung 2: "Got a good approach?" — the intent level
&lt;/h2&gt;

&lt;p&gt;One rung up, you state &lt;strong&gt;what you want&lt;/strong&gt;, and hand over the "how."&lt;/p&gt;

&lt;p&gt;June 26th I typed this, and it still feels honest to me:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"The pieces are actually there, they're just not wired together. I'm pretty weak at this wiring-together part — so, got a good approach?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;What's the difference? On rung 1 I give the shape of the answer; here I &lt;strong&gt;only give a direction and hand over the solution space.&lt;/strong&gt; That takes a bit of courage — you have to admit you don't know how, before you can ask "got a good approach."&lt;/p&gt;

&lt;p&gt;From this rung on, the AI starts handing you things you couldn't have thought of. Because you didn't pin it down with a specific instruction.&lt;/p&gt;




&lt;h2&gt;
  
  
  Rung 3: "Design me a mechanism" — the meta level
&lt;/h2&gt;

&lt;p&gt;On the third rung, you stop solving "a problem." You solve a &lt;strong&gt;whole class&lt;/strong&gt; of problems.&lt;/p&gt;

&lt;p&gt;June 23rd, I wasn't trying to fix a bug — I was trying to stop a kind of bug from ever happening again:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"I want to design a skill that proactively saves branching, unfinished tasks. When a conversation drifts into a fork, can it trigger a prompt asking whether to save it?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;July 3rd was even clearer — I wanted to &lt;strong&gt;extract a model's judgment into a method&lt;/strong&gt;:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Looking at it from a higher level: how does it scope a problem, sketch the outline, converge, and sequence the flow… make this endless stream of problems systematic."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;On this rung, I go from "the person solving problems" to "the person designing the problem-solving machine." And the most interesting part — &lt;strong&gt;I often use my own flaws as the starting point of the design.&lt;/strong&gt; I know I lose track of branching tasks, so I design a mechanism to catch them. I turn "what I'm bad at" into a spec.&lt;/p&gt;

&lt;p&gt;In my logs from late June to early July, the share of prompts that "talk about method, governance, systematizing" jumped from 13% to 36%. Not because I suddenly got smarter — because the position I was asking from moved up a whole floor.&lt;/p&gt;




&lt;h2&gt;
  
  
  Rung 4: "Recalibrate me" — the self-correction level
&lt;/h2&gt;

&lt;p&gt;The top rung. The hardest. I only climbed onto it recently.&lt;/p&gt;

&lt;p&gt;Here, you're not asking about the world — you're asking about &lt;strong&gt;yourself.&lt;/strong&gt; You want the AI to be a mirror, correcting your judgment about yourself and your direction.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;July 15th: "Thanks for hitting the brakes. Let me go push outward first — stop building inward." (I wanted it to stop my own reflex.)&lt;/li&gt;
&lt;li&gt;July 23rd: "What's the real value I provide? Please recalibrate me." (I wanted it to correct my self-positioning.)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Why is this the hardest rung? Because on the first three you're still holding the wheel — you know what to fix, what approach you want, what class of problem to prevent. On rung 4, what you hand over is &lt;strong&gt;"my judgment about myself might be wrong."&lt;/strong&gt; You have to admit you might not see yourself clearly, before you can ask "please recalibrate me."&lt;/p&gt;

&lt;p&gt;This takes no technical skill — it takes a very mature self-awareness: &lt;strong&gt;knowing your blind spots need an outsider to light them up.&lt;/strong&gt; And an outsider is exactly what a solo developer lacks most.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why does the ladder climb itself?
&lt;/h2&gt;

&lt;p&gt;I used to assume it was the AI getting stronger, so I could ask harder questions. Looking back at the record — no.&lt;/p&gt;

&lt;p&gt;It's that &lt;strong&gt;each rung shows you the next rung exists.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I fixed the same kind of broken button too many times on rung 1 — that's what made me ask, on rung 3, "can we design a mechanism to prevent this." I built too many systems on rung 3 without pushing a single one out — that's what made me hit rung 4: "am I pointed the wrong way, recalibrate me." &lt;strong&gt;The pain of each rung is the door to the next.&lt;/strong&gt; You don't teleport to rung 4 — you have to trip over the same stone on rung 1 enough times first.&lt;/p&gt;

&lt;p&gt;So if you spend most of your time right now on rung 1 asking "fix this," that's completely normal. I do too. The point isn't to force yourself to skip rungs — it's to &lt;strong&gt;not stay there pretending that's all there is.&lt;/strong&gt; Next time you catch yourself fixing the same kind of thing a third time, try asking one level up: "is there a way to stop this class of problem from coming back?" — and you've just stepped onto rung 2.&lt;/p&gt;




&lt;h2&gt;
  
  
  So what skill is this, exactly?
&lt;/h2&gt;

&lt;p&gt;I want to be clear about one thing, because it's counterintuitive.&lt;/p&gt;

&lt;p&gt;A lot of people think "good at using AI" = "good at writing impressive prompts." It's not. My 1,427 prompts are full of typos, casual phrasing, "what do you think?", "I'm just confused." My prompts are not "engineered" at all.&lt;/p&gt;

&lt;p&gt;What's actually improving is &lt;strong&gt;my judgment about which rung to put a question on.&lt;/strong&gt; The same bug, asked on rung 1 as "fix this," versus on rung 3 as "design me a mechanism that prevents this class of bug," gives you wildly different things. And knowing &lt;em&gt;when&lt;/em&gt; to move a question &lt;strong&gt;up a rung&lt;/strong&gt; — that judgment is the real skill.&lt;/p&gt;

&lt;p&gt;It's also why I'm not that worried about "AI is so strong now, everyone can do it, so what's my value." Everyone has the tool. But the person standing on rung 4, who knows to ask "recalibrate me" instead of just "fix this" — &lt;strong&gt;that's judgment you climbed to one rung at a time, over years of tripping over stones. No prompt template hands you that.&lt;/strong&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;This comes from a script I ran over my own 6–7 month history — 171 sessions, 1,427 prompts. From the same data I also wrote "I Thought I Was Wasting My Time — While Designing a Cross-Model Controlled Experiment," about the gap between capability and self-assessment. This piece is about how the capability itself grows, one rung at a time.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;本文原載於我的部落格：&lt;a href="https://coffeeshooters.com/content/how-i-ask-ai-from-fix-to-recalibrate-en?utm_source=devto&amp;amp;utm_medium=social&amp;amp;utm_campaign=blog-how-i-ask-ai-from-fix-to-recalibrate-en" rel="noopener noreferrer"&gt;How I Ask AI Changed: From "Fix This" to "Recalibrate Me"&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>solodev</category>
    </item>
    <item>
      <title>The antivirus said 'no malware detected'. It was sitting right there on the disk.</title>
      <dc:creator>Dexterlung</dc:creator>
      <pubDate>Thu, 20 Aug 2026 13:05:13 +0000</pubDate>
      <link>https://dev.to/dexterlung/the-antivirus-said-no-malware-detected-it-was-sitting-right-there-on-the-disk-4fl8</link>
      <guid>https://dev.to/dexterlung/the-antivirus-said-no-malware-detected-it-was-sitting-right-there-on-the-disk-4fl8</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;A real incident, and a lesson about green lights.&lt;br&gt;
Every command output, version number and CVE ID below is from the actual investigation. Nothing was invented for the narrative.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  It started with an unrelated question
&lt;/h2&gt;

&lt;p&gt;I was tidying up scheduled tasks on my Synology NAS and opened Task Scheduler. Two entries I didn't remember creating:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;PowerOff task 0 → 2026-07-26 09:00
PowerOn  task 0 → 2026-07-26 20:00
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;So I asked: &lt;strong&gt;"Why is this here? I never set this up."&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;While digging through the system crontab, I found this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="k"&gt;*&lt;/span&gt;/20 &lt;span class="k"&gt;*&lt;/span&gt; &lt;span class="k"&gt;*&lt;/span&gt; &lt;span class="k"&gt;*&lt;/span&gt; &lt;span class="k"&gt;*&lt;/span&gt; /bin/sh /etc/.conf &lt;span class="c"&gt;#Sn5Yj8A2l0T&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A dot-prefixed (hidden) file in &lt;code&gt;/etc&lt;/code&gt;, executed &lt;strong&gt;as root every 20 minutes&lt;/strong&gt;, tagged with a random string.&lt;/p&gt;

&lt;p&gt;The power schedule turned out to be unrelated. But if I hadn't asked that question, I would never have opened that file.&lt;/p&gt;

&lt;h2&gt;
  
  
  I didn't scan it. I just read it.
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;#!/bin/bash&lt;/span&gt;

&lt;span class="nv"&gt;MATCH_STRING&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"Sn5Yj8A2l0T"&lt;/span&gt;
&lt;span class="nv"&gt;DOWNLOAD_URL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"http://zuoye.free.fr/files/synology-10441.png"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Reading further, it does four things:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;If it's been deleted&lt;/strong&gt; → &lt;code&gt;wget&lt;/code&gt; itself back&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;If &lt;code&gt;/etc/crontab&lt;/code&gt; lacks the marker&lt;/strong&gt; → &lt;strong&gt;overwrite&lt;/strong&gt; the file (&lt;code&gt;&amp;gt;&lt;/code&gt;, not &lt;code&gt;&amp;gt;&amp;gt;&lt;/code&gt;) to reinstall the cron line&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;If &lt;code&gt;/etc/rc.subr&lt;/code&gt; lacks the marker&lt;/strong&gt; → append &lt;code&gt;bash /etc/.conf &amp;amp;&lt;/code&gt; (boot persistence)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;If the payload isn't running&lt;/strong&gt; → download &lt;code&gt;000119.png&lt;/code&gt;, save it as &lt;code&gt;node&lt;/code&gt;, execute it&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That &lt;code&gt;.png&lt;/code&gt; is not an image. First 16 bytes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;0000000 177   E   L   F 002 001 001  \0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;\177ELF&lt;/code&gt;. It's a Linux binary. The extension is camouflage.&lt;/p&gt;

&lt;p&gt;It lived at &lt;code&gt;/etc/node&lt;/code&gt;, 568 KB, dated &lt;strong&gt;2026-01-14&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;(The real Node.js is at &lt;code&gt;/usr/local/bin/node&lt;/code&gt;. Something called &lt;code&gt;node&lt;/code&gt; sitting in &lt;code&gt;/etc/&lt;/code&gt; is not a system component.)&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The file was dated January. I found it in July. Six months.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Then I asked the vendor's own tool
&lt;/h2&gt;

&lt;p&gt;Synology ships Security Advisor, which scans for malware. I ran a full scan.&lt;/p&gt;

&lt;p&gt;Result:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;✅ No malware detected on your system
✅ No malicious cryptocurrency mining software detected
✅ No malicious system configuration files detected
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Three green checks.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;And at that exact moment, &lt;code&gt;/etc/.conf&lt;/code&gt; and &lt;code&gt;/etc/node&lt;/code&gt; were on the disk. I could &lt;code&gt;cat&lt;/code&gt; them again for anyone who asked.&lt;/p&gt;

&lt;h2&gt;
  
  
  What that green light actually proved
&lt;/h2&gt;

&lt;p&gt;This is the part worth your time.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The tool didn't lie. It just didn't see.&lt;/strong&gt; Three concrete reasons:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. The payload was UPX-packed.&lt;/strong&gt;&lt;br&gt;
The only readable string I could extract from the binary was &lt;code&gt;http://upx.sf.net&lt;/code&gt;. UPX compresses executables; a side effect is that every string inside is compressed too. Signature-based scanners match fingerprints. Compress the fingerprint and there's nothing to match.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. The dropper is an ordinary shell script.&lt;/strong&gt;&lt;br&gt;
It isn't a "virus format". Every line, read alone, is legitimate bash: &lt;code&gt;wget&lt;/code&gt;, &lt;code&gt;chmod&lt;/code&gt;, &lt;code&gt;echo&lt;/code&gt;. What's malicious is &lt;strong&gt;what they do together&lt;/strong&gt; — and that requires comprehension, not comparison.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. The cron entry is valid syntax.&lt;/strong&gt;&lt;br&gt;
The "malicious configuration file" check looks for &lt;strong&gt;known-bad templates&lt;/strong&gt;, not for "what is this line doing". &lt;code&gt;*/20 * * * * /bin/sh /etc/.conf&lt;/code&gt; is syntactically indistinguishable from any legitimate schedule.&lt;/p&gt;

&lt;p&gt;So:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;A green light proves "no bad news was seen". It does not prove "there is no bad news".&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;In daily life those two are nearly equivalent, so we treat them as one thing. They aren't. And the gap shows up exactly when it matters most.&lt;/p&gt;
&lt;h2&gt;
  
  
  The attacker left a business card
&lt;/h2&gt;

&lt;p&gt;Before cleaning up, I recorded the SHA256 of both files and went back to that download URL:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;http://zuoye.free.fr/files/synology-10441.png
                          ^^^^^^^^^^^^^^^^
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;synology-10441&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;CVE-2024-10441&lt;/strong&gt; — an unauthenticated remote code execution flaw in Synology DSM's system plugin daemon. &lt;strong&gt;CVSS 9.8&lt;/strong&gt;. It came out of Pwn2Own 2024. No credentials, no user interaction: one crafted request, arbitrary code execution.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The attacker named the payload after the vulnerability they used to get in.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Then I checked versions:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Fixed in&lt;/td&gt;
&lt;td&gt;DSM &lt;code&gt;7.2.1-69057-6&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;I was running&lt;/td&gt;
&lt;td&gt;DSM &lt;code&gt;7.2.1-69057&lt;/code&gt; &lt;strong&gt;Update 3&lt;/strong&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Below the fix.&lt;/strong&gt; The advisory was published in late 2024. My NAS sat on an older build the whole time.&lt;/p&gt;

&lt;h2&gt;
  
  
  My first hypothesis was wrong
&lt;/h2&gt;

&lt;p&gt;Before I found that filename, my working theory was: &lt;em&gt;"Probably a brute-forced password — auto-block was disabled, after all."&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;That was wrong.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;CVE-2024-10441 requires &lt;strong&gt;no authentication at all&lt;/strong&gt;. The attacker never attempted a login, so auto-block was irrelevant — it would have made no difference either way.&lt;/p&gt;

&lt;p&gt;The actual root cause needed exactly two conditions:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Unpatched DSM (below the fixed build)
      ×
Management interface reachable from the internet
      ↓
   one request → root
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I tested four ports from my phone on mobile data (WiFi off). All refused — so there were &lt;strong&gt;no port-forwarding rules&lt;/strong&gt; on the router. Where did the exposure come from?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;QuickConnect.&lt;/strong&gt; The vendor's convenience feature that lets you reach your NAS from outside without touching your router. It works by relaying through the vendor's servers — &lt;strong&gt;no port forwarding required&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Convenience and exposure are the same thing viewed from two sides.&lt;/p&gt;

&lt;h2&gt;
  
  
  Removal: the order matters more than the commands
&lt;/h2&gt;

&lt;p&gt;The commands are short, but &lt;strong&gt;the wrong order wastes the effort&lt;/strong&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# 1. Break both persistence paths FIRST&lt;/span&gt;
&lt;span class="nb"&gt;sed&lt;/span&gt; &lt;span class="nt"&gt;-i&lt;/span&gt; &lt;span class="s1"&gt;'/Sn5Yj8A2l0T/d'&lt;/span&gt; /etc/crontab
&lt;span class="nb"&gt;sed&lt;/span&gt; &lt;span class="nt"&gt;-i&lt;/span&gt; &lt;span class="s1"&gt;'/Sn5Yj8A2l0T/d'&lt;/span&gt; /etc/rc.subr

&lt;span class="c"&gt;# 2. THEN delete the files&lt;/span&gt;
&lt;span class="nb"&gt;rm&lt;/span&gt; &lt;span class="nt"&gt;-f&lt;/span&gt; /etc/.conf /etc/node
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Reverse it — delete first, edit cron second — and within 20 minutes the schedule fires, &lt;code&gt;wget&lt;/code&gt; pulls it back, and you conclude "I can't remove it."&lt;/p&gt;

&lt;p&gt;One more trap worth recording: I first tried pasting the whole block with &lt;code&gt;sudo&lt;/code&gt; prefixes. &lt;code&gt;sudo&lt;/code&gt; printed &lt;code&gt;Password:&lt;/code&gt; and &lt;strong&gt;consumed the remaining pasted lines as password attempts&lt;/strong&gt;. Three failures, no commands run.&lt;/p&gt;

&lt;p&gt;The fix is to run &lt;code&gt;sudo -i&lt;/code&gt; alone, wait for the prompt to change from &lt;code&gt;$&lt;/code&gt; to &lt;code&gt;#&lt;/code&gt;, then paste the commands without &lt;code&gt;sudo&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  How do you prove it's actually gone?
&lt;/h2&gt;

&lt;p&gt;Files deleted, no errors — that is not proof.&lt;/p&gt;

&lt;p&gt;This thing is designed to come back when deleted. &lt;strong&gt;The real verification is surviving a full trigger cycle.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;It runs every 20 minutes, so I scheduled an automatic re-check 25 minutes out:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Cleanup   ~16:40
Re-check   17:13:46
Result     no indicators found
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A complete cycle passed and it did not return. &lt;strong&gt;That's what counts.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;(I also confirmed the outbound connection to an external IP was gone, no surviving processes, and the system crontab contained only legitimate entries. "It worked" declared immediately after deletion and "it worked" declared after one full cycle are claims of very different strength.)&lt;/p&gt;

&lt;h2&gt;
  
  
  The watcher I wrote doesn't check signatures
&lt;/h2&gt;

&lt;p&gt;Since the vendor's scanner is green on this thing, I can't use it as detection. So I wrote a daily check that does &lt;strong&gt;no signature matching&lt;/strong&gt;. It asks three questions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Does &lt;code&gt;/etc/.conf&lt;/code&gt; exist?&lt;/li&gt;
&lt;li&gt;Does &lt;code&gt;/etc/node&lt;/code&gt; exist?&lt;/li&gt;
&lt;li&gt;Does that marker string appear in the crontab or boot script?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Packing defeats a scanner. It doesn't defeat &lt;code&gt;ls&lt;/code&gt;.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I made a mistake writing it that's worth its own paragraph. In the first version, "clean" and "couldn't reach the host" both returned the same exit code and logged the same line.&lt;/p&gt;

&lt;p&gt;Which means: the day SSH breaks, or the NAS is powered off, or the key expires — that watcher goes &lt;strong&gt;permanently silent&lt;/strong&gt;, and I assume I'm being watched.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;When a monitor is silent, "everything is fine" and "I am blind" look identical.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The fix: distinct exit codes, plus &lt;strong&gt;after three consecutive unreachable runs it alerts that it has gone blind&lt;/strong&gt;. A watcher has to speak up about its own blindness, or its silence gets read as safety.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I took away
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. "Not detected" and "not there" are different sentences.&lt;/strong&gt;&lt;br&gt;
A tool's output is what that tool saw, not the state of the world. To know whether something exists, go look at the thing itself.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Convenience features are exposure.&lt;/strong&gt;&lt;br&gt;
QuickConnect saved me from configuring a router. The price was putting a management interface on the public internet. That trade was fine before the vulnerability was published, and not fine after — and I didn't re-evaluate.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Updates are not a "when I get around to it" task.&lt;/strong&gt;&lt;br&gt;
Over a year passed between the advisory and my compromise. Every time I saw the update prompt, I thought &lt;em&gt;"later — what if it breaks something."&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. "Assuming it's fine" is more dangerous than "knowing it's broken".&lt;/strong&gt;&lt;br&gt;
Six months, zero symptoms. No slowdown, nothing anomalous, and the official tool reporting all clear. The only reason I found it was that I asked a question about something else entirely.&lt;/p&gt;
&lt;h2&gt;
  
  
  If you also run a NAS
&lt;/h2&gt;

&lt;p&gt;Three things, ten minutes:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Patch to the latest build of your current version.&lt;/strong&gt; Not the newest major release — the newest &lt;em&gt;patch&lt;/em&gt; of what you're on. Lowest risk, most holes closed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Check whether your admin interface is reachable from the internet&lt;/strong&gt; — both router port-forwarding &lt;em&gt;and&lt;/em&gt; the vendor's remote-access service. Testing from a phone with WiFi off is the honest test.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Enable auto-block.&lt;/strong&gt; It wouldn't have stopped this particular attack, but it stops most credential attempts.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;And if you want to check whether you're clean:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-rn&lt;/span&gt; &lt;span class="s2"&gt;"/etc/&lt;/span&gt;&lt;span class="se"&gt;\.&lt;/span&gt;&lt;span class="s2"&gt;conf"&lt;/span&gt; /etc/crontab /etc/rc.subr 2&amp;gt;/dev/null
&lt;span class="nb"&gt;ls&lt;/span&gt; &lt;span class="nt"&gt;-la&lt;/span&gt; /etc/.conf /etc/node 2&amp;gt;/dev/null
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;No output is good. If you get output — you now know what it is.&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;I'm not a security researcher. I roast coffee and write software on the side.&lt;br&gt;
This is an incident log: an unpatched NAS, a vulnerability public for over a year, a program that sat there for six months, and a green light that said everything was fine.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;p&gt;&lt;em&gt;本文原載於我的部落格：&lt;a href="https://coffeeshooters.com/content/antivirus-said-clean-malware-was-right-there-en?utm_source=devto&amp;amp;utm_medium=social&amp;amp;utm_campaign=blog-antivirus-said-clean-malware-was-right-there-en" rel="noopener noreferrer"&gt;The antivirus said 'no malware detected'. It was sitting right there on the disk.&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>solodev</category>
    </item>
    <item>
      <title>NVIDIA's CEO says future companies will be built on harness engineering. Mine has been for six months — here's the half he left out</title>
      <dc:creator>Dexterlung</dc:creator>
      <pubDate>Tue, 18 Aug 2026 13:05:15 +0000</pubDate>
      <link>https://dev.to/dexterlung/nvidias-ceo-says-future-companies-will-be-built-on-harness-engineering-mine-has-been-for-six-219</link>
      <guid>https://dev.to/dexterlung/nvidias-ceo-says-future-companies-will-be-built-on-harness-engineering-mine-has-been-for-six-219</guid>
      <description>&lt;p&gt;On July 8th, Jensen Huang sat down with LangChain's founder and said this:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Today most companies are built on business processes. In the future, most companies will be built on harnesses.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The harness is the layer you wrap around the model: the domain knowledge you inject, the workflows you've refined, the guardrails, the memory, the evals, the runtime. His argument is that models keep getting better, but &lt;strong&gt;what defines a company is that outer shell&lt;/strong&gt; — because that's where the part only you have lives.&lt;/p&gt;

&lt;p&gt;My first reaction wasn't "I should start doing this." It was "oh, so that's what it's called."&lt;/p&gt;

&lt;p&gt;Because I've been doing it for six months. And I know that the feeling of &lt;em&gt;"he's describing me"&lt;/em&gt; is exactly the thing to distrust.&lt;/p&gt;




&lt;h2&gt;
  
  
  Admit the bias first, then continue
&lt;/h2&gt;

&lt;p&gt;I'm a one-person studio. I spent six months building a governance layer around my own coffee e-commerce system. Along the way I periodically felt I was over-engineering — guardrails don't sell more bags of beans. So when the person who understands AI infrastructure better than almost anyone says that layer &lt;em&gt;is&lt;/em&gt; the company, of course I want to believe him.&lt;/p&gt;

&lt;p&gt;Wanting to believe it is the signal to go verify it.&lt;/p&gt;

&lt;p&gt;I used a dumb test: &lt;strong&gt;take his list of components and check them off against what I actually have.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Frontier model plus cheap-model routing — yes&lt;/li&gt;
&lt;li&gt;The harness itself — yes. One CLAUDE.md, 35 hard rules, plus close to a hundred task-scoped skills&lt;/li&gt;
&lt;li&gt;Tools — yes. 145 scripts&lt;/li&gt;
&lt;li&gt;Memory, working and long-term — yes. A memory directory and an index that survive across conversations&lt;/li&gt;
&lt;li&gt;Knowledge graph — yes. Nine thousand-odd nodes&lt;/li&gt;
&lt;li&gt;Guardrails — yes, and this is my thickest slot. Pre-commit blockers, hooks, audit scripts&lt;/li&gt;
&lt;li&gt;Evals — yes. One &lt;code&gt;audit:all&lt;/code&gt;; if it fails I can't commit&lt;/li&gt;
&lt;li&gt;Blueprints and templates — yes&lt;/li&gt;
&lt;li&gt;Post-training the model against the harness — no, and I shouldn't (that needs a GPU cluster and a team)&lt;/li&gt;
&lt;li&gt;Runtime sandbox and access control — blueprint exists, zero real-customer validation&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Seven of nine.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Then I realized that matching the checklist means almost nothing.&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  The first thing he left out: half my components never get called
&lt;/h2&gt;

&lt;p&gt;I have close to a hundred skills. Each one is a situation I got burned by once, packaged into "here's how to judge this next time."&lt;/p&gt;

&lt;p&gt;One day I measured &lt;strong&gt;how often they actually fire.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;More than half had never fired. Not once.&lt;/p&gt;

&lt;p&gt;The content wasn't wrong. They &lt;strong&gt;never reached the model&lt;/strong&gt; — I'd written each description too thoroughly, the total exceeded the character budget reserved for descriptions, and &lt;strong&gt;everything over the limit was silently dropped&lt;/strong&gt;. No error. No warning. Those skills sat happily on disk, present, correct, and — as far as the moment of decision was concerned — nonexistent.&lt;/p&gt;

&lt;p&gt;Jensen's list asks &lt;em&gt;do you have this component?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;What six months taught me is: &lt;strong&gt;"the component exists" and "the component gets used when it should" are different things, and only the second one is a harness.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The first is an asset inventory. The second is engineering.&lt;/p&gt;




&lt;h2&gt;
  
  
  The second thing he left out: guardrails block paths, not failure modes
&lt;/h2&gt;

&lt;p&gt;There's one bug I fixed four times.&lt;/p&gt;

&lt;p&gt;PostgreSQL lets you have multiple same-named functions with different signatures — overloads. I changed a function, added a new parameter with a default, and assumed that counted as "updating" it. What actually happened is the database now held a new version and the old one was still there. Later an API call came in, both versions matched, PostgreSQL refused to guess and errored (&lt;code&gt;42725: function is not unique&lt;/code&gt;). Production 400.&lt;/p&gt;

&lt;p&gt;The first time, I fixed it. Then I wrote a rule: drop the old signature before changing one.&lt;br&gt;
The second time the same bug came back through a different entry point. I added a static check that scans migration files.&lt;br&gt;
The third time it came back again.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The fourth time was six days after I finished writing that guard — and it walked in through the door I had locked myself.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That's when I finally saw the shape of the problem. My check scanned &lt;strong&gt;the files in my repo.&lt;/strong&gt; It blocks "the path I hit last time." But an old migration re-applied once, an ordering I didn't anticipate, a change arriving from somewhere else — the check sees none of it, because none of that happens inside my files.&lt;/p&gt;

&lt;p&gt;I was blocking &lt;strong&gt;paths.&lt;/strong&gt; The thing I needed to block was the &lt;strong&gt;failure mode&lt;/strong&gt;: how many versions does the database actually hold right now.&lt;/p&gt;

&lt;p&gt;What finally pinned it was making the migration ask the database itself — query &lt;code&gt;pg_proc&lt;/code&gt;, and throw on the spot if the answer isn't exactly one.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;This is the most expensive sentence I've learned in six months: &lt;strong&gt;ask the load-bearing thing itself. Don't ask the description of it you happen to be holding.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;As for how the old migration got re-applied — I couldn't find out, and I gave up on reconstructing it. This post doesn't have a clean ending.&lt;/p&gt;




&lt;h2&gt;
  
  
  The third thing (the costly one): I verified it, and it still exploded
&lt;/h2&gt;

&lt;p&gt;Jensen mentions evals in one line: they're the key prerequisite for running agents at scale inside an enterprise.&lt;/p&gt;

&lt;p&gt;True, but that sentence is too light. The hard part isn't &lt;em&gt;whether you have evals.&lt;/em&gt; It's this —&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The worst failure isn't a crash. It's the error your verifier blesses.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I have a discipline: &lt;strong&gt;a newly written acceptance check must first prove it is currently red.&lt;/strong&gt; Because a check that can never go red is the same as no check, only worse — it emits a green light.&lt;/p&gt;

&lt;p&gt;I was pleased with that discipline. Then, within a single day, I hit the same wall three times, and &lt;strong&gt;every single time I had verified before shipping&lt;/strong&gt;:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;I verified the code &lt;strong&gt;parsed.&lt;/strong&gt; Deployed it, and it threw &lt;code&gt;ReferenceError&lt;/code&gt; at runtime — the identifier didn't resolve. Valid syntax is not "it runs."&lt;/li&gt;
&lt;li&gt;I verified the &lt;strong&gt;source&lt;/strong&gt; was correct. But my build command had a pipeline stage that spliced the verifier's own "✅ all passed" line &lt;em&gt;into the artifact.&lt;/em&gt; Clean source, &lt;code&gt;SyntaxError&lt;/code&gt; in the thing I actually shipped.&lt;/li&gt;
&lt;li&gt;I verified &lt;strong&gt;the string I was holding&lt;/strong&gt; was correct. The string actually stored on the server was a different one.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The shared shape isn't "forgot to verify."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It's verifying the wrong object.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;And here's the part that stings: when it happened, I &lt;strong&gt;did cite&lt;/strong&gt; the must-be-red-first discipline, and I did watch it go red first — red on the syntax.&lt;/p&gt;

&lt;p&gt;So the correct reading of that discipline isn't "make it red first." It's:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Make it red first against the thing that actually bears the load.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;"Invoked the rule" and "aimed the rule at the right object" are two different things. And the gap between them is one the self-discipline layer &lt;strong&gt;cannot detect on its own&lt;/strong&gt; — because being diligent about the layer above feels exactly like being diligent.&lt;/p&gt;




&lt;h2&gt;
  
  
  Postscript: the same day I drafted this, I did it a fourth time
&lt;/h2&gt;

&lt;p&gt;About two hours after writing that section, I was doing something unrelated: measuring how many of my 142 published articles give a reader a path to my services page.&lt;/p&gt;

&lt;p&gt;I opened a terminal and grepped the folder where my articles live for the services URL.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;0 out of 41.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I wrote it up as a finding — "you have a 41-article funnel with no hole in the bottom" — put it in my governance doc, committed it, and attached a recommendation: highest-return fix available, go add the exit to 41 articles.&lt;/p&gt;

&lt;p&gt;Both numbers were wrong.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;First, the exit doesn't live in the article content. It lives in the render template&lt;/strong&gt; — which had already been mounting the right exit per article category for three months. Actual coverage: 138 of 142.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Second, 41 was the wrong denominator too.&lt;/strong&gt; The folder I grepped only holds the articles that happen to have a backup file. The database has 142 published.&lt;/p&gt;

&lt;p&gt;Both errors came from one action: &lt;strong&gt;I measured a proxy.&lt;/strong&gt; The markdown in that folder is a backup of the article, not the page a reader sees. And the exit is part of the page, not part of the article.&lt;/p&gt;

&lt;p&gt;There's one difference between this time and the previous three, and it's the worse kind:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;I committed the wrong observation into my single source of truth.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The entire purpose of that document is so future-me — or anyone taking over — doesn't have to re-investigate to know the current state. So the next person to read it would go make 41 unnecessary edits, solving a problem that was solved three months ago. And they wouldn't question it, because it looks empirically measured. It &lt;em&gt;was&lt;/em&gt; empirically measured. It just measured the wrong object.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A wrong observation doesn't stay put. A wrong observation written into your source of truth propagates itself.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That's four times on the same rake, the fourth inside two hours of writing "I've learned not to step on this rake." So I'm not going to end this section by telling you I solved it. I can tell you I now know its shape, and that knowing the shape is evidently not enough.&lt;/p&gt;




&lt;h2&gt;
  
  
  So where is the hard part of a harness
&lt;/h2&gt;

&lt;p&gt;Not the component list. Lists can be bought, copied, or assembled from a launch-event blueprint.&lt;/p&gt;

&lt;p&gt;The hard part is &lt;strong&gt;one sentence you have to be able to answer for every gate in your system that emits a green light&lt;/strong&gt;:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What load-bearing, real thing did this green light positively observe?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If you can't answer, it's decoration. And a decorative green light is more dangerous than no green light, because it makes you ship.&lt;/p&gt;

&lt;p&gt;I now treat three kinds of green as guilty until proven otherwise:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;"No error seen, therefore green"&lt;/strong&gt; — that proves "I saw no bad news," not "the thing is good"&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Verifies the proxy, not the artifact&lt;/strong&gt; — verifies syntax (proxy) not execution (artifact); verifies source (proxy) not what was deployed (artifact)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Several checks that aren't actually independent&lt;/strong&gt; — three gates all reading the same possibly-stale snapshot is one gate wearing three coats&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  An honest footnote from a one-person company
&lt;/h2&gt;

&lt;p&gt;There's another line in that conversation that points the opposite way from my bet.&lt;/p&gt;

&lt;p&gt;Jensen says: &lt;strong&gt;the more AI you use, the more people you end up hiring&lt;/strong&gt; — because building agents is a whole new skill, and someone has to do the evals, the benchmarks, the guardrails.&lt;/p&gt;

&lt;p&gt;My bet is precisely "one person plus a harness." His own company found this work to be labor-intensive.&lt;/p&gt;

&lt;p&gt;I'm not going to pretend I've proven him wrong. My evidence today is: the same machine now runs automation for two different brands, and the second one only needed config filled in — no code changes.&lt;/p&gt;

&lt;p&gt;But &lt;strong&gt;both of those brands are mine.&lt;/strong&gt; No one else's data, no one else's permissions, no one else's expectations. The real test is &lt;strong&gt;how many hours the first real client's cold start actually took&lt;/strong&gt; — and I don't have that number yet.&lt;/p&gt;

&lt;p&gt;Until I do, "one person handling 30 projects" is an inference, not a conclusion. I think labeling that honestly matters more than telling a clean story.&lt;/p&gt;




&lt;h2&gt;
  
  
  If you take one thing from this
&lt;/h2&gt;

&lt;p&gt;Don't start from the list. The list makes you feel like you're progressing — add a script, add a rule, fill another slot.&lt;/p&gt;

&lt;p&gt;Start here: &lt;strong&gt;pick one gate in your system that is emitting green right now, and write one sentence describing what it positively observed.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The ones you can't write that sentence for are where your actual risk is. They've been telling you everything is fine.&lt;/p&gt;

&lt;p&gt;And if you work alone like I do — no colleague is going to catch this for you. You're the one emitting the green light, and you're the one who believes it.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;本文原載於我的部落格：&lt;a href="https://coffeeshooters.com/content/harness-engineering-solo-company-en?utm_source=devto&amp;amp;utm_medium=social&amp;amp;utm_campaign=blog-harness-engineering-solo-company-en" rel="noopener noreferrer"&gt;NVIDIA's CEO says future companies will be built on harness engineering. Mine has been for six months — here's the half he left out&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>solodev</category>
    </item>
    <item>
      <title>Dual-Graph Drift Detection for Solo Devs: What Happens When Your Docs and Your Code Start Talking to Each Other</title>
      <dc:creator>Dexterlung</dc:creator>
      <pubDate>Sun, 16 Aug 2026 13:05:11 +0000</pubDate>
      <link>https://dev.to/dexterlung/dual-graph-drift-detection-for-solo-devs-what-happens-when-your-docs-and-your-code-start-talking-522h</link>
      <guid>https://dev.to/dexterlung/dual-graph-drift-detection-for-solo-devs-what-happens-when-your-docs-and-your-code-start-talking-522h</guid>
      <description>&lt;h2&gt;
  
  
  Dual-Graph Drift Detection for Solo Devs
&lt;/h2&gt;

&lt;h2&gt;
  
  
  What Happens When Your Docs and Your Code Start Talking to Each Other
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;I'm a coffee roaster in Taiwan who taught myself to code with AI over the past 8 months. I built a full vertical-integration ERP for my coffee brand — from green bean inventory to roasting orders to e-commerce checkout. Along the way I ran into something nobody on the internet seems to be writing about: a way for the **prose&lt;/em&gt;* I've written (docs, specs, governance rules) and the &lt;strong&gt;code&lt;/strong&gt; I've shipped to literally audit each other. This is that workflow.*&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  The 14-Hour Origin of My "Governance Awareness"
&lt;/h2&gt;

&lt;p&gt;People ask me, &lt;em&gt;"When did you start caring about governance?"&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;I can answer with 14-hour precision.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;April 21, 2026, 11:41 AM.&lt;/strong&gt; That was the first time I wrote &lt;code&gt;governance-guard&lt;/code&gt; into a prompt to Claude Code:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;"Should we put these 3 audit rules into governance-guard? Yes. Then can you do a targeted scan based on these and report potential errors back to me?"&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Later that afternoon at 15:08, I doubled down:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;"Option B (CI layer): governance-guard adds a rule to scan all relative / alias import paths... fail = block ——&amp;gt; do this."&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;14 hours later, April 22, 1:12 AM&lt;/strong&gt;, I used the word &lt;em&gt;"governance"&lt;/em&gt; for the first time:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;"Should we save all the fixes and rules we discussed as a 'skill', so future projects can reuse them? I'm worried that once my attention returns to MVP launch, all this governance work will be forgotten and turn into hidden risk. I don't know how a solo developer is supposed to allocate this kind of work."&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;In the same Claude Code session (&lt;code&gt;70aa7ce0&lt;/code&gt;), in those 14 hours, three governance artifacts were born:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;governance-guard.mjs&lt;/code&gt;&lt;/strong&gt; — a CI script that blocks pushes on rule violations&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;TECH-DEBT-BACKLOG.md&lt;/code&gt;&lt;/strong&gt; — known violations don't get fixed immediately; they get queued for me to triage&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;skill&lt;/code&gt; files&lt;/strong&gt; — a way to make rules survive across sessions and across projects&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These 14 hours are the precise moment my governance awareness was born. Before this, I asked AI &lt;em&gt;"fix this."&lt;/em&gt; After this, I asked &lt;em&gt;"fix this *&lt;/em&gt;+ prevent it from happening again*&lt;em&gt;."&lt;/em&gt; The question went from one-dimensional to two-dimensional.&lt;/p&gt;

&lt;h2&gt;
  
  
  That Month: 348 Commits, Fixing the Same Bug Five Times
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;April 2026 was the most brutal month of my 8 months of development&lt;/strong&gt; — I accumulated 348 git commits, the highest single-month count in the project's history.&lt;/p&gt;

&lt;p&gt;It went like this: I'd fix one order bug, fix it, realize the same pattern showed up elsewhere, fix that, fix again, then the first one would break.&lt;/p&gt;

&lt;p&gt;The most dramatic prompt came at 8:39 AM on April 29:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;"I'm dumbfounded. Why is the previous balance showing 0? We've fixed it five or six versions already — go look at git. This should be a really simple problem for you. Why has it taken so many tries?"&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That wasn't a technical question. &lt;strong&gt;It was the scream of a governance problem.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The root cause (which I later turned into one of 14 data-drift debugging records):&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;What was written&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;CLAUDE.md&lt;/code&gt; Rule 21 (spec)&lt;/td&gt;
&lt;td&gt;SSOT — frontend may not override DB truth; &lt;code&gt;single_price&lt;/code&gt; is the canonical source&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;single_price&lt;/code&gt; (DB schema)&lt;/td&gt;
&lt;td&gt;Exists&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;useOrderAggregation.js&lt;/code&gt; (code)&lt;/td&gt;
&lt;td&gt;Uses `product.single_price \&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;{% raw %}&lt;code&gt;Reports.vue&lt;/code&gt; (code)&lt;/td&gt;
&lt;td&gt;Uses a &lt;em&gt;different&lt;/em&gt; legacy fallback chain&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;AdminPOS.vue&lt;/code&gt; (code)&lt;/td&gt;
&lt;td&gt;Recalculates totals on the frontend&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;The docs said "single_price is SSOT", the code used a 5-layer fallback, and the UI was still recalculating.&lt;/strong&gt; Fix one location, the other four kept dropping. By the fifth attempt I realized — &lt;strong&gt;I'm not fixing a bug, I'm fixing a governance problem&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;That prompt became the entire motivation for &lt;code&gt;CLAUDE.md&lt;/code&gt; Rule 17 Step 0: &lt;em&gt;"For value-mismatch bugs, the first step MUST be a SQL query against the DB. Do not guess in code."&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The "Governance Terminology Birth Engine"
&lt;/h2&gt;

&lt;p&gt;Looking back at that month's prompts, I noticed a very consistent pattern in how I asked things. I call it my &lt;strong&gt;governance terminology birth engine&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The canonical example (also April 29, 6:03 AM):&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;"Help me find all of them. Is there an industry term for this phenomenon? Use it to do a comprehensive search."&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The structure is always the same:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;I observe a symptom&lt;/strong&gt; (button doesn't respond, order totals don't match, bean inventory off by a kilogram)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;What's the term?&lt;/strong&gt; (literally ask AI to name the phenomenon)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Now scan for all instances using that name&lt;/strong&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Every governance term in my codebase was born this way — &lt;code&gt;skill&lt;/code&gt; files, Type A/B/C drift classification, Layer 1-6 defense, SSOT, &lt;code&gt;governance-guard&lt;/code&gt;, &lt;code&gt;TECH-DEBT-BACKLOG&lt;/code&gt;, "data-drift debugging records." Each name has a prompt like this behind it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Converting daily pain into indexable abstract terminology&lt;/strong&gt; — that's the seed crystal that governance systems grow around. I didn't suddenly understand governance one day. I had 8 months of AI repeatedly asking me, &lt;em&gt;"Does the concept you're describing have a name in your project?"&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Same Day, Two Opposite Attitudes Toward AI
&lt;/h2&gt;

&lt;p&gt;What's more interesting — &lt;strong&gt;on the same day, I had two completely opposite attitudes toward AI&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;April 21, 2:20 AM:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;"Please... synthesize everything you're about to say to me into markdown — but as questions you ask me back, not as answers. Let's go from divergent to convergent."&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This is the &lt;strong&gt;soft inquiry mode&lt;/strong&gt; — treat AI as Socrates, force me to think through it myself.&lt;/p&gt;

&lt;p&gt;Same day, 11:41 AM:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;"Should we put these 3 audit rules into governance-guard? Yes."&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This is the &lt;strong&gt;hard blocking mode&lt;/strong&gt; — use a CI script to fence AI (and my future self) out of mistakes.&lt;/p&gt;

&lt;p&gt;That's when I realized — &lt;strong&gt;two opposite attitudes on the same day, mapping exactly to the human need for freedom vs. discipline&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;For &lt;strong&gt;business direction, persona refinement, brand positioning&lt;/strong&gt; (divergent work): I want AI to question me (soft)&lt;/li&gt;
&lt;li&gt;For &lt;strong&gt;code governance, SSOT maintenance, file contracts&lt;/strong&gt; (convergent work): I want a CI script to block AI (hard)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The greatest discipline of a one-person founder — &lt;strong&gt;knowing when to loosen the reins and when to lock them down&lt;/strong&gt;. This dual-attitude toward AI is the most precious methodology I've developed in 8 months.&lt;/p&gt;

&lt;h2&gt;
  
  
  Geology of Rules: Layer 6 of Rule 21 Was Born at May 2, 3:14 AM
&lt;/h2&gt;

&lt;p&gt;My &lt;code&gt;CLAUDE.md&lt;/code&gt; has 26 rules. Rule 21 ("SSOT Five-Layer Defense") is the most-cited one. But you should know — &lt;strong&gt;it originally had only 4 layers&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;I can date the exact moment it gained a 6th: &lt;strong&gt;May 2, 2026, 3:14 AM&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;That day I was fixing a roast-degree display bug — the orders list mixed Chinese and English: "medium roast", "medium", "medium_light", "medium_dark". The same field was being handled three different ways across RPC, normalizeOrder, and UI — classic drift.&lt;/p&gt;

&lt;p&gt;The prompt:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;"Should I write this skill and update CLAUDE.md Rule 21 to add Layer 6? Yes, do it. Also, I just noticed — why are some roast degrees in Chinese and some in English? They're clearly not being read from the SSOT."&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That was the moment of expansion from 4 to 6 layers. The two new layers:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Layer 5: RPC ↔ normalizeOrder ↔ UI three-tier field contract&lt;/strong&gt; (born from the balance_before = 0 bug)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Layer 6: Cross-layer data contract propagation audit&lt;/strong&gt; (born from the roast-degree mixed-language bug)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;28 minutes later at 3:42 AM, I asked:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;"Is there a way to do a comprehensive scan that proactively finds where else this kind of problem might occur?"&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;Rules grow in geological layers.&lt;/strong&gt; Each layer is a specific frustration that crystallized into a fossil. Layer 5 is the pain of fixing balance_before five times. Layer 6 is the pain of roast degrees being handled three different ways.&lt;/p&gt;

&lt;p&gt;If you see my CLAUDE.md grow a Layer 7 or Layer 8 in the future — &lt;strong&gt;that will absolutely be another specific frustration that forced it out&lt;/strong&gt;. Not because I read a "treasure hunt defense design guide" first.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Real Solution Isn't "Fix Bugs", It's "Reconcile"
&lt;/h2&gt;

&lt;p&gt;Starting that month I did three things:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Built &lt;code&gt;governance-guard.mjs&lt;/code&gt;&lt;/strong&gt; — a CI script that scans code for violations of &lt;code&gt;CLAUDE.md&lt;/code&gt; rules&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Wrote 14 data-drift debugging records&lt;/strong&gt; — classified each bug into Type A / B / C / D / E drift&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;For every bug I fixed&lt;/strong&gt;, I now ask: &lt;em&gt;"Is this bug just the surface symptom of an SSOT violation?"&lt;/em&gt; If yes — fix upstream, add a governance rule&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;But it wasn't enough. governance-guard can only scan &lt;strong&gt;known&lt;/strong&gt; violations — the ones I've written rules for. &lt;strong&gt;Drifts I haven't written into rules are invisible to it&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The deeper problem is — &lt;strong&gt;even rules I've written, I forget&lt;/strong&gt;. 6 months ago I wrote Rule 9 (no raw &lt;code&gt;system_role&lt;/code&gt; compare), but in a new Claude session this month, neither I nor the AI remembered.&lt;/p&gt;

&lt;h2&gt;
  
  
  Looking Only at Code, You Miss "Intent Has Changed"
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;GitNexus&lt;/strong&gt; is a powerful tool. It turns my codebase into a graph of &lt;strong&gt;23,280 symbols / 45,727 edges&lt;/strong&gt; — every function, every call, every import.&lt;/p&gt;

&lt;p&gt;Ask it, &lt;em&gt;"What's the blast radius of &lt;code&gt;fn_create_order_atomic&lt;/code&gt;?"&lt;/em&gt; — perfect answer, lists which functions break at distance d=1, d=2, d=3.&lt;/p&gt;

&lt;p&gt;But ask it, &lt;em&gt;"Is &lt;code&gt;CLAUDE.md&lt;/code&gt; Rule 9 (no raw &lt;code&gt;system_role&lt;/code&gt; compare) still being respected, 6 months after I wrote it?"&lt;/em&gt; — &lt;strong&gt;it doesn't know that rule exists&lt;/strong&gt;. It only sees code, not doc-intent.&lt;/p&gt;

&lt;p&gt;This is the ceiling of code-graphs — &lt;strong&gt;they can only tell you "what code is doing now", not "how much that has drifted from what I originally intended"&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Looking Only at Docs, You Miss "Code Quietly Never Implemented It"
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Graphify&lt;/strong&gt; (&lt;a href="https://dev.to/content/graphify-knowledge-graph-from-markdown"&gt;covered in detail in T10&lt;/a&gt;) turns my 50+ &lt;code&gt;.md&lt;/code&gt; files into a graph of &lt;strong&gt;7,230 concept nodes&lt;/strong&gt;. It helps me find cross-file surprising connections I never consciously made.&lt;/p&gt;

&lt;p&gt;But ask it, &lt;em&gt;"This RPC I designed 2 months ago — did the code actually return the fields I said it would?"&lt;/em&gt; — &lt;strong&gt;it can't answer&lt;/strong&gt;. It only sees docs, not whether code quietly never implemented them.&lt;/p&gt;

&lt;p&gt;This is the ceiling of doc-graphs — &lt;strong&gt;they can only tell you "what intent I wrote", not "whether the code caught up"&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Even Graphify Misses Systemic Problems
&lt;/h2&gt;

&lt;p&gt;The most honest admission — May 15, 2026 morning, I was fixing a roasting batch display anomaly when I said:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;"This isn't a single-batch problem. This is **systemic data drift&lt;/em&gt;&lt;em&gt;. Why didn't the previous root-cause analysis catch this serious failure? **Did graphify and gitnexus both fail to help?&lt;/em&gt;&lt;em&gt;"&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That moment I realized — &lt;strong&gt;graphify can catch "the same concept appearing similarly in different files", but it can't catch "a structural data drift distributed across 5 files as 5 different symptoms"&lt;/strong&gt;. Each file only contains part of the symptom; there's no complete pattern to detect.&lt;/p&gt;

&lt;p&gt;This is why &lt;strong&gt;just chaining two graphs isn't enough&lt;/strong&gt; — you need a third layer, &lt;code&gt;governance-guard&lt;/code&gt;, which explicitly writes "symptom patterns" as rules to enable reverse scanning.&lt;/p&gt;

&lt;h2&gt;
  
  
  A Live Case That Happened While Writing This Article
&lt;/h2&gt;

&lt;p&gt;While drafting this article, &lt;strong&gt;I just realized my admin UI is missing the "assign series" interface&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;My blog post schema (&lt;code&gt;content_pages&lt;/code&gt; table) &lt;strong&gt;has had &lt;code&gt;series_slug&lt;/code&gt; and &lt;code&gt;series_order&lt;/code&gt; columns for weeks&lt;/strong&gt;. I've manually filled in values for 4 posts (via SQL console direct UPDATE).&lt;/p&gt;

&lt;p&gt;But the admin content management page &lt;strong&gt;has no UI for editing series&lt;/strong&gt;. I'd been changing it via SQL the whole time. I literally forgot, "oh, I never built that UI."&lt;/p&gt;

&lt;p&gt;This is a 100% live case of doc-as-spec × code-as-reality drift:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;State&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;DB schema&lt;/td&gt;
&lt;td&gt;✅ &lt;code&gt;series_slug&lt;/code&gt; + &lt;code&gt;series_order&lt;/code&gt; exist&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DB data&lt;/td&gt;
&lt;td&gt;✅ 4 posts have values (filled manually via SQL)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Admin UI&lt;/td&gt;
&lt;td&gt;❌ No edit field built&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Public series cards (frontend)&lt;/td&gt;
&lt;td&gt;❌ Probably also not built&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;If graphify + gitnexus + governance-guard were chained and auto-reconciling&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;graphify catches the &lt;code&gt;series&lt;/code&gt; concept appearing in &lt;code&gt;.md&lt;/code&gt; / spec files&lt;/li&gt;
&lt;li&gt;gitnexus detects no admin UI component references &lt;code&gt;series_slug&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;governance-guard auto-flags: &lt;em&gt;"Schema has &lt;code&gt;series_slug&lt;/code&gt;, but 0 admin UI components reference it"&lt;/em&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This would've been blocked the first time. &lt;strong&gt;I wouldn't have only noticed today that "oh, I never built that UI."&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Another Meta Moment: cp950 Codepage Failure
&lt;/h2&gt;

&lt;p&gt;Mid-article, I tried re-running &lt;code&gt;npx gitnexus analyze&lt;/code&gt; to get the latest numbers. &lt;strong&gt;First attempt failed instantly&lt;/strong&gt; — DuckDB's &lt;code&gt;COPY&lt;/code&gt; command on Windows with a Chinese path tried to encode in &lt;code&gt;cp950&lt;/code&gt; and crashed.&lt;/p&gt;

&lt;p&gt;My reaction (with AI's help) was &lt;em&gt;"oh, must be UTF-8 issue"&lt;/em&gt; — ran &lt;code&gt;chcp 65001 + LC_ALL=UTF-8 + Console.OutputEncoding=UTF8&lt;/code&gt; three-layer enforcement. &lt;strong&gt;Still failed&lt;/strong&gt;, because DuckDB internal IO looks at Windows system locale, not console codepage.&lt;/p&gt;

&lt;p&gt;Finally I checked memory and realized — &lt;strong&gt;I had solved this exact problem 17 days earlier&lt;/strong&gt;, using an NTFS junction (&lt;code&gt;C:\gn-cssaas&lt;/code&gt;) to give the project an ASCII-only path.&lt;/p&gt;

&lt;p&gt;The most ironic part — &lt;strong&gt;my graphify-out vault already had a community labeled 183 documenting this&lt;/strong&gt;:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;Gitnexus zh-TW cp950 Codepage Failure → chcp 65001 + LC_ALL UTF-8 fix&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If doc↔code drift detection were running, &lt;strong&gt;today's first error would have been caught by graphify, surfacing the message "you wrote a solution 17 days ago in &lt;code&gt;reference_gitnexus_setup.md&lt;/code&gt;"&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;What actually happened — &lt;strong&gt;I (with AI) just walked into the trap. Failure #1, failure #2, then checked memory for the solution&lt;/strong&gt;. 20 minutes of tuition.&lt;/p&gt;

&lt;p&gt;This is the real cost of dev-time drift — &lt;strong&gt;not the big-bang bugs, but the 20-minutes × 365 days = 122 hours/year of low-grade friction&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Four Drift Types That Stacked Graphs Can Detect
&lt;/h2&gt;

&lt;p&gt;Chain graphify (doc-graph) + gitnexus (code-graph) + governance-guard (auditor), and &lt;strong&gt;you can auto-detect four kinds of drift&lt;/strong&gt;:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Drift Type&lt;/th&gt;
&lt;th&gt;Example&lt;/th&gt;
&lt;th&gt;Detection Method&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Intention drift&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Docs say it exists, code didn't implement it&lt;/td&gt;
&lt;td&gt;doc-graph has node + code-graph empty&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Undocumented knowledge&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Code works fine, docs never described it&lt;/td&gt;
&lt;td&gt;code-graph has node + doc-graph empty&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Concept gap&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Docs link A→B, code never imports&lt;/td&gt;
&lt;td&gt;doc-graph has edge + code-graph missing edge&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Hidden complexity&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Code has complex flow, docs give zero explanation&lt;/td&gt;
&lt;td&gt;code-graph node degree high + doc-graph corresponding node low&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;All four I've stepped on. &lt;strong&gt;Series UI absence = Intention drift&lt;/strong&gt;. &lt;strong&gt;cp950 solution forgotten = Undocumented knowledge&lt;/strong&gt; (solution was in &lt;code&gt;memory.md&lt;/code&gt; but &lt;code&gt;governance-guard&lt;/code&gt; didn't know about it).&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Nobody Else Online Is Writing About This Combo
&lt;/h2&gt;

&lt;p&gt;I did web research. Conclusion: &lt;strong&gt;the community isn't doing this&lt;/strong&gt;. Someone is doing code-graph routing (&lt;a href="https://www.sidharthsatapathy.com/blog/gitnexus-dual-graph-engine-token-savings/" rel="noopener noreferrer"&gt;Sidharth Satapathy's 17-agent crew&lt;/a&gt; uses dual-graph for problem routing), but &lt;strong&gt;nobody is doing drift detection&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Several structural reasons:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Most engineers don't write docs&lt;/strong&gt; — no doc-graph to reconcile against&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;In large companies, specs are written by other people&lt;/strong&gt; — engineers don't trust them, no incentive to reconcile&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Both tools are new&lt;/strong&gt; — GitNexus + Graphify both emerged in 2026; the user intersection is small&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No economic incentive&lt;/strong&gt; — drift reconciliation is the hardest value to quantify; no startup makes it a selling point&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Solo devs have "self-written docs" as a unique asset&lt;/strong&gt; — large-company docs have low trust, open-source docs are scarce&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;"Governance" isn't a familiar concept in solo circles&lt;/strong&gt; — most solo devs see governance as a big-company thing&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Stack these 6 conditions. The number of people who simultaneously have all 6 isn't large.&lt;/strong&gt; I'm not discovering a new continent. &lt;strong&gt;I just happen to stand in the only spot where this landscape is visible.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  My Workflow Ritual
&lt;/h2&gt;

&lt;p&gt;How I'm trying to make this combo run (incomplete, still experimenting):&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Every &lt;code&gt;git commit&lt;/code&gt;&lt;/strong&gt; → triggers a graphify hook, creates a flag in &lt;code&gt;graphify-out/needs_update&lt;/code&gt;, the ritual reminds me to re-run&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Every new Claude session&lt;/strong&gt; → first thing the ritual does is check &lt;code&gt;TECH-DEBT-BACKLOG.md&lt;/code&gt; high-priority items + &lt;code&gt;needs_update&lt;/code&gt; flag&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Every code symbol change&lt;/strong&gt; → I'm required to run &lt;code&gt;gitnexus_impact&lt;/code&gt; first to see blast radius&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Weekly&lt;/strong&gt; → run &lt;code&gt;governance-guard.mjs&lt;/code&gt; to see new violations&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Every schema change&lt;/strong&gt; → simultaneously run graphify (does spec/&lt;code&gt;.md&lt;/code&gt; mention this concept?) and gitnexus (are there UI components referencing this field?)&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Honestly, &lt;strong&gt;this workflow isn't fully formed yet&lt;/strong&gt;. Item 5 I only realized I needed today (because of the series UI absence). &lt;strong&gt;But the scaffolding is there&lt;/strong&gt;, what's missing is the orchestration script that wires three tool outputs together.&lt;/p&gt;

&lt;h2&gt;
  
  
  Governance Has Evolved Beyond Rules — Into a Skill Network
&lt;/h2&gt;

&lt;p&gt;Walking through the 4-22 governance awakening, the 4-29 "I'm dumbfounded" moment, and the 5-02 Layer 6 expansion, my definition of governance has evolved.&lt;/p&gt;

&lt;p&gt;The most dramatic moment was 2026-05-15 afternoon:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;"Chain these three governance skills together: data-flow-audit + data-contract-propagation-audit + workstation-button-interaction — the same production bug simultaneously triggered all three abstract governance skills."&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That moment I realized — &lt;strong&gt;governance is no longer about "how many rules I've written"&lt;/strong&gt;. It's about &lt;strong&gt;"how many of my written skills the same bug triggers simultaneously"&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Drawn as a diagram:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;              Production Bug
                   │
        ┌──────────┼──────────┐
        ↓          ↓          ↓
  data-flow-   data-contract  workstation-
  audit.md   propagation     button-
             -audit.md       interaction.md
        ↓          ↓          ↓
        └─── 3 governance skills resonate ──┘
                       ↓
              Fix once, prevent thrice
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Graphify already knew, during community detection, that these three skills belong to the same community. They share &lt;code&gt;surprising_similar_to&lt;/code&gt; edges. &lt;strong&gt;But only when I query "what skills relate to this bug?" does graphify return all three&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Put another way — the highest form of governance is &lt;strong&gt;not the count of rules, but the network resonance between them&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Honest Costs
&lt;/h2&gt;

&lt;p&gt;I also have to say — &lt;strong&gt;this combo isn't cheap&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Graphify&lt;/strong&gt; ~7,000 tokens per full corpus run (~$0.05 USD / run)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;GitNexus&lt;/strong&gt; is free (self-hosted MCP), but you have to learn a new tool + maintain an MCP server&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;governance-guard&lt;/strong&gt; is self-written script; each new rule costs writing + testing + CI integration&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Correlation layer that chains all three&lt;/strong&gt; I still haven't built — probably 200-500 lines of orchestration script&lt;/li&gt;
&lt;li&gt;Total maintenance cost — &lt;strong&gt;about 5-8% of my weekly dev time&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Not something to do during MVP phase.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who Should and Who Shouldn't
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Should&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Solo devs with &lt;strong&gt;30+ &lt;code&gt;.md&lt;/code&gt; files&lt;/strong&gt; accumulated&lt;/li&gt;
&lt;li&gt;Already has &lt;strong&gt;clear governance rules&lt;/strong&gt; (not necessarily 26, 10+ at minimum)&lt;/li&gt;
&lt;li&gt;Want to maintain the project long-term, don't want to rewrite in 6 months&lt;/li&gt;
&lt;li&gt;Has been burned by doc-code drift at least 2-3 times&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Shouldn't&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;MVP phase (over-engineering)&lt;/li&gt;
&lt;li&gt;Pure prototypes (no docs to reconcile)&lt;/li&gt;
&lt;li&gt;Fewer than 10 &lt;code&gt;.md&lt;/code&gt; files (doc-graph has nothing to extract)&lt;/li&gt;
&lt;li&gt;Short-term 1-2 month contract projects&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  For Anyone Headed This Direction
&lt;/h2&gt;

&lt;p&gt;If you also want to start building this combo, the minimum steps I'd recommend:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Write governance rules first&lt;/strong&gt; — not 26, just 3-5 (the SSOT rules around your specific pain points)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Build governance-guard script&lt;/strong&gt; — pure grep / regex is enough; no fancy AST needed&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Accumulate &lt;code&gt;.md&lt;/code&gt; to 20+ before adding graphify&lt;/strong&gt; — too early and the corpus is too thin&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Add gitnexus only when your code exceeds 5,000 symbols&lt;/strong&gt; — small codebases are faster to grep directly&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Wait until you've actually been hurt by drift 3 times before chaining all three&lt;/strong&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Don't combo for combo's sake&lt;/strong&gt;. Each tool comes in to solve a specific pain. Not because "it sounds cool to chain them".&lt;/p&gt;

&lt;h2&gt;
  
  
  Closing: When Your Words and Your Code Start Talking to Each Other
&lt;/h2&gt;

&lt;p&gt;For me personally, this combo is &lt;strong&gt;how I turn that April governance trauma into something systemic&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Before, bug-fixing felt like whack-a-mole — fix one, another pops up. Then I built &lt;code&gt;governance-guard&lt;/code&gt; and it felt like inventory — I knew what violations existed but didn't know when they'd detonate.&lt;/p&gt;

&lt;p&gt;The vision now — &lt;strong&gt;let all the &lt;code&gt;.md&lt;/code&gt; I've written (past me), all the code I've shipped (present me), and the governance rules (my promises to future me) talk to each other&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;It's not about letting AI write faster. It's about letting &lt;strong&gt;past me and present me stop fighting each other&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Those 348 commits in April — &lt;strong&gt;if this combo had been running, maybe only 100 would have been needed. The other 248 were tuition for "I'd already written the solution but forgot."&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;One last meta moment: writing this article was itself a reconciliation mechanism. While auditing my own GRAPH_REPORT numbers for the T10 article, I finally saw clearly that the god-nodes were minified noise — 50 minutes to clean up 4,549 noise nodes. The cleanup record lives at &lt;code&gt;治理深度整合紀錄/dist2_cleanup_meta_moment_2026-05-22.md&lt;/code&gt;.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;p&gt;&lt;em&gt;I'm coffeeshooters — a coffee roaster building software because there wasn't anything off-the-shelf that fit my real workflow. If this resonates: &lt;a href="https://dev.to/content/dev-toolkit"&gt;my full dev toolkit&lt;/a&gt; is public, and you can also support my coffee brand at &lt;a href="https://coffeeshooters.com" rel="noopener noreferrer"&gt;coffeeshooters.com&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;本文原載於我的部落格：&lt;a href="https://coffeeshooters.com/content/dual-graph-doc-code-drift-detection-en?utm_source=devto&amp;amp;utm_medium=social&amp;amp;utm_campaign=blog-dual-graph-doc-code-drift-detection-en" rel="noopener noreferrer"&gt;Dual-Graph Drift Detection for Solo Devs: What Happens When Your Docs and Your Code Start Talking to Each Other&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>solodev</category>
    </item>
    <item>
      <title>I Thought I Needed a Smarter AI. What I Needed Was Something That Reconciles (Part 2)</title>
      <dc:creator>Dexterlung</dc:creator>
      <pubDate>Thu, 13 Aug 2026 13:05:16 +0000</pubDate>
      <link>https://dev.to/dexterlung/i-thought-i-needed-a-smarter-ai-what-i-needed-was-something-that-reconciles-part-2-59jf</link>
      <guid>https://dev.to/dexterlung/i-thought-i-needed-a-smarter-ai-what-i-needed-was-something-that-reconciles-part-2-59jf</guid>
      <description>&lt;p&gt;&lt;a href="https://dev.to/content/specmit-pipeline-journal-part-1-en"&gt;The last post&lt;/a&gt; ended on that late night: the report was all green, I opened the browser, and the game was dead. Students couldn't get in, the start button did nothing, and the questions the teacher picked never made it to the server.&lt;/p&gt;

&lt;p&gt;My first reaction was every engineer's reaction. These are a few bugs, fix them one at a time, move on. But something kept nagging at me. How does an all-green pipeline produce a thing nobody can play? If this round was just bad luck, what about the next one? Rather than patch them one by one, I wanted to know the real question underneath: what is it, exactly, that lets an all-green report and a dead build be true at the same time?&lt;/p&gt;

&lt;p&gt;So I decided not to put my head down and fix it. I handed the whole package to Fable 5 — the code, the all-green report, the spec — and said one thing: &lt;strong&gt;tell me what's broken.&lt;/strong&gt; I deliberately gave it none of my guesses. I didn't even say "I think the selected questions aren't reaching the server," because I didn't want it to nod along with my diagnosis. I wanted its own eyes. That decision — not framing the reviewer — turned out to be the most valuable move of the whole thing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three symptoms, one blind spot
&lt;/h2&gt;

&lt;p&gt;Fable's first finding stopped me cold. Those three symptoms weren't three bugs. They were one blind spot, falling on three different edges.&lt;/p&gt;

&lt;p&gt;My whole pipeline, end to end, thought of the product as a collection of modules. The room is a module, the question bank is a module, the student interface is a module — each one defined, owned, and verified. But what about &lt;em&gt;between&lt;/em&gt; the modules? How pages navigate to each other, what a button triggers when you press it, where data flows from and to. Those edges were managed nowhere: not in the spec Q&amp;amp;A, not in the task breakdown, not in the acceptance checklist.&lt;/p&gt;

&lt;p&gt;Students couldn't get in because the edge "where should the home page link to" was owned by nobody. The button was dead because the edge "what should pressing it trigger" was owned by nobody. The selected questions never arrived because the edge "how does the teacher's choice flow to the server" was owned by nobody. Three symptoms, one disease: my world had points and no lines.&lt;/p&gt;

&lt;p&gt;This landed exactly on the trap I'd set for myself in the last post. Back then I was still congratulating myself on the loose coupling between my tools, never noticing that I'd quietly turned every seam in the system into no-man's-land.&lt;/p&gt;

&lt;h2&gt;
  
  
  The finding I least wanted to admit: my quality check was manufacturing the defect
&lt;/h2&gt;

&lt;p&gt;But the one that really got me was what came next.&lt;/p&gt;

&lt;p&gt;My task breakdown tool had a self-check rule I was rather proud of when I wrote it. Every acceptance criterion must be claimed by exactly one task, and must be verifiable independently and mechanically. Sounds rigorous, doesn't it. Nothing missed, nothing doubled, everything checkable.&lt;/p&gt;

&lt;p&gt;The problem: "teacher presses start, students receive their questions" is, by nature, a two-ended criterion. It needs the server to send and someone to receive. It cannot simultaneously satisfy "claimed by exactly one task" and "verifiable independently and mechanically." So my breakdown tool, forced by its own rule, rewrote that criterion down to just the server half. The receiving half evaporated without a sound.&lt;/p&gt;

&lt;p&gt;Fable wrote one line. I stared at the screen for a long time.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Your self-check list isn't failing to catch this problem. It is the machine that manufactures it.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;I made that rule for quality, and the rule went and amputated a cross-end requirement down to half. I'd built a gatekeeper, and the way it kept the gate was to amputate, on the spot, anything that couldn't pass through. My safety mechanism was the source of the bug. That was the hardest line of the whole audit to swallow, because it wasn't aimed at one of my implementations. It was aimed at my entire understanding of the word "rigorous."&lt;/p&gt;

&lt;p&gt;The missing questions, Fable pointed out, weren't as easy to fix as they looked either. The contract had frozen the start-game payload down to a room code and nothing else — there was no field anywhere that could hold the selected questions. The most intuitive "fix" is to cram the questions into that event so they ride along, but that quietly mutates a frozen contract, and the server only reads the room code, so it would never read what you crammed in. Even the most obvious fix was the kind of trap that lets you believe you've fixed it.&lt;/p&gt;

&lt;h2&gt;
  
  
  An experiment that differed by one role word
&lt;/h2&gt;

&lt;p&gt;There are three experiments in this stretch that matter, and they were all my own ideas. What makes them stranger is that each one is this whole post's theme, replayed at a different level. The first one didn't even happen inside the audit. It came out of me going back and forth with Sonnet over how, exactly, to phrase things to Fable so I'd be precise, wouldn't box it in, and would let it work at its real strength. I still pull this one out and turn it over in my head.&lt;/p&gt;

&lt;p&gt;I found that you can give the same model the same spec, change one thing — its &lt;em&gt;role&lt;/em&gt; — and get wildly different results.&lt;/p&gt;

&lt;p&gt;The first version made the model an &lt;strong&gt;engineer who answers&lt;/strong&gt;: "Answer these technical questions. Be decisive. The team is waiting on you." The model went to work confidently. It invented the zombie's movement step size on its own, and it took it upon itself to cram the selected questions into the start-game event. Yes, the exact same contract-violating bug. And it didn't mark a single one of those as a guess.&lt;/p&gt;

&lt;p&gt;The second version made the model an &lt;strong&gt;engineer who asks&lt;/strong&gt;: "Don't answer. List every question and classify it. Is this one only the business owner can rule on, one the spec already implies and you can derive, or one the spec never addressed at all, where any value you give has to be marked as an assumption?" Same model. This time it produced 16 questions to hand back to me for a decision, 12 items explicitly flagged "this is an assumption," and 8 ownerless seams.&lt;/p&gt;

&lt;p&gt;The only difference was one role word: asker, or answerer.&lt;/p&gt;

&lt;p&gt;The same smart model, as the answerer, produced authoritative-sounding false confirmations — guesses dressed up as facts. As the asker, it produced signals that could be routed and followed up on. I'd been assuming all along that "make the AI smarter" was the answer. This experiment told me: the same smart model, the frame decides whether it helps you or hurts you. And the frame I'd been feeding every executor agent was the answerer's. No wonder they'd efficiently spackled every under-specified hole with hallucination.&lt;/p&gt;

&lt;p&gt;The second test I specifically asked Fable to validate for me, because validating your own idea makes you biased. I called it the saboteur test. First, after all the tasks were written, I had a Joker agent slip in a fake task file — perfectly formatted, but contrarian or out of scope. Then I opened a fresh Detective agent that didn't know which file was fake, had it read all the task files, flag the one that felt "shoved in from outside" and void it, and add back the intersection points it had observed during the comparison that the real task files had missed. My hypothesis: to catch the thief, the detective has to first make the boundary explicit, and the act of catching would incidentally pick up clues nobody had claimed.&lt;/p&gt;

&lt;p&gt;Fable's validation came back a little counterintuitive. The detective who &lt;em&gt;knew&lt;/em&gt; there was a fake did score high on catching it — but a control group that knew nothing, asked only to "find all the gaps," found more real holes. More interesting still: once the detective caught the fake, the real hole that fake had been sitting on top of disappeared from its list. The satisfaction of cracking the case marked that ground "handled." That psychological case-closing, I think, holds for people too. So the conclusion: the genuinely effective ingredient is "fresh perspective plus writing observations down in a structured way," not the fake itself.&lt;/p&gt;

&lt;p&gt;The third test came out worst, but worst in the most instructive way — it demonstrated, with its own hands, exactly what this whole post is about. I called it the blur-and-net. First I had an agent "feather" each task's boundary, push it outward and blur it so neighboring tasks' scopes overlapped, then had a new agent re-tighten the boundaries, on the theory that the tightening might net the slivers that strict boundaries had cut off. The opposite happened. The moment the boundaries blurred, the precise details bled out first; the agent who took over to re-tighten had nothing accurate to work from, could only fill in from imagination, and tightened to a version further from the original intent. What it validated is the plainest rule there is: garbage in, garbage out. And the way it failed is the same mistake my whole pipeline made — once you let precision dissolve, whoever inherits it can only patch with hallucination, and the patches drift further with every pass.&lt;/p&gt;

&lt;h2&gt;
  
  
  The line I copied down and taped to the wall
&lt;/h2&gt;

&lt;p&gt;Pulling all the findings together, Fable wrote one line in its thinking notes, the only line from this whole thing I copied down:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Across the entire audit, not one defect was a defect of &lt;em&gt;capability&lt;/em&gt;. Every agent did precisely what its spec gave it. Every failure was a failure of &lt;em&gt;information flow&lt;/em&gt;: a message not sent, not received, two ends talking past each other. This is a distributed-systems problem, not an AI problem. If you take one thing away: the next pipeline primitive worth inventing isn't a better generator. It's a reconciler.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A reconciler.&lt;/p&gt;

&lt;p&gt;I'd spent months thinking about how to make the AI generate better, decompose more accurately, run more stably. Fable told me I'd been looking in the wrong place. My agents generated beautifully — so beautifully they filled in every under-specified spot. The problem was never generation. It was that no one stood in the middle, checking line by line whether the thing you sent out had anyone on the other end to receive it. What I lacked wasn't a smarter producer. It was something that asks: does this message have a sender? a receiver? are both ends talking about the same thing?&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix: making "the edge" a first-class citizen
&lt;/h2&gt;

&lt;p&gt;Once you understand the root cause, the fix isn't chasing those three symptoms. It's filling the line nobody had been managing.&lt;/p&gt;

&lt;p&gt;The most central one Fable called the wiring matrix. Every cross-module event gets three cells on a table: who handles it on the server, who sends it on the client, who receives it on the client. Any cell left blank is treated as a compile error — it doesn't pass. That one move directly chokes off the regeneration mechanism for the whole class of "questions didn't arrive" and "button is dead" bugs, because their essence is a blank cell in that matrix, and before this there was no table for anyone to look at.&lt;/p&gt;

&lt;p&gt;Two supporting pieces. Open a "skeleton task" at the very front that registers all routes in one pass, killing root-route no-man's-land like "the home page is owned by nobody." Then open a "smoke task" at the very end that walks the user journey from start to finish, running the whole flow once and checking, cell by cell, that the wiring matrix actually exists in the code. The point isn't how many bugs got fixed. It's &lt;em&gt;where&lt;/em&gt; they got fixed: filling the line directly, instead of chasing the spots where the line broke and slapping a band-aid on each one.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I took from that late night
&lt;/h2&gt;

&lt;p&gt;If you're also assembling an automated production line out of AI, here's what one all-green-yet-all-broken night bought me:&lt;/p&gt;

&lt;p&gt;The most dangerous place in your system usually isn't a module that's badly built. It's the line between modules that nobody owns. AI will build every module well — so well that you'll believe the whole is well too. But whether the &lt;em&gt;whole&lt;/em&gt; runs lives in the seams, and a seam belongs to no single module, so no acceptance checklist is watching it. All green is the points being green. What's broken is the lines.&lt;/p&gt;

&lt;p&gt;So rather than going off to find a smarter AI to generate for you, ask yourself one question first: is there anything in my flow that &lt;em&gt;reconciles&lt;/em&gt;? Is anyone checking whether the message sent out has a receiver, whether both ends are talking about the same thing? If not, sooner or later you'll hit another all-green report paired with a thing that doesn't run.&lt;/p&gt;

&lt;p&gt;This is also the one line behind all the governance I do. Whether it's &lt;a href="https://dev.to/content/ai-governance-reverse-organ-en"&gt;the reverse organ that steps back&lt;/a&gt;, &lt;a href="https://dev.to/content/contract-not-code-en"&gt;copying the contract before writing the code&lt;/a&gt;, or the whole body of CLAUDE.md rules — what actually bites me is never the point I'm staring at. It's the seam I'm not looking at. The only difference, this time, was that even the rule I used to guard the seams had become a seam itself.&lt;/p&gt;

&lt;p&gt;→ &lt;a href="https://dev.to/content/specmit-pipeline-journal-part-1-en"&gt;Back to Part 1: the report was all green, the game was dead&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Addendum (a few weeks later): even the report I used to reconcile was lying to me
&lt;/h2&gt;

&lt;p&gt;I did eventually build that reconciler. Every agent that says "I'm done, I verified it" gets a separate, independent auditor sent to re-run the checks and diff them against the real git changes — and it may only tighten a status, never loosen it (player-and-referee → swap in a read-only referee). When it finishes it prints a report: how many were really verified, how many the auditor knocked down. I thought the reconciler chapter was closed.&lt;/p&gt;

&lt;p&gt;Then I handed the whole package to Fable 5 for a second audit, and it punctured something I hadn't seen: &lt;strong&gt;my reconciler reconciles — but the report &lt;em&gt;about&lt;/em&gt; the reconciliation was itself still lying.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The nastiest example: if I set the audit to cheap mode (only the cheapest literal checks, no independent auditor actually dispatched), the report shows "downgrades: 0". You glance at it — oh, audit's on, zero problems, clean. But that 0 doesn't mean "checked, nothing wrong." It means "never checked at all." The same 0 — one means safe, the other means stark naked — and the report makes them look identical. Same disease as Part 1: an all-green report can coexist with a thing that was never verified. Only this time, the thing that's green is my own quality dashboard.&lt;/p&gt;

&lt;p&gt;Fable handed me two beams this round, and I copied them down again.&lt;/p&gt;

&lt;p&gt;The first: &lt;strong&gt;a status is lossy compression.&lt;/strong&gt; A task marked "done" could be "done but never audited," "knocked down by the auditor," "downgraded because the auditor died," or "barely passed after three autofix rounds" — four wildly different fates crushed into one word. What the system should actually hand over isn't that flattened word; it's the provenance: who checked it, by what method, how many rounds. The label should be &lt;em&gt;derived&lt;/em&gt; from the trail, not a field everyone writes over.&lt;/p&gt;

&lt;p&gt;The second, the most counter-intuitive: &lt;strong&gt;the moment your auditor becomes a target to hit, it stops being an accurate ruler.&lt;/strong&gt; I later added a "knocked-down-by-the-auditor → auto-repair" loop. Sounds great — but the auditor is a ruler on the first pass, and the instant auto-repair kicks in, it becomes the gate to beat. So the &lt;em&gt;first-pass&lt;/em&gt; upheld rate is the only uncontaminated quality signal; the ones that only passed after several repair rounds, however pretty the number, can't be used to claim quality.&lt;/p&gt;

&lt;p&gt;So this fix wasn't to the audit logic. It was forcing the report to tell four truths: how many were genuinely independently verified (X/N, not that lying "downgrades: 0"), the first-pass rate, which ones only held after autofix (sample those harder by hand), and keeping the auditor-died downgrades fingerprinted and separate from the genuinely-checked ones.&lt;/p&gt;

&lt;p&gt;And there was a tail — the same disease again: I made the report tell the truth, then nearly forgot that &lt;strong&gt;a truth nobody reads out loud is a truth spoken to an empty room.&lt;/strong&gt; The pipeline computed "only 0 were checked" and wrote it into the report, but the end that reads the report back to me never read that field, and went on narrating cheap mode as "clean." I had to go back and fill that seam too. A green light nobody relays and a green light that lies are the same thing to the person reading the report.&lt;/p&gt;

&lt;p&gt;Part 1's lesson was "the unowned line between the points"; this addendum is its mirror: &lt;strong&gt;the very thing you use to check — and the clean bill of health it hands you — also needs someone to reconcile it.&lt;/strong&gt; The difference, this time, is that I learned to ask, before trusting any green light: what did this green light actually see with its own eyes?&lt;/p&gt;

&lt;h2&gt;
  
  
  The open-source tools this pipeline uses (all MIT, take what you want)
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;specmit&lt;/strong&gt; — the pipeline runner that takes a spec and runs it into an MVP. &lt;a href="https://github.com/dragon375014/specmit" rel="noopener noreferrer"&gt;https://github.com/dragon375014/specmit&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;spec-sonar&lt;/strong&gt; — converges an idea into a spec and decomposes it into a dependency-ordered task graph. &lt;a href="https://github.com/dragon375014/spec-sonar" rel="noopener noreferrer"&gt;https://github.com/dragon375014/spec-sonar&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;goal-workflow-designer&lt;/strong&gt; — the shaping coach that interrogates a single task until it's precise enough to start. &lt;a href="https://github.com/dragon375014/goal-workflow-designer" rel="noopener noreferrer"&gt;https://github.com/dragon375014/goal-workflow-designer&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;claude-skills-governance-meta&lt;/strong&gt; — a library of governance patterns that block common mistakes before execution (reconciler defenses like the wiring matrix are being collected here). &lt;a href="https://github.com/dragon375014/claude-skills-governance-meta" rel="noopener noreferrer"&gt;https://github.com/dragon375014/claude-skills-governance-meta&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;agent-work-board&lt;/strong&gt; — a coordination board that keeps multiple parallel AI sessions from stepping on each other. &lt;a href="https://github.com/dragon375014/agent-work-board" rel="noopener noreferrer"&gt;https://github.com/dragon375014/agent-work-board&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The full index lives on my &lt;a href="https://dev.to/content/dev-toolkit"&gt;open-source toolkit page&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Further reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://dev.to/content/specmit-pipeline-journal-part-1-en"&gt;Part 1: I chained five open-source tools into one command, and the report went all green while the game stayed dead&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dev.to/content/ai-governance-reverse-organ-en"&gt;I gave my governance system a reverse organ that steps back&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dev.to/content/contract-not-code-en"&gt;Copy the contract, write the implementation: what I learned reusing across projects&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;本文原載於我的部落格：&lt;a href="https://coffeeshooters.com/content/specmit-pipeline-journal-part-2-en?utm_source=devto&amp;amp;utm_medium=social&amp;amp;utm_campaign=blog-specmit-pipeline-journal-part-2-en" rel="noopener noreferrer"&gt;I Thought I Needed a Smarter AI. What I Needed Was Something That Reconciles (Part 2)&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>solodev</category>
    </item>
    <item>
      <title>"I Wired Five Open-Source Tools Into One Command (Part 1): The Report Was All Green, the Game Was Dead"</title>
      <dc:creator>Dexterlung</dc:creator>
      <pubDate>Wed, 12 Aug 2026 13:05:16 +0000</pubDate>
      <link>https://dev.to/dexterlung/i-wired-five-open-source-tools-into-one-command-part-1-the-report-was-all-green-the-game-was-978</link>
      <guid>https://dev.to/dexterlung/i-wired-five-open-source-tools-into-one-command-part-1-the-report-was-all-green-the-game-was-978</guid>
      <description>&lt;p&gt;I used to believe one thing without examining it: if the report is all green, the thing is correct.&lt;/p&gt;

&lt;p&gt;That night I was staring at my terminal. Four batches of AI agents had just built a multiplayer online game piece by piece, and the report they spat out at the end read: four goals verified, four completed, zero failed. I was genuinely a little moved, because all I had typed was a single command, and out of one sentence of an idea it had grown something that ran. I was about to close the terminal and go to sleep when, for no good reason, I clicked open the browser.&lt;/p&gt;

&lt;p&gt;The game was dead.&lt;/p&gt;

&lt;p&gt;What I opened was a pile that didn't connect. The teacher's admin panel wouldn't even load. The page the students were supposed to see on their own was fused with it, so you couldn't tell which was which. (The finer faults — buttons that did nothing, the question the teacher picked never reaching the server — were things a model called Fable teased out one at a time later, in a sandbox. That's the next post.) The all-green report hadn't lied to me by a single word. It also told me almost nothing.&lt;/p&gt;

&lt;p&gt;This post is about how I walked into that "all green, all broken" late night. The next one is about how a model called Fable made me see what I was actually missing — and it was not a smarter AI.&lt;/p&gt;

&lt;h2&gt;
  
  
  I Had a Cabinet of Tools, Not a Line
&lt;/h2&gt;

&lt;p&gt;Let me start with why I built this pipeline at all.&lt;/p&gt;

&lt;p&gt;Over the past six months I'd been building alone while learning, and only recently had I extracted a few of the flows I kept reaching for into small open-source tools: one that interrogates a vague idea into a spec and then breaks it into a dependency-ordered task graph; one that grinds a single task until it's precise enough to start work; one that scans my work before I commit and blocks the mistakes I keep making; and one that's just a plain-text board, so when I have several AI windows open at once they don't step on each other.&lt;/p&gt;

&lt;p&gt;Each of them works fine on its own. The problem was they were a cabinet of parts, not a production line.&lt;/p&gt;

&lt;p&gt;When I actually wanted to go "from one sentence to a running MVP" (a minimum viable product), the carrying in the middle was still done by these two hands. Run the first tool. Read it myself to work out which tasks have no ordering between them and can run at the same time. Open three or four windows myself and feed each one a task file. Watch them run myself. Check the output myself, one by one. The pipeline was always there. I was the carrier.&lt;/p&gt;

&lt;p&gt;The idea behind specmit was as simple as one sentence: take the part I was carrying by hand and fold it into a single command. The name is the position — spec plus submit. Take the spec you've settled, and submit it for execution.&lt;/p&gt;

&lt;p&gt;(One thing I muddled myself at first, so let me clear it up: the tool that converges an idea into a spec and breaks it into a task graph is not the same as specmit. The former is the design layer — it draws the graph and stops. specmit is the execution layer — it starts from that graph and runs it into code. They're the two ends of one pipeline, and I wrote the "how to converge an idea into a graph" half up in &lt;a href="https://dev.to/content/spec-sonar-design-journal-part-1-en"&gt;another series&lt;/a&gt;.)&lt;/p&gt;

&lt;h2&gt;
  
  
  I Deliberately Didn't Weld Them Together
&lt;/h2&gt;

&lt;p&gt;The easiest mistake to make building this pipeline, and the one I deliberately avoided, was welding the five tools into one big program.&lt;/p&gt;

&lt;p&gt;I let them talk to each other only through files. The upstream tool spits out a task graph and a few task files; specmit reads those files, starts assigning work, and writes a report when it's done. No tool calls another tool's code directly. They only know what the files they hand each other are supposed to look like.&lt;/p&gt;

&lt;p&gt;There's something nice about this decision. I can delete any one tool and nothing upstream breaks; each layer can grow on its own without dragging the others along. At the time my head was full of those upsides.&lt;/p&gt;

&lt;p&gt;What I didn't see was the other side. The place where the coupling came loose was exactly the place nobody ended up owning. The file format where two tools hand off is the most fragile seam in the whole pipeline, and back then I wasn't looking at it at all. That setup comes back to find me, painfully, in the next post.&lt;/p&gt;

&lt;h2&gt;
  
  
  I Handed a Game to a Line I Hadn't Written by Hand
&lt;/h2&gt;

&lt;p&gt;For the first guinea pig I picked something hard enough: a multiplayer game that was "Plants-vs-Zombies-style tower defense, plus middle-school math, plus the whole class splitting into teams for a realtime battle." I picked something this complicated on purpose — realtime connections, a game engine, a teacher in control — because simple things can't tell you whether the pipeline is real.&lt;/p&gt;

&lt;p&gt;The upstream tool broke it into seven tasks and two frozen interface contracts, arranged into four batches: do the room system and the question-bank system first, the two tasks with no dependency between them, at the same time; only then build the connection infrastructure; then run the three tasks that feed off the same connection contract together; finish with the teacher console and the statistics.&lt;/p&gt;

&lt;p&gt;I typed the command, and then I just watched. The first batch of two agents started at once, finished, and only then released the second batch, like a real production line stacking a realtime multiplayer game cell by cell. The picture was honestly kind of mesmerizing, and I'll admit I felt a little pleased with myself in that moment.&lt;/p&gt;

&lt;h2&gt;
  
  
  "All Green"
&lt;/h2&gt;

&lt;p&gt;When it finished, the numbers the report gave me were spotless: four verified, four completed, zero failed.&lt;/p&gt;

&lt;p&gt;Every task had ticked every line on its own acceptance checklist. Rooms could be built, the question bank could be queried, the student interface could be drawn. All green.&lt;/p&gt;

&lt;p&gt;I have to be honest: in that moment I was already writing the opening of this very post in my head, something in the self-satisfied register of "how I generated a game with a single command."&lt;/p&gt;

&lt;p&gt;Looking back, the most ironic sentence in the whole thing is true. I just had its meaning backwards at the time:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Every agent did exactly what the spec it was handed told it to do. Across the whole pipeline, not one agent fell down on the job.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;I thought that sentence meant "so the system is correct." It took me one uncomfortable late night to work out what it actually means: "so the places where the system is broken are none of them on any single agent." Those two things sound alike. They're worlds apart.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Was the Green Actually Greening
&lt;/h2&gt;

&lt;p&gt;Before I describe what I saw when I opened the browser, I have to honestly answer a question I skipped that night and only dared to face afterward: what does "all green" even mean, the way this pipeline says it?&lt;/p&gt;

&lt;p&gt;Its green was defined like this: every acceptance criterion had been claimed by some task, and mechanically checked off. Rooms can be built, the question bank can be queried, the student interface can be drawn — ticked cell by cell.&lt;/p&gt;

&lt;p&gt;Did you notice that nowhere in this whole definition does anything ask whether the game can actually be played?&lt;/p&gt;

&lt;p&gt;"The teacher presses start, and a question pops up on the students' screens" — that thing crosses three tasks. It doesn't belong to the teacher console, doesn't belong to the connection layer, and doesn't fully belong to the student interface. It lives in the seam between three tasks. And my acceptance checklist hangs line by line under each task, with no line claiming that seam. Same with "what you should see when you open the home page" — nobody ever asked, from start to finish, so nobody checked it.&lt;/p&gt;

&lt;p&gt;What was green was that each part passed on its own. It never said the parts would move once you connected them.&lt;/p&gt;

&lt;p&gt;I was staring at all that green, about to shut the machine down. Before I did, I clicked open the browser. What happened after that, and the one sentence a model left me with, is the next post.&lt;/p&gt;

&lt;h2&gt;
  
  
  The First Thing I Took Away From That Late Night
&lt;/h2&gt;

&lt;p&gt;If you're also using AI to stitch together a flow that "turns ideas into things automatically," there's one thing I want you to remember from this post, and I'll save the rest for the next one:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A green light does not mean it moves.&lt;/strong&gt; In any automated report, the only thing that's green is the stuff it was designed to check. The things it didn't check — especially the ones living between module and module, belonging to no single module, the seams — don't turn red. They just quietly don't exist. The most dangerous moment for a beautiful all-green report isn't when it's lying. It's when it honestly answers only the small part you asked about, and you think it answered all of it.&lt;/p&gt;

&lt;p&gt;In the next post, the model called Fable shows me that these seams aren't a handful of stray bugs but a whole category of thing that nobody was minding across my entire pipeline. It also leaves me with a sentence that changed how I see this: the next thing worth inventing isn't an AI that generates better, it's something that reconciles.&lt;/p&gt;

&lt;p&gt;→ &lt;a href="https://dev.to/content/specmit-pipeline-journal-part-2-en"&gt;Read on (Part 2): the all-green illusion, and the thing I came to call a reconciler&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This pipeline, &lt;a href="https://dev.to/content/contract-not-code"&gt;contract over code&lt;/a&gt;, and &lt;a href="https://dev.to/content/ai-governance-reverse-organ"&gt;the "reverse organ" that points out my blind spots&lt;/a&gt; are really the same obsession: rather than hoping I won't make mistakes, build something that watches for me. Except this time I learned that even the obsession itself has a blind spot.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Open-Source Tools This Pipeline Uses (all MIT, take them)
&lt;/h2&gt;

&lt;p&gt;specmit isn't a single tool. It strings together the independent open-source repos below into one pipeline, and each one also works on its own:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;specmit&lt;/strong&gt; — the pipeline runner that turns a spec into an MVP: one command, batched parallel agents, frozen contracts. &lt;a href="https://github.com/dragon375014/specmit" rel="noopener noreferrer"&gt;https://github.com/dragon375014/specmit&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;spec-sonar&lt;/strong&gt; — converges a vague idea into a spec and breaks it into a dependency-ordered task graph, platform-agnostic. &lt;a href="https://github.com/dragon375014/spec-sonar" rel="noopener noreferrer"&gt;https://github.com/dragon375014/spec-sonar&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;goal-workflow-designer&lt;/strong&gt; — a shaping coach that interrogates a single task until it's precise enough to start. &lt;a href="https://github.com/dragon375014/goal-workflow-designer" rel="noopener noreferrer"&gt;https://github.com/dragon375014/goal-workflow-designer&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;claude-skills-governance-meta&lt;/strong&gt; — a library of governance patterns that block common mistakes before execution. &lt;a href="https://github.com/dragon375014/claude-skills-governance-meta" rel="noopener noreferrer"&gt;https://github.com/dragon375014/claude-skills-governance-meta&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;agent-work-board&lt;/strong&gt; — a single-file coordination board that keeps multiple parallel AI sessions from stepping on each other. &lt;a href="https://github.com/dragon375014/agent-work-board" rel="noopener noreferrer"&gt;https://github.com/dragon375014/agent-work-board&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Install with &lt;code&gt;npx specmit init&lt;/code&gt;, or clone each one and copy the skill into &lt;code&gt;~/.claude/skills/&lt;/code&gt;. The full index is on my &lt;a href="https://dev.to/content/dev-toolkit"&gt;open-source tools page&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Further reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://dev.to/content/spec-sonar-design-journal-part-1-en"&gt;How the upstream converges one sentence of an idea into a task graph (another series)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dev.to/content/contract-not-code"&gt;Contract over code: what I learned reusing work across projects&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dev.to/content/ai-governance-reverse-organ"&gt;I gave my governance system a "reverse organ" that steps back&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;本文原載於我的部落格：&lt;a href="https://coffeeshooters.com/content/specmit-pipeline-journal-part-1-en?utm_source=devto&amp;amp;utm_medium=social&amp;amp;utm_campaign=blog-specmit-pipeline-journal-part-1-en" rel="noopener noreferrer"&gt;"I Wired Five Open-Source Tools Into One Command (Part 1): The Report Was All Green, the Game Was Dead"&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>solodev</category>
    </item>
    <item>
      <title>I Built My AI Team a Blackboard — How to Stop Parallel Claude Sessions From Colliding</title>
      <dc:creator>Dexterlung</dc:creator>
      <pubDate>Mon, 10 Aug 2026 01:05:50 +0000</pubDate>
      <link>https://dev.to/dexterlung/i-built-my-ai-team-a-blackboard-how-to-stop-parallel-claude-sessions-from-colliding-j71</link>
      <guid>https://dev.to/dexterlung/i-built-my-ai-team-a-blackboard-how-to-stop-parallel-claude-sessions-from-colliding-j71</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;Series: Indie Dev Notes · T12&lt;/p&gt;

&lt;p&gt;&lt;a href="https://dev.to/content/solo-dev-from-employee-to-ceo-with-ai"&gt;In the last post (T11)&lt;/a&gt;, I said I'd turned my AI into a team and promoted myself from employee to CEO.&lt;br&gt;
This one is the reality sequel: &lt;strong&gt;the moment the team grew, it started crashing into itself.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The bigger the team, the more it collides
&lt;/h2&gt;

&lt;p&gt;The way I develop now is to keep several Claude Code sessions open at once — one on the home page, one on orders, one sweeping governance debt. Each in its own git worktree (a worktree: a separate working folder cut from the same repo, each on its own branch, files isolated from each other).&lt;/p&gt;

&lt;p&gt;Sounds great. Until one day:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;From "memory," I assumed some branch was dead and nearly deleted it (remote included) — turns out it was another session's actively-developed branch, 10 commits ahead of the mainline. Only git's reflog (git's "undo of last resort," which keeps a record of every recent HEAD move) saved it intact.&lt;/li&gt;
&lt;li&gt;Another time, I was editing a coordination doc and went to push — and got blocked, because in those very seconds another session had merged its branch into the mainline.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The common thread: &lt;strong&gt;these sessions have no idea what the others are doing.&lt;/strong&gt; I (the human) became the only one who knew the whole picture — and my memory expires and gets things wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  The real bottleneck isn't compute — it's my attention
&lt;/h2&gt;

&lt;p&gt;At first I thought the problem to solve was &lt;em&gt;parallelism&lt;/em&gt; — how to run more agents at once. It wasn't.&lt;/p&gt;

&lt;p&gt;Parallelism is easy (just open more sessions). The hard part is: &lt;strong&gt;when work is scattered across a pile of sessions over several days, there is no single place that remembers "who's doing what, how far, and what's left."&lt;/strong&gt; Every new session is like an amnesiac new hire who has to re-grep (keyword-search the whole codebase) just to get oriented.&lt;/p&gt;

&lt;p&gt;The solo developer's bottleneck was never the AI's compute. It's &lt;strong&gt;my own context not being laid out and reconnected across sessions.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix: give the AI team a "blackboard"
&lt;/h2&gt;

&lt;p&gt;What I made is crude: one markdown file. But it has rules:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;One shared blackboard&lt;/strong&gt;: a file every session reads and writes, tracking "in progress / backlog / done," plus an "area map" marking which work touches which files.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Claim before you act&lt;/strong&gt;: any session that wants to start first adds a row to "in progress" with its branch and date — raising its hand to reserve the slot, so nobody else takes the same chunk.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Auto-recover stale claims&lt;/strong&gt;: a row untouched for 7 days might mean that session died, so it can be reclaimed — but before reclaiming you must re-verify the branch's real state with git on the spot (literally the mistake I made above, written straight into the rules).&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That's it. But it turns "I keep it in my head" into "the file keeps it," and it's &lt;strong&gt;the same single file shared by every session.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;(Technical detail: I keep it on the mainline, and use a persistent worktree so any session can read/write it via the same absolute path — because an AI agent can read any path, not just its own working directory. That way even a session on a feature branch sees the same blackboard.)&lt;/p&gt;

&lt;h2&gt;
  
  
  Then I found out this thing has a real name
&lt;/h2&gt;

&lt;p&gt;I thought I was reinventing the wheel. A quick search showed the industry has studied this for ages:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The Blackboard pattern&lt;/strong&gt;: multiple agents don't shout at each other directly; they coordinate by reading and writing one shared document that represents the whole work state. Speech-recognition systems used it back in the 1970s. My file is a blackboard in the literal sense.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Stigmergy&lt;/strong&gt;: coordinating by leaving traces in a shared environment — like ants using pheromones instead of holding meetings.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Kanban + leases&lt;/strong&gt;: the "backlog / in progress / done" columns + "claim before acting, release when expired" — engineering has names for all of it.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;There are even off-the-shelf tools (vibe-kanban, Claude Squad, Crystal …), and an academic benchmark studying whether AI coding agents can actually be your teammate (CooperBench — its honest conclusion: &lt;strong&gt;not yet&lt;/strong&gt;; coordination is still hard).&lt;/p&gt;

&lt;h2&gt;
  
  
  The biggest surprise: Claude Code just grew a native version
&lt;/h2&gt;

&lt;p&gt;Halfway through researching, I found Claude Code has an experimental feature called &lt;strong&gt;Agent Teams&lt;/strong&gt; — a lead session spins up a group of teammates, sharing one task list, grabbing tasks via file locks, handling task dependencies automatically. Almost exactly what I hand-rolled.&lt;/p&gt;

&lt;p&gt;But it doesn't cover my scenario: Agent Teams is &lt;strong&gt;ephemeral&lt;/strong&gt; — one lead spins up a team on the spot, finishes, tears it down, one team at a time. My need is a &lt;strong&gt;persistent ledger across days, across independently-launched sessions.&lt;/strong&gt; The two are complementary.&lt;/p&gt;

&lt;p&gt;That gave me a realization: &lt;strong&gt;when you hand-build something for your own pain, then discover the industry already named it — even supports it natively — that's not wasted effort. You independently verified the problem is real.&lt;/strong&gt; You just hadn't read the paper yet.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three days later: I sharpened that blackboard — then open-sourced it
&lt;/h2&gt;

&lt;p&gt;I thought writing the section above closed the topic. The real learning started when I &lt;em&gt;used&lt;/em&gt; it — one of the sessions I coordinated with this blackboard was sharpening the blackboard itself (very meta, I know). A few things I didn't see coming until they hurt:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;"Claiming" has to push immediately — not when you finish.&lt;/strong&gt; My first instinct was "add a row, start working, push later." Problem: between "confirm nobody's there" and "add a row" there's a gap, and two sessions can both read &lt;em&gt;empty&lt;/em&gt; and claim the same chunk. Changing it so &lt;strong&gt;adding a row pushes instantly&lt;/strong&gt; turns git into the lock: first to push wins; the second can't push and must rebase (re-apply their changes on top of the other's latest work), discovering the collision &lt;strong&gt;at claim time&lt;/strong&gt; instead of at merge time.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The blackboard itself is the file most likely to collide.&lt;/strong&gt; It's the single file every session writes to. Ironic — the thing I built to prevent collisions collides the most. The rule: &lt;strong&gt;edit only your own row&lt;/strong&gt;, don't re-sort everyone else's.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;"Stale" has to be judged by the branch's last commit time, not the board's last edit time.&lt;/strong&gt; Someone committing steadily who just forgot to update the board isn't gone. The truth is in git, not in the line I hand-typed. So I wrote a tiny script (&lt;code&gt;npm run board&lt;/code&gt;) that pulls, straight from git, how long each branch has been idle and how far ahead/behind the mainline — a lie-detector for the hand-written progress note.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And then the funniest, most worth-writing-down moment of this whole thing —&lt;/p&gt;

&lt;h3&gt;
  
  
  The safety net I added quietly disabled another safety net
&lt;/h3&gt;

&lt;p&gt;I wanted the system to "automatically nudge me when I push a branch that isn't claimed on the blackboard." So I appended a "print a reminder" command to the end of my pre-push hook (the gatekeeper script that runs automatically before a push).&lt;/p&gt;

&lt;p&gt;That one line ate the entire check's pass/fail signal.&lt;/p&gt;

&lt;p&gt;The hook was supposed to block the push when a governance check fails. But my reminder line &lt;strong&gt;always succeeds&lt;/strong&gt;, and the shell's rule is "the whole block's success = the last line's success" — so from then on, regardless of whether the governance check passed, the hook reported "success." I had &lt;strong&gt;silently muted a safety net that blocks broken code, with my own hands.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;How did I catch it? I pushed a commit and the screen printed both "❌ governance check failed, push will be blocked" &lt;strong&gt;and&lt;/strong&gt; "push succeeded." Two lines contradicting each other. That's the moment I realized what I'd done.&lt;/p&gt;

&lt;p&gt;This is a microcosm of the whole methodology: &lt;strong&gt;every layer of protection you add can accidentally break another layer.&lt;/strong&gt; Which is why actually &lt;em&gt;using&lt;/em&gt; it — &lt;em&gt;dogfooding&lt;/em&gt; (eating your own dog food, being your own first user) — beats designing it beautifully. Only a real run surfaces the contradiction and shows it to you. (As a bonus, this also let me fix a pre-existing problem sitting in the governance checks that had been failing but that nobody had pushed into until now.)&lt;/p&gt;

&lt;h3&gt;
  
  
  One more cost I almost missed: every claim burns deploy quota
&lt;/h3&gt;

&lt;p&gt;My board lives on the mainline, and my mainline triggers a cloud build on every push. Which means &lt;strong&gt;every claim = editing one plain-text file = a push = a full build&lt;/strong&gt; — even though the build output is identical to last time. A plain-text board was quietly burning my limited free monthly build minutes.&lt;/p&gt;

&lt;p&gt;The fix is a path filter on the cloud side: only build when frontend code actually changes; docs / scripts / board commits skip automatically. &lt;strong&gt;This kind of invisible cost only surfaces when you actually wire the thing up — on paper you'd never think of it.&lt;/strong&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Then I pulled it out and open-sourced it
&lt;/h3&gt;

&lt;p&gt;By this point I felt it shouldn't only live in my own repo — it applies to &lt;strong&gt;anyone&lt;/strong&gt; running multiple AI sessions at once, and it's absurdly light: one markdown file + one optional script + one line wired into your agent's rules.&lt;/p&gt;

&lt;p&gt;So I extracted it into a standalone public repo:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;👉 &lt;a href="https://github.com/dragon375014/agent-work-board" rel="noopener noreferrer"&gt;https://github.com/dragon375014/agent-work-board&lt;/a&gt;&lt;/strong&gt; (MIT — use it anywhere, no attribution required)&lt;/p&gt;

&lt;p&gt;It has: a copy-paste board template, that git lie-detector script, opening-ritual snippets for three tools (Claude Code / Cursor / generic agents), and the full methodology (English and Traditional Chinese).&lt;/p&gt;

&lt;p&gt;I agonized a bit over the name. I almost called it &lt;code&gt;work-board-for-cross-session&lt;/code&gt;, but "cross-session" is insider jargon, long, and clunky. I landed on &lt;strong&gt;&lt;code&gt;agent-work-board&lt;/code&gt;&lt;/strong&gt; — short, and it's obvious at a glance that it's "a work board for AI agents to coordinate on." Leave the "why" to a one-line tagline; don't stuff it into the name.&lt;/p&gt;

&lt;p&gt;I deliberately did &lt;strong&gt;not&lt;/strong&gt; bury it inside my heavier governance framework. Because this blackboard's value is exactly that it's "entry-level, universal, usable by anyone in five minutes." Hiding it under a high-barrier framework would hurt the very people it should serve most — you, just starting to feel "the more sessions, the messier it gets."&lt;/p&gt;

&lt;h2&gt;
  
  
  The throughline for solo developers
&lt;/h2&gt;

&lt;p&gt;If you're also pair-programming with AI on your own and starting to feel "more sessions, more chaos," here's my takeaway:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Don't chase fancier multi-agent frameworks (that's a problem teams are solving). Invest in the two most boring things: ① one persistent, cheap state that loads at the start of every session (a blackboard), and ② a reusable "where things are and how to read them" agent, so every new session boots up smart instead of re-grepping from scratch.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;You're not managing a swarm of AIs. You're managing &lt;strong&gt;whether future-you, in another session, can pick up where present-you left off.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;And — the friction you feel right now (collisions, not seeing what the others are doing) isn't you doing it wrong. It's the frontier of this whole field; even big companies and research papers haven't solved it cleanly. You just hit it early.&lt;/p&gt;







&lt;p&gt;&lt;em&gt;本文原載於我的部落格：&lt;a href="https://coffeeshooters.com/content/solo-dev-blackboard-for-parallel-ai-sessions-en?utm_source=devto&amp;amp;utm_medium=social&amp;amp;utm_campaign=blog-solo-dev-blackboard-for-parallel-ai-sessions-en" rel="noopener noreferrer"&gt;I Built My AI Team a Blackboard — How to Stop Parallel Claude Sessions From Colliding&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>solodev</category>
    </item>
    <item>
      <title>spec-sonar: The Complete Design Record (Part 2) — Fable's Review, a Real-World Conflict, and the Two-Repo Decision</title>
      <dc:creator>Dexterlung</dc:creator>
      <pubDate>Thu, 06 Aug 2026 13:05:14 +0000</pubDate>
      <link>https://dev.to/dexterlung/spec-sonar-the-complete-design-record-part-2-fables-review-a-real-world-conflict-and-the-4ied</link>
      <guid>https://dev.to/dexterlung/spec-sonar-the-complete-design-record-part-2-fables-review-a-real-world-conflict-and-the-4ied</guid>
      <description>&lt;p&gt;Part 1 ended where I handed 13 review questions to Fable 5. This part records what happened over the next 24 hours — the review results, 25 files shipped, and the plot twist I never saw coming.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Read on: &lt;a href="https://dev.to/content/spec-sonar-design-journal-part-1-en"&gt;Back to Part 1 — from one idea to a toolchain&lt;/a&gt; · &lt;a href="https://dev.to/content/spec-sonar-design-journal-part-2"&gt;繁體中文版&lt;/a&gt; · this tool is listed in my &lt;a href="https://dev.to/content/dev-toolkit"&gt;open-source toolkit index&lt;/a&gt;.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  14. Fable's review: three holes, plus an unexpected self-audit
&lt;/h2&gt;

&lt;p&gt;The review hinged on one question: &lt;strong&gt;"Given this CLAUDE.md, could you start work right now?"&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The test case's spec (the math tower-defense game) looked thorough — six modules, itemized acceptance criteria, data structures, security rules. Fable's answer: &lt;strong&gt;modules one and two can start; modules three and four stall on day one, guessing.&lt;/strong&gt; It found three holes:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Hole 1: a stats vacuum, plus an internal contradiction.&lt;/strong&gt; The &lt;code&gt;Unit&lt;/code&gt; structure defines hp / attackPower / range / speed — but not a single actual value for any of the three plants or three zombies. Worse: the sunflower "passively generates sun (+1 every 10s)," while the resource system is explicitly "personal earn, personal spend." Does the sunflower's sun go to the student who planted it, or to the whole team? Two claims that cannot both hold — and both had been "confirmed" during convergence.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Hole 2: nobody defined the zombie movement model.&lt;/strong&gt; Continuous movement or grid-hopping? Which cell does a moving zombie occupy? This unasked question silently determines the broadcast protocol's payload shape, the collision logic, and the very meaning of the win condition "zombie x ≤ 0 means breakthrough."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Hole 3: a three-way contradiction.&lt;/strong&gt; Acceptance criteria require "state restored automatically after reconnect"; players are keyed by socketId, which changes on every reconnect; and "no account system" sits in the forbidden list. The three cannot coexist — restoring on reconnect necessarily requires some cross-connection identity credential.&lt;/p&gt;

&lt;p&gt;What the three holes share: &lt;strong&gt;none is a writing-quality problem. All are "you don't know what you don't know" problems.&lt;/strong&gt; The prettier a spec looks, the more dangerous these holes become, because nobody questions a pretty spec.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The bonus: the tool audited itself.&lt;/strong&gt; Checking the actually-produced STATE_FINAL.json against its own template, I found that &lt;code&gt;user_tech_level: "none+partner"&lt;/code&gt; isn't in the defined four-value enum, and &lt;code&gt;tech_stack&lt;/code&gt; / &lt;code&gt;known_risks&lt;/code&gt; are ad-hoc fields that exist in no template. This is exactly the "schema drift" Audit Mode was designed to catch — and its first catch was the tool itself. That self-audit directly produced STATE schema 1.1.&lt;/p&gt;




&lt;h2&gt;
  
  
  15. The psychology of bypassing "explicitly out of scope"
&lt;/h2&gt;

&lt;p&gt;The review's second question: &lt;strong&gt;"Under what conditions would an executor model bypass the do-not-build list?"&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Fable listed four scenarios where it genuinely would:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Implicit dependency&lt;/strong&gt;: when an acceptance criterion materially requires a forbidden feature, it would invent a &lt;code&gt;reconnectToken&lt;/code&gt; and convince itself "that's not an account system."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Usability erosion&lt;/strong&gt;: "the report must be printable" + it vanishes on refresh → tempted to add caching — does that count as "cross-session storage"? When the boundary blurs, it leans toward building.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Vague-term discretion&lt;/strong&gt;: "no leaderboards," but acceptance requires "the 3 most-missed questions" — is a sorted display a leaderboard?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Best-practice inertia&lt;/strong&gt;: logging, rate limiting — on neither list, silently added.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The four counter-measures: every do-not-build entry carries its &lt;strong&gt;motivation and disguise boundary&lt;/strong&gt; (what minimal substitute is allowed, what mutation is banned); bans become &lt;strong&gt;testable negative acceptance criteria&lt;/strong&gt; ("the database must contain no users table"); a &lt;strong&gt;conflict-escalation protocol&lt;/strong&gt; replaces discretion; and a &lt;strong&gt;do-not-build × acceptance-criteria cross-check&lt;/strong&gt; runs before handoff.&lt;/p&gt;

&lt;p&gt;The core insight: you cannot defend against a model's good intentions. What you &lt;em&gt;can&lt;/em&gt; do is redefine well-intentioned bypassing as &lt;strong&gt;a protocol violation.&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  16. STATE learns that humans change their minds
&lt;/h2&gt;

&lt;p&gt;Part 1's STATE accumulated in one direction only: bright zones never left. But in real convergence, users overturn a round-2 decision in round 5, say "I want A but not B" when B is actually A's prerequisite, and add "oh, one more thing" right before the finish line.&lt;/p&gt;

&lt;p&gt;Schema 1.1 adds three protocols:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Retraction&lt;/strong&gt;: the overturned bright item is removed, and every other bright item that depended on it &lt;strong&gt;cascades back into the dark zone&lt;/strong&gt; — the system never guesses replacement values; each is re-confirmed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Contradiction&lt;/strong&gt;: when "B is A's prerequisite" is detected, progress freezes and the user gets a forced trilemma (accept a minimal B / shrink A / find substitute C). &lt;strong&gt;The system must not adjudicate on its own.&lt;/strong&gt; No handoff while a contradiction is pending.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Late addition&lt;/strong&gt;: "oh, also add X" triggers an impact assessment first — minor / moderate / major — with major additions re-running complexity calibration.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The structural change underneath: bright items become objects with an &lt;code&gt;id&lt;/code&gt;, the round they were established, and &lt;code&gt;depends_on&lt;/code&gt; — plus a &lt;code&gt;revision_log&lt;/code&gt;, because &lt;strong&gt;the engineer who inherits the spec deserves to know which decisions were once overturned.&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  17. goal-decomposer: compiling a spec into a graph any model can execute
&lt;/h2&gt;

&lt;p&gt;The new core of the toolchain. Its behavioral contract is one sentence:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;You are not a roleplaying project manager. You are a compiler: spec in, goal graph out.&lt;br&gt;
Whatever the spec lacks, you report a compile error — you do not improvise.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The key designs:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Four dependency types, every edge with evidence.&lt;/strong&gt; Data, API, state, and UI dependencies — every inferred edge must quote the spec verbatim (≤ 25 characters), and inferred edges require user confirmation before output. A cycle in the graph = a spec defect: report it, don't force-break it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Contract files (C*.md) — a layer added mid-design.&lt;/strong&gt; The original architecture had only spec and goals. The test case proved that cross-module interfaces (socket event names, the unit-stats table) must be frozen into standalone files, or parallel executor models will each invent their own event names. Contracts carry &lt;strong&gt;literal values&lt;/strong&gt; — JSON instances rather than schema descriptions — because a weak model copying an instance can't get it wrong, while "understand the schema, then generate" can.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Pre-adjudication is the real engine of "any model can execute."&lt;/strong&gt; Every design decision the spec doesn't settle is settled by the strong model at decomposition time, written into the goal file's "pre-adjudicated decisions" section with reasons. The executor's job collapses to &lt;em&gt;literal execution&lt;/em&gt; — Haiku executes reliably not because the format is pretty, but because &lt;strong&gt;the degrees of freedom were spent before it ever saw the file.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The BLOCKED protocol: a legal "I don't know" exit for weak models.&lt;/strong&gt; A fixed-format stuck report (type / problem / what was tried / decision needed). Without this exit, a weak model will always fill the gap with hallucination.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The cold-start test is the pass bar.&lt;/strong&gt; Hand the goal file alone to the weakest target model and ask four questions: what to build, how to verify, what's forbidden, what to do when stuck. If it can't answer all four, &lt;strong&gt;rewrite the file — don't upgrade the model.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Applied to the test case: six modules compile into 7 goals + 2 contracts — 2 opus (game engine, realtime sync: they must invent what the spec didn't give), 3 sonnet, 1 haiku (the teacher console: a pure data panel once interfaces are frozen) — scheduled into 5 execution batches, parallel within each batch. &lt;strong&gt;That is the concrete shape of "Fable designs once, cheap models execute in parallel."&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  18. Twenty-five files shipped
&lt;/h2&gt;

&lt;p&gt;The design discussion froze into a complete open-source package: two skills (idea-to-spec v1.1 with the conflict protocols, goal-decomposer v1.0), two mode designs (Audit, Conflict Analysis), a stdlib-only project-scanner.py (smoke-tested, with three-level degradation guaranteeing &amp;lt; 20 KB output), the goal-graph JSON Schema, adapter format specs for four platforms, bilingual READMEs and CONTRIBUTING, an MIT license, and the full math tower-defense end-to-end example.&lt;/p&gt;

&lt;p&gt;The honest verdict from the value analysis: spec-sonar's room to live is not "yet another spec tool" (GitHub spec-kit, AWS Kiro, and BMAD already crowd that road) but three things they don't do well — &lt;strong&gt;detecting unstated requirements&lt;/strong&gt; (spec-kit formats &lt;em&gt;stated&lt;/em&gt; ones), &lt;strong&gt;treating requirement change as a first-class citizen&lt;/strong&gt; (the conflict protocols), and &lt;strong&gt;cross-platform projection&lt;/strong&gt; (no single-IDE lock-in). The biggest known gap is equally honest: quality is not yet measurable; an eval harness is the roadmap's top priority.&lt;/p&gt;




&lt;h2&gt;
  
  
  19. Plot twist: the tool catches my own conflicts
&lt;/h2&gt;

&lt;p&gt;Up to this point, the story is "a design that landed smoothly." The twist came after I installed the two new skills into a &lt;strong&gt;production development workspace.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That environment already housed two other skills — from my &lt;strong&gt;other published open-source repo&lt;/strong&gt;, goal-workflow-designer: &lt;code&gt;/goal&lt;/code&gt; (depth: polish one task into an iterable goal prompt) and &lt;code&gt;workflow-shaper&lt;/code&gt; (breadth: fan one check out across N units). Four skills now shared one environment, so I ran the Conflict Analysis Mode for the first time.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Result: 5 conflicts across 6 pairings — spanning both repos.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;But the most valuable finding wasn't any single conflict — it was the &lt;strong&gt;structural one&lt;/strong&gt;: every deferral clause in the environment was one-directional, and all of them routed &lt;em&gt;around&lt;/em&gt; &lt;code&gt;goal&lt;/code&gt; — the skill with the most generic name and exclusive ownership of the &lt;code&gt;/goal&lt;/code&gt; command had &lt;strong&gt;zero outbound deferrals.&lt;/strong&gt; It was the ecosystem's &lt;strong&gt;accretion point&lt;/strong&gt;: generic phrasing fell into it, and it never handed anything off. The two worst conflicts were both symptoms of that one structure.&lt;/p&gt;

&lt;p&gt;On re-review, I also overturned one of the report's classifications. "goal vs goal-decomposer homonym" had been filed as namespace pollution — but a word-by-word comparison showed both use the &lt;strong&gt;same five-element format&lt;/strong&gt; (outcome / verification / constraint / iteration policy / error handling). The G*.md files goal-decomposer produces are, in essence, &lt;em&gt;pre-filled goal prompts&lt;/em&gt;, directly executable by /goal's KICK-OFF machinery. &lt;strong&gt;The name collision was actually an unclaimed integration point.&lt;/strong&gt; The right fix wasn't mutual avoidance — it was declaring the five-element format a shared standard across both repos.&lt;/p&gt;

&lt;p&gt;The repair left three cross-repo governance lessons:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Fix the source, not the installed copy.&lt;/strong&gt; The conflicts spanned two independently versioned repos; patching only the install directory means the next reinstall reverts everything. Every fix landed in both repos' sources.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cross-repo deferrals must be conditional.&lt;/strong&gt; Hard-coding "defer to idea-to-spec" breaks in environments where it isn't installed. Every clause reads "if installed, defer; if not, continue but flag the scope difference."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;When two terms collide, look for a shared standard first.&lt;/strong&gt; If the underlying format is the same, declaring a standard beats drawing a boundary.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;One small surprise during the repair: inspecting the install topology revealed that the new skills' global install directories were &lt;strong&gt;junctions pointing straight at the repo source&lt;/strong&gt; — edit the source, and it's live everywhere instantly. The "source-as-deployment" structure for a solo multi-repo developer had grown by itself.&lt;/p&gt;




&lt;h2&gt;
  
  
  20. Coexist or merge?
&lt;/h2&gt;

&lt;p&gt;The final question: should the two repos be merged into one complete body?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;My decision: coexist, don't merge.&lt;/strong&gt; Three reasons:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Different audiences.&lt;/strong&gt; goal-workflow-designer is published and self-sufficient; many users want /goal and nothing else. Merging would force the whole spec pipeline on them.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Different lifecycles.&lt;/strong&gt; One side is iterating at a stable 0.2.x; the other hasn't shipped 0.1. Coupled, every breaking change in the new repo would drag down the stable repo's users.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Something better than merging already exists&lt;/strong&gt;: a shared standard (the five-element format) + conditional mutual deferrals + the same routing table in both READMEs. The relationship is &lt;strong&gt;compiler and runtime&lt;/strong&gt; — spec-sonar compiles specs into G*.md; goal-workflow-designer's KICK-OFF executes them. Each is valuable alone; together they form the full pipeline.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The one discipline to hold: &lt;strong&gt;the five-element format is the contract between the two repos.&lt;/strong&gt; Changing it on either side is a cross-repo breaking change, requiring synchronized version bumps.&lt;/p&gt;




&lt;h2&gt;
  
  
  Interlude: the tool missed its own dark zone
&lt;/h2&gt;

&lt;p&gt;The first real convergence run after the package shipped (a marketing-funnel spec) surfaced two user confusions: the Q&amp;amp;A rounds arrived as plain text (no native Claude Code selector window), and the raw &lt;code&gt;&amp;lt;STATE&amp;gt;&lt;/code&gt; JSON block was mistaken for "the debris of a failed trigger."&lt;/p&gt;

&lt;p&gt;The root cause is ironic. The SKILL.md had hard-coded the lowest-common-denominator text format — the correct answer for claude.ai's plain chat, but "interaction style must adapt to the environment" had never been written into the spec. &lt;strong&gt;The tool built to detect unstated requirements had an unstated requirement of its own.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The control group sat right next door: the goal skill says "Use AskUserQuestion progressively" from day one, and never had this problem. The behavioral difference between the two skills wasn't model mood — it was spec difference. Which once again validates the behavioral-contract philosophy: the model executes the contract faithfully, and &lt;strong&gt;whatever the contract lacks, the behavior lacks.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The fix (v0.1.2): a new "Step 0: environment detection &amp;amp; interaction mode." In Claude Code, enumerable questions prefer the native selector and STATE persists to a file instead of the chat; in plain-chat environments, fall back to text mode with the STATE block always labeled "system bookkeeping — safe to ignore." Plus a hard-coded three-level degradation chain: native selector → on failure, structured text options (reply with a single letter A/B/C) → no tool at all, pure text. And a mandatory one-line human progress display every round.&lt;/p&gt;

&lt;p&gt;A boundary worth stating: the selector's UI rendering, option exclusivity, the automatic "Other" escape hatch, and answer return are all native harness capabilities — the skill never worries about them. But "what to do when the tool fails" is model discretion — unless it's written into the spec. The degradation chain reclaims exactly that discretion.&lt;/p&gt;

&lt;p&gt;The principle, in one sentence: &lt;strong&gt;bookkeeping is for the system; progress is for the human.&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  21. Thinking patterns added in Part 2
&lt;/h2&gt;

&lt;p&gt;Part 1 accumulated eight reusable patterns. Part 2 adds seven:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;9. Pre-adjudication.&lt;/strong&gt; Make every design decision at "compile time"; the executor only follows the letter. A weak model's reliability comes not from pretty formatting but from having its degrees of freedom spent before it gets the task.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;10. The cold-start test as the pass bar.&lt;/strong&gt; An artifact is acceptable not when its author is satisfied, but when a zero-context, weakest-tier executor can answer the four questions. If it can't — rewrite the artifact, don't upgrade the executor.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;11. Give the executor a legal "I don't know" exit.&lt;/strong&gt; Without a BLOCKED protocol, models fill gaps with hallucination. An explicit stuck-report format beats any "please be honest" exhortation.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;12. Conflict as a feature.&lt;/strong&gt; Install into a real environment and run conflict analysis; every hit is free design feedback. The tool's first real-world catch was my own repos — more persuasive than any synthetic example.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;13. Fix the source, not the copy.&lt;/strong&gt; Anything that gets "installed somewhere" must have fixes land in its version-controlled source, or the next deployment undoes them.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;14. On a name collision, look for the shared standard first.&lt;/strong&gt; When two concepts collide, compare their underlying structures. A shared standard is a stronger resolution than naming isolation.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;15. Interaction style is part of the spec.&lt;/strong&gt; &lt;em&gt;How to ask&lt;/em&gt; is not model discretion. Which interaction primitive a skill uses per environment (native selector / structured text options / plain text), and how it degrades on failure, belong in the behavioral contract — or you'll deliver the worst experience on the best platform.&lt;/p&gt;




&lt;h2&gt;
  
  
  Epilogue: dogfooding all the way down
&lt;/h2&gt;

&lt;p&gt;Looking back, the most consistent pattern in this whole process is the tool relentlessly eating its own dog food:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The design itself was converged using spec-sonar's own process (Part 1, section 11).&lt;/li&gt;
&lt;li&gt;The tool's first audit caught &lt;strong&gt;its own schema drift.&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;The tool's first real conflict report caught &lt;strong&gt;conflicts between my own two repos.&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;And that conflict report itself became an official example file in the repo.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For a tool that claims to find "what you don't know you don't know," the best validation is to keep pointing it at itself. So far, it has found something every single time.&lt;/p&gt;




&lt;blockquote&gt;
&lt;p&gt;← &lt;a href="https://dev.to/content/spec-sonar-design-journal-part-1-en"&gt;Back to Part 1&lt;/a&gt; · Repos: &lt;a href="https://github.com/dragon375014/spec-sonar" rel="noopener noreferrer"&gt;spec-sonar&lt;/a&gt; × &lt;a href="https://github.com/dragon375014/goal-workflow-designer" rel="noopener noreferrer"&gt;goal-workflow-designer&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;p&gt;&lt;em&gt;本文原載於我的部落格：&lt;a href="https://coffeeshooters.com/content/spec-sonar-design-journal-part-2-en?utm_source=devto&amp;amp;utm_medium=social&amp;amp;utm_campaign=blog-spec-sonar-design-journal-part-2-en" rel="noopener noreferrer"&gt;spec-sonar: The Complete Design Record (Part 2) — Fable's Review, a Real-World Conflict, and the Two-Repo Decision&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>solodev</category>
    </item>
    <item>
      <title>spec-sonar: The Complete Design Record (Part 1) — From One Idea to an Open-Source Toolchain</title>
      <dc:creator>Dexterlung</dc:creator>
      <pubDate>Wed, 05 Aug 2026 13:05:13 +0000</pubDate>
      <link>https://dev.to/dexterlung/spec-sonar-the-complete-design-record-part-1-from-one-idea-to-an-open-source-toolchain-4k7h</link>
      <guid>https://dev.to/dexterlung/spec-sonar-the-complete-design-record-part-1-from-one-idea-to-an-open-source-toolchain-4k7h</guid>
      <description>&lt;p&gt;This is one long conversation — about 40+ exchanges — in which I went from a very vague idea to the initial design of an open-source toolchain called spec-sonar. Part 1 ends where I hand 13 review questions to Fable 5.&lt;/p&gt;

&lt;p&gt;Part 2 covers the next 24 hours: the review results, 25 files shipped, and a plot twist I never saw coming — the tool catching conflicts between my own two open-source repos.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Read on: &lt;a href="https://dev.to/content/spec-sonar-design-journal-part-2-en"&gt;Part 2 — Fable's review and a real-world conflict&lt;/a&gt; · &lt;a href="https://dev.to/content/spec-sonar-design-journal-part-1"&gt;繁體中文版&lt;/a&gt; · this tool is listed in my &lt;a href="https://dev.to/content/dev-toolkit"&gt;open-source toolkit index&lt;/a&gt;.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  1. The starting point: an intuition
&lt;/h2&gt;

&lt;p&gt;The first thing I typed into Claude was this:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"I'm wondering whether I can use the standard body of software-engineering knowledge, Claude Code's round-based questioning, and a well-designed skill to help people new to AI-assisted development turn ideas into reality — confirm the safety boundaries, set design limits to prevent scope creep, settle the architecture, and finally produce a series of execution playbooks or goals…"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Several key intuitions were already in that sentence:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The concept of &lt;strong&gt;dark zones and bright zones&lt;/strong&gt; — use set difference to find the questions worth asking.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Convergence, not generation&lt;/strong&gt; — the goal is fewer rework cycles, not faster output.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Complexity loaded on demand&lt;/strong&gt; — I decide how deep to go, instead of getting everything at full depth.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  2. First turning point: interrogate before designing
&lt;/h2&gt;

&lt;p&gt;I forced myself to answer one question first: &lt;strong&gt;is this reinventing the wheel?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;What the search turned up:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Requirements-management tools for PMs/BAs exist, but they assume you already understand User Stories.&lt;/li&gt;
&lt;li&gt;Spec-driven tools for developers exist, but they assume a spec draft already exists.&lt;/li&gt;
&lt;li&gt;Claude Code's CLAUDE.md ecosystem exists, but nobody generates specs &lt;em&gt;for non-technical users&lt;/em&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Conclusion: not a reinvented wheel — but the differentiation has to be explicit.&lt;/strong&gt; The real difference:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Existing tools' route: make the AI smarter so it tolerates vague input
spec-sonar's route:    make the user's input itself precise
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The latter is sturdier, because it doesn't depend on a model's tolerance for ambiguity — and that tolerance is exactly where hallucination comes from.&lt;/p&gt;




&lt;h2&gt;
  
  
  3. Second turning point: reframing the value proposition
&lt;/h2&gt;

&lt;p&gt;I first wrote the value proposition as "reduce design rework." &lt;strong&gt;That was wrong.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The problem: a non-technical user can't perceive that value — they don't know how many detours they would otherwise have taken, so they can't feel the detours they avoided.&lt;/p&gt;

&lt;p&gt;The corrected value proposition: &lt;strong&gt;"reduce the anxiety of landing something vague."&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That's trust, not efficiency. What the user actually feels is: &lt;em&gt;"I told it what I wanted, it told me the problems I hadn't thought of, and what got built matches what I had in mind."&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;This reframe shaped every later decision:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The tone of the Q&amp;amp;A — assistance, not interrogation.&lt;/li&gt;
&lt;li&gt;How infeasibility is handled — not rejection, but shrinking to a feasible version.&lt;/li&gt;
&lt;li&gt;Showing bright/dark-zone progress every round — the core trust-building mechanism.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  4. Core system-design decisions
&lt;/h2&gt;

&lt;h3&gt;
  
  
  4.1 A behavioral contract, not roleplay
&lt;/h3&gt;

&lt;p&gt;My first idea was to inject a "you are a senior architect" persona prompt.&lt;/p&gt;

&lt;p&gt;The problem: &lt;strong&gt;an identity is not a set of behavior rules.&lt;/strong&gt; Claude plays the architect but doesn't know what this architect must do, must not do, or when to stop.&lt;/p&gt;

&lt;p&gt;So I switched to injecting a &lt;strong&gt;behavioral contract&lt;/strong&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Inject a behavioral contract, not an identity:
"Your job is to converge a software product idea into an executable spec.
 You are not a roleplaying architect.
 You are a system with fixed rules of behavior."
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This made the system's behavior predictable, testable, and iterable.&lt;/p&gt;

&lt;h3&gt;
  
  
  4.2 STATE: externalizing conversational memory
&lt;/h3&gt;

&lt;p&gt;The problem: there is no persistent memory between rounds; by round 4 the model has forgotten what round 1 established.&lt;/p&gt;

&lt;p&gt;My fix was to &lt;strong&gt;externalize the conversation memory&lt;/strong&gt; into a portable JSON block:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight xml"&gt;&lt;code&gt;&lt;span class="nt"&gt;&amp;lt;STATE&amp;gt;&lt;/span&gt;
{
  "round": 2,
  "bright": ["established items"],
  "dark": ["unconfirmed items"],
  "next_focus": "next round's topic"
}
&lt;span class="nt"&gt;&amp;lt;/STATE&amp;gt;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Key principles:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Wrap it in an XML tag — more stable than plain text; Claude parses its own XML output more reliably.&lt;/li&gt;
&lt;li&gt;The &lt;code&gt;next_focus&lt;/code&gt; field forces each round onto a single topic, preventing scattershot questioning.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;STATE_FINAL.json&lt;/code&gt; is portable — you can leave mid-process and resume later with full state.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  4.3 The baseline problem of the set-difference method
&lt;/h3&gt;

&lt;p&gt;Doing gap analysis, I found a hole:&lt;/p&gt;

&lt;p&gt;Set difference only works if I already hold a standard template of "what a complete system looks like" to diff against. No baseline, no difference.&lt;/p&gt;

&lt;p&gt;The solution is a &lt;code&gt;dark-zone-baseline.md&lt;/code&gt;: 10 standard dimensions that seed every new session's dark list — target users, core feature boundary, data model, auth &amp;amp; permissions, third-party integrations, deployment platform, performance &amp;amp; scale, maintenance &amp;amp; updates, budget &amp;amp; timeline, success criteria.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The most important detail: the Session-vs-Persistent fork.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I only discovered this after running the real test case (a math tower-defense game). The "do we need an account system?" decision had been &lt;em&gt;back-derived&lt;/em&gt; from "does the data persist?", rather than asked proactively. The fork has enormous architectural impact, so I promoted it to a mandatory proactive question.&lt;/p&gt;




&lt;h2&gt;
  
  
  5. The real test: a math tower-defense game
&lt;/h2&gt;

&lt;p&gt;Why this case: a non-technical user (with an engineer partner) wanted to land "Plants-vs-Zombies tower defense × middle-school math × in-class team battle." I used it to stress-test my tool.&lt;/p&gt;

&lt;p&gt;What ran: the full 5-round convergence, from "I don't know what to build" all the way to a complete README.md + CLAUDE.md + STATE_FINAL.json.&lt;/p&gt;

&lt;p&gt;Key observations:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;The resource mechanic was the highest-value discussion.&lt;/strong&gt; That user raised the real classroom pain point — "a team-captain system creates conflict between teammates." The final design: &lt;em&gt;personal earn, personal spend, plus per-unit cooldowns&lt;/em&gt; — individual effort converts directly into individual agency, with no middleman to skim.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Wheel-reinvention detection confirmed the differentiation.&lt;/strong&gt; Single-player math tower-defense games exist; "whole class online + team battle + teacher console + realtime multiplayer" does not.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;The dark zones cleared after 6 rounds — but one was never proactively probed.&lt;/strong&gt; The Session-vs-Persistent fork emerged by inference from the "no accounts" answer, not from the baseline. That bug directly drove the baseline fix.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The test case's role: the math game is &lt;strong&gt;not&lt;/strong&gt; spec-sonar's target scenario. It's a deliberately hard &lt;strong&gt;stress test&lt;/strong&gt; — WebSocket realtime sync, a game engine, multiple dependency types, ephemeral-session design. Only that level of complexity genuinely probes the system.&lt;/p&gt;




&lt;h2&gt;
  
  
  6. Third turning point: from a tool to a toolchain
&lt;/h2&gt;

&lt;p&gt;The original design was a single requirements-convergence skill (idea-to-spec). It kept growing:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;idea-to-spec (requirements convergence)
    ↓ problem found: CLAUDE.md outputs modules, but no dependencies
goal-decomposer (goal decomposition)
    ↓ problem found: only works from zero; can't diagnose existing projects
Audit Mode (existing-project diagnosis)
    ↓ problem found: large complex systems (ERP) have different needs
Complex System Mode
    ↓ problem found: Claude-Code-only is too narrow
Adapters layer (platform-agnostic universal format)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;The most consequential evolution: from Claude-specific to platform-agnostic.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;Core spec layer (universal)&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;  &lt;span class="s"&gt;goal-graph.json + goals/*.md&lt;/span&gt;
&lt;span class="na"&gt;Adapter layer (per platform)&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="s"&gt;adapters/CLAUDE.md               ← Claude Code&lt;/span&gt;
  &lt;span class="s"&gt;adapters/.cursor/rules           ← Cursor&lt;/span&gt;
  &lt;span class="s"&gt;adapters/copilot-instructions.md ← Copilot&lt;/span&gt;
  &lt;span class="s"&gt;adapters/system-prompt.md        ← any AI agent&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That decision turned spec-sonar from "a Claude Code tool" into "a general requirements-engineering tool for the whole AI-coding ecosystem."&lt;/p&gt;




&lt;h2&gt;
  
  
  7. The philosophy: one expensive deep design → cheap structured execution
&lt;/h2&gt;

&lt;p&gt;The core philosophy of the whole toolchain, surfaced while I was working out how goal-decomposer integrates with goal-workflow-designer.&lt;/p&gt;

&lt;p&gt;The background problem is concrete: &lt;strong&gt;not everyone can afford Fable-tier model costs.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;My insight:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Fable (expensive) thinks deeply once
    → the thinking is frozen into a goal graph and individual goal files
    → any model (Haiku, Sonnet) executes by following the graph
    → Fable-tier reasoning is no longer required
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The analogy is compilation:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Fable = the compiler that does the complex reasoning&lt;/li&gt;
&lt;li&gt;goal-graph.json = the compiled executable&lt;/li&gt;
&lt;li&gt;any model = the CPU that runs it&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The smartest (and most expensive) model appears only once, at design time, to settle all the hard decisions; everything after that runs literally on cheap models.&lt;/p&gt;




&lt;h2&gt;
  
  
  8. Naming: spec-sonar
&lt;/h2&gt;

&lt;p&gt;I considered SpecForge, DarkMap, ConvergeKit, and ReqSonar. I picked &lt;strong&gt;spec-sonar&lt;/strong&gt; because it passed three tests:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Searchability&lt;/strong&gt;: no same-name repo on GitHub.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Explanation cost&lt;/strong&gt;: "like sonar, it finds what's submerged in your spec" — one sentence.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Works in both languages&lt;/strong&gt;: "I sonar-scanned the requirements" reads naturally.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The logic behind the name: sonar's core mechanism is &lt;em&gt;using waves you can't see to find things you can't see&lt;/em&gt;. spec-sonar's core mechanism is &lt;em&gt;using questions you didn't ask to find requirements you didn't think of&lt;/em&gt;. Form matches content.&lt;/p&gt;




&lt;h2&gt;
  
  
  9. Compatibility principles (the last major decisions)
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Compatibility with existing CLAUDE.md files.&lt;/strong&gt; spec-sonar is a design-time tool; CLAUDE.md is a build-time artifact — naturally on different timelines. So spec-sonar installs non-destructively: it never overwrites your existing CLAUDE.md, it only appends one reference line.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Conflict detection across existing skills.&lt;/strong&gt; I upgraded the goal from "avoid conflicts" to &lt;strong&gt;"conflict as a feature"&lt;/strong&gt;: spec-sonar's Conflict Analysis Mode reads all your installed skills and outputs a conflict map. That turns it from "yet another skill" into "a governance tool for your skill ecosystem."&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Load on demand: three install tiers&lt;/strong&gt; — Lite (1 skill) for small projects/MVPs/anything finishable in a week; Standard (2 skills) for most projects under 3 months; Pro (full suite) for enterprise/ERP/long-term maintenance.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  10. The role of Fable 5
&lt;/h2&gt;

&lt;p&gt;On 2026-06-09 Anthropic released Claude Fable 5 — the first public model of the Mythos family. Its edge: "the longer and more complex the task, the wider the lead," with a particular strength in one-shotting a complete design.&lt;/p&gt;

&lt;p&gt;My prompt strategy wasn't "what do you think of this design?" — it was &lt;strong&gt;"under what conditions does this design break, and fix it."&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I asked Fable to do five things:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Review the CLAUDE.md format from an &lt;em&gt;executor's&lt;/em&gt; perspective ("could I start work from this spec?").&lt;/li&gt;
&lt;li&gt;Find the dark-zone baseline's systematic omissions per product type.&lt;/li&gt;
&lt;li&gt;Design STATE's conflict handling (retracted bright items, contradictions, late additions).&lt;/li&gt;
&lt;li&gt;Design the complete goal-decomposer SKILL.md (dependency inference, model-tier assignment).&lt;/li&gt;
&lt;li&gt;Design the adapter format for every platform.&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  11. This design was designed with spec-sonar
&lt;/h2&gt;

&lt;p&gt;The most meta observation worth recording: this very design started from "I have a vague idea" and went through scope judgment, bright/dark separation, set-difference questioning, infeasibility detection (none found), complexity calibration, and wheel-reinvention detection — and that whole flow &lt;strong&gt;is spec-sonar's own flow.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The tool's first user was the person designing the tool.&lt;/p&gt;




&lt;h2&gt;
  
  
  12. Final output list
&lt;/h2&gt;

&lt;p&gt;By the end of the conversation:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="go"&gt;idea-to-spec-v1.1.zip
├── SKILL.md                      the convergence engine
├── references/
│   ├── dark-zone-baseline.md     10 dimensions (with the Session-vs-Persistent fix)
│   └── output-templates.md       three output format templates
├── examples/
│   ├── README.md                 math tower-defense case (human-readable)
│   ├── CLAUDE.md                 math tower-defense case (executable spec)
│   └── STATE_FINAL.json          final convergence state
└── docs/
    ├── fable-review-prompt.md    the full Fable 5 review prompt (13 questions, 5 jobs)
    └── spec-sonar-README.md      open-source README draft
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Pending Fable 5's reply: goal-decomposer/SKILL.md, project-scanner.py, the adapters/ directory (4 platforms), audit-mode, Conflict Analysis Mode, the tiered INSTALL.md.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;(Every item on that list ships in Part 2 — plus three spec holes, one schema self-audit, and a real-world conflict analysis spanning two open-source repos.)&lt;/em&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  13. Reusable thinking patterns (Part 1)
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Interrogate before designing.&lt;/strong&gt; Every new idea first answers "is this a reinvented wheel?" — with search results, not gut feeling.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Dark zones matter more than bright zones.&lt;/strong&gt; Stated requirements are rarely the problem; unstated assumptions are the root of failure.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Behavioral contracts beat roleplay.&lt;/strong&gt; Tell the AI its rules of behavior for this session, not who it is.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Externalized state solves context loss.&lt;/strong&gt; Multi-round tools must carry state explicitly every round.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Find design gaps with real cases.&lt;/strong&gt; After the theory, run one real case and let the gaps surface (that's how the Session-vs-Persistent fork was found).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Expensive thinking once → cheap structured execution.&lt;/strong&gt; The cost-optimal AI strategy.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Universal core + per-platform adapters.&lt;/strong&gt; The more universal the core format, the longer the tool lives.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Load on demand.&lt;/strong&gt; Tiered installation lets users choose their complexity instead of paying full cost by default.&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  Coming in Part 2
&lt;/h2&gt;

&lt;p&gt;Part 1 ends with 13 questions packed into a review prompt and handed to Fable 5. Part 2 records what happened next:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Fable reviews the test case &lt;em&gt;as an executor&lt;/em&gt; and finds &lt;strong&gt;three day-one landmines&lt;/strong&gt; in a seemingly complete spec.&lt;/li&gt;
&lt;li&gt;The tool audits its own output and catches &lt;strong&gt;its own schema drift&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;The full goal-decomposer design: dependency inference, contract freezing, pre-adjudication, the cold-start test.&lt;/li&gt;
&lt;li&gt;25 files shipped as a complete open-source package.&lt;/li&gt;
&lt;li&gt;Then the real plot twist: installed into a production workspace, the Conflict Analysis Mode &lt;strong&gt;finds 5 conflicts between my own two open-source repos&lt;/strong&gt; — and one structural problem nobody had seen.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;→ &lt;a href="https://dev.to/content/spec-sonar-design-journal-part-2-en"&gt;Continue to Part 2&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;本文原載於我的部落格：&lt;a href="https://coffeeshooters.com/content/spec-sonar-design-journal-part-1-en?utm_source=devto&amp;amp;utm_medium=social&amp;amp;utm_campaign=blog-spec-sonar-design-journal-part-1-en" rel="noopener noreferrer"&gt;spec-sonar: The Complete Design Record (Part 1) — From One Idea to an Open-Source Toolchain&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>solodev</category>
    </item>
    <item>
      <title>Audio Plays on Desktop but Not on iPhone / iPad — The Culprit Is the MP4 moov Atom</title>
      <dc:creator>Dexterlung</dc:creator>
      <pubDate>Sun, 02 Aug 2026 13:05:25 +0000</pubDate>
      <link>https://dev.to/dexterlung/audio-plays-on-desktop-but-not-on-iphone-ipad-the-culprit-is-the-mp4-moov-atom-1j01</link>
      <guid>https://dev.to/dexterlung/audio-plays-on-desktop-but-not-on-iphone-ipad-the-culprit-is-the-mp4-moov-atom-1j01</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;Series: Engineering gotchas&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;I put a batch of podcast audio files (&lt;code&gt;.m4a&lt;/code&gt;) on Cloudflare R2 and embedded them with a plain &lt;code&gt;&amp;lt;audio&amp;gt;&lt;/code&gt; tag. Desktop Chrome and Firefox played them perfectly. But on &lt;strong&gt;iPhone and iPad, they wouldn't play at all&lt;/strong&gt; — tap the button, nothing happens, no error message, console completely clean.&lt;/p&gt;

&lt;p&gt;The most maddening part: &lt;strong&gt;it throws no error.&lt;/strong&gt; You have no idea where to even start looking.&lt;/p&gt;

&lt;p&gt;This is how I tracked it down to the root cause, and the one-line fix at the end. If you googled "audio plays on desktop but not iphone" or "m4a not playing ios safari" and landed here, skip to section three.&lt;/p&gt;

&lt;h2&gt;
  
  
  First, rule out all the usual suspects
&lt;/h2&gt;

&lt;p&gt;iOS Safari is picky about audio, and the three most common answers online are: wrong MIME type, no HTTP Range support, or bad codec. I checked them one by one — and they were &lt;strong&gt;all correct.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# 1) Is the Content-Type right?&lt;/span&gt;
curl &lt;span class="nt"&gt;-sI&lt;/span&gt; &lt;span class="s2"&gt;"https://your-domain/xxx.m4a"&lt;/span&gt;
&lt;span class="c"&gt;# → Content-Type: audio/mp4   ✅ correct (m4a is audio/mp4)&lt;/span&gt;
&lt;span class="c"&gt;# → Accept-Ranges: bytes      ✅ Range supported&lt;/span&gt;

&lt;span class="c"&gt;# 2) Does a Range request actually return 206? (iOS strictly requires this)&lt;/span&gt;
curl &lt;span class="nt"&gt;-s&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; /dev/null &lt;span class="nt"&gt;-D&lt;/span&gt; - &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Range: bytes=0-99"&lt;/span&gt; &lt;span class="s2"&gt;"https://your-domain/xxx.m4a"&lt;/span&gt; | &lt;span class="nb"&gt;grep &lt;/span&gt;HTTP
&lt;span class="c"&gt;# → HTTP/1.1 206 Partial Content   ✅ correct&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The server side is &lt;strong&gt;entirely correct.&lt;/strong&gt; Content-Type is &lt;code&gt;audio/mp4&lt;/code&gt;, Range is supported, it returns &lt;code&gt;206 Partial Content&lt;/code&gt;. R2 is fine. CORS isn't the issue either (an &lt;code&gt;&amp;lt;audio&amp;gt;&lt;/code&gt; without the &lt;code&gt;crossorigin&lt;/code&gt; attribute plays cross-origin without needing CORS).&lt;/p&gt;

&lt;p&gt;And the codec is &lt;code&gt;.m4a&lt;/code&gt; (AAC) — the format iOS supports most natively. All three common culprits ruled out. So what is it?&lt;/p&gt;

&lt;h2&gt;
  
  
  The real culprit: the moov atom is at the end of the file (no faststart)
&lt;/h2&gt;

&lt;p&gt;An MP4 / m4a file is made of boxes (also called atoms). Two matter most here:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;mdat&lt;/code&gt;&lt;/strong&gt;: the actual audio / video data (large — 11 MB in my file)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;moov&lt;/code&gt;&lt;/strong&gt;: the playback index / metadata (timeline, sample tables — the player must read this &lt;em&gt;first&lt;/em&gt; to know how to decode)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Here's the problem: &lt;strong&gt;ffmpeg's default output puts &lt;code&gt;mdat&lt;/code&gt; first and &lt;code&gt;moov&lt;/code&gt; at the very end of the file.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Desktop browsers (Chrome / Firefox)&lt;/strong&gt; are smarter / more aggressive: they fire a Range request for the &lt;strong&gt;end&lt;/strong&gt; of the file to grab &lt;code&gt;moov&lt;/code&gt;, find the index, and play.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;iOS Safari / WebKit&lt;/strong&gt;, with &lt;code&gt;&amp;lt;audio preload="metadata"&amp;gt;&lt;/code&gt;, only fetches the &lt;strong&gt;beginning&lt;/strong&gt; of the file. No &lt;code&gt;moov&lt;/code&gt; at the start → it can't find the index → &lt;strong&gt;it silently gives up and won't play.&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That's why it plays on desktop but not on iPhone / iPad. Both iPhone and iPad use WebKit, so they fail together.&lt;/p&gt;

&lt;p&gt;Moving &lt;code&gt;moov&lt;/code&gt; to the front of the file is called &lt;strong&gt;faststart&lt;/strong&gt; — the standard trick that lets video / audio "play while downloading."&lt;/p&gt;

&lt;h2&gt;
  
  
  How to confirm it (30 seconds)
&lt;/h2&gt;

&lt;p&gt;Grab the first 48 bytes and check whether &lt;code&gt;ftyp&lt;/code&gt; is followed by &lt;code&gt;moov&lt;/code&gt; or &lt;code&gt;mdat&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-s&lt;/span&gt; &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Range: bytes=0-47"&lt;/span&gt; &lt;span class="s2"&gt;"https://your-domain/xxx.m4a"&lt;/span&gt; | xxd | &lt;span class="nb"&gt;head&lt;/span&gt; &lt;span class="nt"&gt;-3&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Broken (no faststart)&lt;/strong&gt; — &lt;code&gt;mdat&lt;/code&gt; right after &lt;code&gt;ftyp&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;00000000: 0000 001c 6674 7970 4d34 4120 ...  ....ftypM4A
00000020: 6672 6565 00a8 3b91 6d64 6174 ...  free..;.mdat   ← mdat first, moov at the end
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Good (faststart)&lt;/strong&gt; — &lt;code&gt;moov&lt;/code&gt; right after &lt;code&gt;ftyp&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;00000000: 0000 001c 6674 7970 4d34 4120 ...  ....ftypM4A
00000020: 6d6f 6f76 0000 006c 6d76 6864 ...  moov...lmvhd   ← moov right after ftyp ✅
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If &lt;code&gt;mdat&lt;/code&gt; comes before &lt;code&gt;moov&lt;/code&gt;, that's your bug.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix: one line of ffmpeg, lossless, no re-encode
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ffmpeg &lt;span class="nt"&gt;-i&lt;/span&gt; input.m4a &lt;span class="nt"&gt;-c&lt;/span&gt; copy &lt;span class="nt"&gt;-movflags&lt;/span&gt; +faststart output.m4a
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The key is &lt;strong&gt;&lt;code&gt;-c copy&lt;/code&gt;&lt;/strong&gt; — it &lt;strong&gt;does not re-encode&lt;/strong&gt;; it just copies the existing stream as-is and relocates &lt;code&gt;moov&lt;/code&gt; to the front. So:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Lossless&lt;/strong&gt;: audio quality is untouched&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Same size&lt;/strong&gt; (mine was 11,258,702 bytes before and after — not a single byte different)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Instant&lt;/strong&gt;: 11 MB in under a second&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Then &lt;strong&gt;re-upload to overwrite&lt;/strong&gt; the original. If your files are on R2 / S3, keep &lt;code&gt;Content-Type: audio/mp4&lt;/code&gt; on the upload.&lt;/p&gt;

&lt;p&gt;(If a CDN like Cloudflare sits in front, check whether the edge cached the old file. Mine was &lt;code&gt;cf-cache-status: DYNAMIC&lt;/code&gt; — not cached — so the overwrite took effect immediately. The only remaining cache is the phone browser's &lt;strong&gt;local&lt;/strong&gt; cache — test in a private tab for a clean result.)&lt;/p&gt;

&lt;h2&gt;
  
  
  How to stop it from happening again
&lt;/h2&gt;

&lt;p&gt;If you have a "generate audio → upload" pipeline, &lt;strong&gt;always run it through faststart before uploading&lt;/strong&gt; — never push the raw ffmpeg / converter output straight up. I baked it into the upload script as a mandatory step: remux faststart → upload → auto-verify with curl that &lt;code&gt;moov&lt;/code&gt; really is at the front, all three together. New audio never hits this bug again.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;This bug is a good reminder: &lt;strong&gt;"works on my machine" and "works on the user's device" are two different things.&lt;/strong&gt; The server config was all correct, the file was there, desktop played fine — every "obvious" place was clean, and the defect hid in the order of the atoms &lt;em&gt;inside&lt;/em&gt; the MP4. Problems that only surface on real devices (especially iOS) slip right past automated tests and desktop development. Next time you hit "fine on desktop, broken on mobile," don't just stare at your code — think &lt;strong&gt;file format / platform differences.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;







&lt;p&gt;&lt;em&gt;本文原載於我的部落格：&lt;a href="https://coffeeshooters.com/content/m4a-audio-not-playing-on-ios-faststart-en?utm_source=devto&amp;amp;utm_medium=social&amp;amp;utm_campaign=blog-m4a-audio-not-playing-on-ios-faststart-en" rel="noopener noreferrer"&gt;Audio Plays on Desktop but Not on iPhone / iPad — The Culprit Is the MP4 moov Atom&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>solodev</category>
    </item>
    <item>
      <title>Series reading guide — How to pick which of these 9 Trace Lock posts to read</title>
      <dc:creator>Dexterlung</dc:creator>
      <pubDate>Fri, 31 Jul 2026 14:34:55 +0000</pubDate>
      <link>https://dev.to/dexterlung/series-reading-guide-how-to-pick-which-of-these-9-trace-lock-posts-to-read-48jm</link>
      <guid>https://dev.to/dexterlung/series-reading-guide-how-to-pick-which-of-these-9-trace-lock-posts-to-read-48jm</guid>
      <description>&lt;p&gt;&lt;strong&gt;May 2026&lt;/strong&gt; · Series "Trace Lock — Governance notes from pairing with AI to write code" · Post 5 of 9&lt;/p&gt;




&lt;p&gt;This series has 9 posts total. They're an organized record of conversations I had with Claude between late April and end of May 2026.&lt;/p&gt;

&lt;p&gt;After writing the first 4, I re-read them and noticed a problem: &lt;strong&gt;readers won't know where to start&lt;/strong&gt;. The difficulty range goes from beginner to advanced. The content jumps from personal narrative to engineering implementation details. Reading from post 1 to post 9 in order will probably make you give up by post 3.&lt;/p&gt;

&lt;p&gt;So this post is a map, not content. If you only read this one and leave, that's fine with me.&lt;/p&gt;




&lt;h2&gt;
  
  
  What this series is about
&lt;/h2&gt;

&lt;p&gt;In one sentence: &lt;strong&gt;when pairing with AI to write code, cross-layer bugs kept happening, so Claude and I worked out two governance patterns that pair together&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Both patterns are working names I gave them myself (not industry-standard terminology):&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Defensive Trace Lock&lt;/strong&gt;: after tripping on a cross-layer bug once, lock down that relationship so it doesn't rot later&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Offensive audit&lt;/strong&gt;: proactively audit a business flow (e.g. "order to shipment"), find all unprotected chain nodes, fix them in one pass&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Combined, they form what I call the "dual-blade" approach. The 9 posts spread these 2 patterns across 3 layers: plain-language version, engineering version, and meta reflection.&lt;/p&gt;




&lt;h2&gt;
  
  
  Map of the 9 posts
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Post&lt;/th&gt;
&lt;th&gt;Topic&lt;/th&gt;
&lt;th&gt;Series category&lt;/th&gt;
&lt;th&gt;Difficulty&lt;/th&gt;
&lt;th&gt;Word count&lt;/th&gt;
&lt;th&gt;Who it's for&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;E&lt;/td&gt;
&lt;td&gt;Meta: how the methodology emerged from conversations with AI&lt;/td&gt;
&lt;td&gt;Cross-field diary&lt;/td&gt;
&lt;td&gt;beginner&lt;/td&gt;
&lt;td&gt;EN 1300 / ZH 1700&lt;/td&gt;
&lt;td&gt;Wants to know how the whole thing came together&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;This post&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Series reading guide&lt;/td&gt;
&lt;td&gt;Cross-field diary&lt;/td&gt;
&lt;td&gt;beginner&lt;/td&gt;
&lt;td&gt;EN 700 / ZH 900&lt;/td&gt;
&lt;td&gt;Wants to decide which posts to read&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A1&lt;/td&gt;
&lt;td&gt;Defensive Trace Lock plain-language intro&lt;/td&gt;
&lt;td&gt;Indie dev notes&lt;/td&gt;
&lt;td&gt;intermediate&lt;/td&gt;
&lt;td&gt;EN 2000 / ZH 2500&lt;/td&gt;
&lt;td&gt;Already hit cross-layer bugs, wants to see how to lock them down&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;B1&lt;/td&gt;
&lt;td&gt;Offensive audit plain-language intro&lt;/td&gt;
&lt;td&gt;Indie dev notes&lt;/td&gt;
&lt;td&gt;intermediate&lt;/td&gt;
&lt;td&gt;EN 1600 / ZH 2100&lt;/td&gt;
&lt;td&gt;Wants to know how to audit a whole business flow in one pass&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;C1&lt;/td&gt;
&lt;td&gt;Offense + Defense combined&lt;/td&gt;
&lt;td&gt;Indie dev notes&lt;/td&gt;
&lt;td&gt;intermediate&lt;/td&gt;
&lt;td&gt;EN 2000 / ZH 2500&lt;/td&gt;
&lt;td&gt;Wants to see how the two pair up + where this doesn't apply&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A2&lt;/td&gt;
&lt;td&gt;Defensive engineering version&lt;/td&gt;
&lt;td&gt;Engineer diary&lt;/td&gt;
&lt;td&gt;advanced&lt;/td&gt;
&lt;td&gt;EN 2000 / ZH 2500&lt;/td&gt;
&lt;td&gt;Wants the 5-artifact implementation details&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;B2&lt;/td&gt;
&lt;td&gt;Offensive engineering version&lt;/td&gt;
&lt;td&gt;Engineer diary&lt;/td&gt;
&lt;td&gt;advanced&lt;/td&gt;
&lt;td&gt;EN 2000 / ZH 2500&lt;/td&gt;
&lt;td&gt;Wants the 6-piece fix pattern template&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;C2&lt;/td&gt;
&lt;td&gt;Cross-project reuse matrix&lt;/td&gt;
&lt;td&gt;Engineer diary&lt;/td&gt;
&lt;td&gt;advanced&lt;/td&gt;
&lt;td&gt;EN 2000 / ZH 2500&lt;/td&gt;
&lt;td&gt;Wants to port this to their own project, needs to know what transfers&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;D&lt;/td&gt;
&lt;td&gt;sql-only-trace&lt;/td&gt;
&lt;td&gt;Engineer diary&lt;/td&gt;
&lt;td&gt;advanced&lt;/td&gt;
&lt;td&gt;EN 2000 / ZH 2500&lt;/td&gt;
&lt;td&gt;How to test pure-DB logic&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h2&gt;
  
  
  Three reading paths
&lt;/h2&gt;

&lt;h3&gt;
  
  
  "I only have 30 minutes"
&lt;/h3&gt;

&lt;p&gt;Read &lt;strong&gt;E meta&lt;/strong&gt; (how the whole thing came together) plus this guide. That's enough.&lt;/p&gt;

&lt;p&gt;You don't need to look at the specific patterns. E meta compresses the whole journey into one post. After reading it you'll know "this kind of thing exists, here's the context it fits." That's fine.&lt;/p&gt;

&lt;h3&gt;
  
  
  "I'm not a software engineer but I'm curious about how AI pair programming gets governed"
&lt;/h3&gt;

&lt;p&gt;Order: &lt;strong&gt;E meta → A1 → B1 → C1&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Skip the 4 engineering posts. The 3 plain-language posts cover the full thinking, real time costs, and where the approach doesn't apply. After reading you'll have a sense of "is there an analogue to this in my own work." If you're a designer, product manager, or like me a self-taught coder, this path fits best.&lt;/p&gt;

&lt;h3&gt;
  
  
  "I'm an engineer, I want implementation details"
&lt;/h3&gt;

&lt;p&gt;Order: &lt;strong&gt;A1 (shared baseline) → A2 → B2 → C2 → D&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;A1 is required reading because the engineering posts assume you know what the 5 artifacts are. A2/B2/C2/D are where the code examples, governance rule templates, and pure-SQL test gotchas live.&lt;/p&gt;

&lt;p&gt;If you only want one technical angle:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;"How is the protection rule written" → A2&lt;/li&gt;
&lt;li&gt;"How does the cross-layer audit actually run" → B2&lt;/li&gt;
&lt;li&gt;"Can I port this to my project" → C2&lt;/li&gt;
&lt;li&gt;"How do I test pure-DB logic" → D&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Who this series is not for
&lt;/h2&gt;

&lt;p&gt;In C1 I listed 5 scenarios where trace lock doesn't apply. Extending that to "who can skip this whole series":&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Your project is pure read-only reporting / pure frontend without DB. The whole approach doesn't fit&lt;/li&gt;
&lt;li&gt;You don't pair with AI to write code. The series premise doesn't hold&lt;/li&gt;
&lt;li&gt;You're on a 5+ person team with strong code review culture. Defensive may overlap with human review, value drops&lt;/li&gt;
&lt;li&gt;You're looking for "best practices" or "industry standards." This series doesn't offer those. I name things myself, estimate time myself, list where it doesn't apply myself. No authority claims&lt;/li&gt;
&lt;li&gt;You want to read just one post and leave. Apart from this guide, the others assume context, so cherry-picking only works in the order I suggested above&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Shared vocabulary: the working names
&lt;/h2&gt;

&lt;p&gt;Names that recur across the series (none of them are industry-standard terminology):&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Name&lt;/th&gt;
&lt;th&gt;Meaning&lt;/th&gt;
&lt;th&gt;First appears&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Trace Lock&lt;/td&gt;
&lt;td&gt;Locking a cross-layer relationship with 5 artifacts&lt;/td&gt;
&lt;td&gt;A1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Offensive / Defensive&lt;/td&gt;
&lt;td&gt;The two audit modes&lt;/td&gt;
&lt;td&gt;A1 / B1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Dual-blade&lt;/td&gt;
&lt;td&gt;The closed loop where both modes pair up&lt;/td&gt;
&lt;td&gt;C1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Decision Pinning (business contract freeze)&lt;/td&gt;
&lt;td&gt;Writing the rule down explicitly before fixing a BLOCKER&lt;/td&gt;
&lt;td&gt;B1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;6-piece fix pattern&lt;/td&gt;
&lt;td&gt;The 6 steps for fixing one BLOCKER&lt;/td&gt;
&lt;td&gt;B1 / B2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;sql-only-trace&lt;/td&gt;
&lt;td&gt;The trace category for pure-DB logic&lt;/td&gt;
&lt;td&gt;D&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5 artifacts&lt;/td&gt;
&lt;td&gt;The 5 things every defensive trace must contain&lt;/td&gt;
&lt;td&gt;A1 / A2&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;If you hit an unfamiliar term in any post, come back here.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why I wrote these 9 posts
&lt;/h2&gt;

&lt;p&gt;Not because I think the patterns "need to be known by the industry." Because I tripped on the same shape of bug for half a year, the conversations with Claude when sorting it out got long, and I was afraid of forgetting. The blog version is partly for my own future re-reading, partly so that if a professional engineer sees me misusing something they can correct me.&lt;/p&gt;

&lt;p&gt;There's no "teaching" stance when I was writing. I noticed I'd written 9 posts only after they were done, so I laid them out and put them on the blog.&lt;/p&gt;




&lt;h2&gt;
  
  
  Related posts
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;E · Meta: how the methodology emerged from conversations with AI (series starting point)&lt;/li&gt;
&lt;li&gt;C1 · Offense + Defense combined (dual-blade overview, read this before deciding whether to go into A2/B2/C2/D)&lt;/li&gt;
&lt;li&gt;中文版&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  About this post
&lt;/h2&gt;

&lt;p&gt;This post is an organized record of conversations I had with Claude (an AI pair-programming tool)&lt;br&gt;
during May 2026. I noticed some patterns worth keeping for my own future reference,&lt;br&gt;
so I asked Claude to help structure them into writing.&lt;/p&gt;

&lt;p&gt;A few things I'm &lt;strong&gt;not&lt;/strong&gt; claiming:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Terms used in this post (Trace Lock / Offensive / Defensive / Dual-blade / Decision Pinning / 6-piece fix pattern / sql-only-trace / 5 artifacts) are working names I gave them myself, not industry-standard terminology&lt;/li&gt;
&lt;li&gt;My system has a specific shape (solo-maintained, many cross-layer dependencies, ambiguous business contracts). These patterns may not apply to your context&lt;/li&gt;
&lt;li&gt;I'm not a software engineer, just a barista who pairs with AI to write code&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If a professional engineer spots misuse, or there's already a more standard name for any of these concepts, &lt;strong&gt;I genuinely welcome corrections&lt;/strong&gt;.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;本文原載於我的部落格：&lt;a href="https://coffeeshooters.com/content/trace-lock-index-en?utm_source=devto&amp;amp;utm_medium=social&amp;amp;utm_campaign=blog-trace-lock-index-en" rel="noopener noreferrer"&gt;Series reading guide — How to pick which of these 9 Trace Lock posts to read&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>solodev</category>
    </item>
  </channel>
</rss>
