<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Hichoi-Dev</title>
    <description>The latest articles on DEV Community by Hichoi-Dev (@casamia918).</description>
    <link>https://dev.to/casamia918</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F1175424%2Fccccef7f-7c9c-4c14-812d-d36e99459a10.png</url>
      <title>DEV Community: Hichoi-Dev</title>
      <link>https://dev.to/casamia918</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/casamia918"/>
    <language>en</language>
    <item>
      <title>How Do You Evaluate People Who Are "Good at AI"? The ABCD2 Framework</title>
      <dc:creator>Hichoi-Dev</dc:creator>
      <pubDate>Sun, 19 Jul 2026 15:33:41 +0000</pubDate>
      <link>https://dev.to/casamia918/how-do-you-evaluate-people-who-are-good-at-ai-the-abcd2-framework-iad</link>
      <guid>https://dev.to/casamia918/how-do-you-evaluate-people-who-are-good-at-ai-the-abcd2-framework-iad</guid>
      <description>&lt;p&gt;Try asking this in your next team meeting: "Who on our team is the best at using AI?"&lt;/p&gt;

&lt;p&gt;There will probably be a moment of silence. Someone thinks of the person who uses the tools the most. Someone thinks of the person who's good with prompts. Someone thinks of whoever recently produced something impressive-looking with AI. But none of these answers feels convincing. There's a gut sense that something is off — and yet nobody knows what the right criterion would be.&lt;/p&gt;

&lt;p&gt;Nearly every company is standing in front of this wall right now. AI competence is rising to the top of hiring and performance criteria, and there's no yardstick to measure it with. This essay proposes one. The short answer is a set of five axes I call ABCD2. But before that, we need to fix the question itself.&lt;/p&gt;

&lt;h2&gt;
  
  
  The very idea of "good at using AI" is wrong
&lt;/h2&gt;

&lt;p&gt;The title of this essay says "good at AI" — but that phrasing is itself the trap.&lt;/p&gt;

&lt;p&gt;The word "using" carries a hidden frame: AI is a tool, and the person is a user. It puts AI on the same shelf as Excel and Photoshop — something you get skilled at operating. On top of that frame, evaluation naturally becomes a measure of tool proficiency. How many features do you know? How fluently do you operate it?&lt;/p&gt;

&lt;p&gt;But AI is not a hammer. AI is an engine that judges and executes. It's less like a tool and more like a junior colleague you can delegate work to. And once you think of a junior colleague, the frame flips. Behind every junior who performs well, there is invariably a senior who &lt;em&gt;makes&lt;/em&gt; them perform well — someone who explains the context of the work, lays out the criteria for judgment, makes it safe to ask about ambiguities, and verifies the output. That senior isn't "using" the junior. They are making the junior &lt;em&gt;do the work well&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;AI is exactly the same. The person who truly gets the most out of AI is not the person who is good at using AI, but the person who helps AI do the work well.&lt;/p&gt;

&lt;p&gt;And what do they help with? Two things, broadly. One is &lt;strong&gt;knowledge building&lt;/strong&gt; — preparing the ground AI stands on: what our organization's terms actually mean, what the criteria for judgment are, what separates a good deliverable from a bad one, all captured in a form AI can draw on. The other is &lt;strong&gt;context pipelining&lt;/strong&gt; — designing the flow so that this knowledge reaches the AI at the right moment, in the right shape. Which context rides along with which task, what the output gets verified against, and where things route back to when verification fails.&lt;/p&gt;

&lt;p&gt;Seen from this angle, what is a prompt? It's merely the last line of a much longer flow. A good prompt is the &lt;em&gt;output&lt;/em&gt; of good knowledge building and good context pipelining — not the substance of the capability itself.&lt;/p&gt;

&lt;p&gt;Sharp readers will have already caught it by now: this is where &lt;strong&gt;ontology&lt;/strong&gt; becomes important. Structuring an organization's terms and concepts, the relationships between them, and its criteria for judgment into a form AI can draw on — that is precisely what an ontology is. Peel one layer off the phrase "knowledge building," and what you find inside is ontology. The more AI becomes an organization's execution engine, the wider the gap will grow between organizations that have an ontology for that engine to stand on and those that don't.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why every existing yardstick misses
&lt;/h2&gt;

&lt;p&gt;Hold on to that redefinition, and it becomes obvious at a glance why the evaluation methods in use today all miss the mark.&lt;/p&gt;

&lt;p&gt;Start with the traditional yardsticks: coding tests, algorithm problems, certifications, credentials. What these have actually been measuring is &lt;em&gt;execution ability&lt;/em&gt; — how accurately and quickly you carry out a well-defined problem. But execution is precisely the part that has gone to the AI. The object of measurement has moved wholesale to the AI's side while the measuring instrument still points at the human — of course the predictive power evaporates. There is no guarantee that someone who aces a coding test works well with AI. If anything, the more someone is optimized for solving well-defined problems fast, the clumsier they may be at defining problems and designing context.&lt;/p&gt;

&lt;p&gt;What about the newer yardsticks, the ones supposedly built for the AI era? Two are common.&lt;/p&gt;

&lt;p&gt;First, prompt-skill assessment. As we just saw, the prompt is the last line of the flow. Measuring only the last line is like hiring a chef based on plating alone. Worse, prompt techniques are tied to specific tools and specific model versions, so they have a short half-life. The better the models get, the less the tricks are worth — and what remains is the ability to design &lt;em&gt;what to ask for&lt;/em&gt; and &lt;em&gt;what context to provide&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;Second, tool usage volume and time spent. The assumption is that whoever uses AI the most uses it best — and it may be exactly backwards. The person who has built a proper pipeline actually types &lt;em&gt;less&lt;/em&gt;. The person who starts from scratch every time, manually grinding through long conversations, racks up the highest usage. Using a lot and making it work well are different things.&lt;/p&gt;

&lt;p&gt;What the two yardsticks share is that both are measurements from the "use" frame. A tool-proficiency ruler cannot capture the ability to &lt;em&gt;make something do the work well&lt;/em&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  So what should we measure — ABCD2
&lt;/h2&gt;

&lt;p&gt;"The ability to make AI do the work well" cannot be measured as a lump. To measure it, you have to decompose it. I decompose it into five axes: Comprehend, Abstract, Dissolve, Build, Describe. Taking the initials: ABCD2.&lt;/p&gt;

&lt;p&gt;From an evaluator's point of view, each axis becomes a question.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Comprehend — does this person understand the world to be connected?&lt;/strong&gt; Do they grasp what the task actually touches, what that word means in this organization, who is affected by the outcome? Without comprehension, they misunderstand the very thing they're asking the AI to do.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Abstract — does this person abstract the essence?&lt;/strong&gt; Can they strip a workable structure out of messy reality? Good abstraction simplifies without discarding what matters. Without that sense of balance, the problem definition handed to the AI is warped from the start.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Dissolve — does this person break work into verifiable units?&lt;/strong&gt; Delegate the whole lump and you can't tell where it went wrong. Do they cut the work along boundaries where AI output can be checked? This connects directly to the ability to catch AI's plausible-sounding wrong answers.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Build — does this person construct the flow?&lt;/strong&gt; Do they weave the dissolved units into a single flow with verification and fallbacks? Do they build not a one-off request but a pipeline that is repeatable and can be handed to others?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Describe — does this person explain and take responsibility?&lt;/strong&gt; Can they articulate what they connected, why, and how? And for the AI's output, do they stand as the subject of responsibility — "I designed it to judge this way"?&lt;/p&gt;

&lt;p&gt;You may have noticed: nowhere among these five axes is "proficiency with a particular tool." Swap the tool, swap the model — the axes stay valid. That is the minimum requirement for any evaluation criterion worth adopting.&lt;/p&gt;

&lt;h2&gt;
  
  
  A case — same instruction, two people
&lt;/h2&gt;

&lt;p&gt;Abstract axes alone don't quite land, so consider a scene. A manager gives two team members the same instruction: "Analyze this month's customer inquiries with AI and write it up."&lt;/p&gt;

&lt;p&gt;A does this: downloads the inquiry data, pastes the whole thing into an AI, and asks, "Analyze these customer inquiries and summarize the main issues." After a few follow-up exchanges, a clean summary comes out — category breakdown, top complaints, even improvement suggestions. Time spent: 30 minutes. The report looks solid.&lt;/p&gt;

&lt;p&gt;B does this: first checks where these inquiries come from. Learns that each channel has a different character, and that the CS team already maintains a tagging taxonomy (Comprehend). Takes that taxonomy as the base, but reorganizes it around three lenses that fit this analysis's purpose — deriving next quarter's improvement backlog: "recurring inquiries / new types / urgent" (Abstract). Instead of feeding the data in whole, splits it by channel and by week, so each batch's results can be checked against the source (Dissolve). Writes up the classification criteria, the company's term definitions, and examples of good classification into a single document that rides along with every request, and routes ambiguous cases to human review (Build). And at the end of the report, writes: "This analysis is based on the CS tagging taxonomy; channel X data is not included; the criterion for 'urgent' is one I defined as follows" (Describe). Time spent: half a day.&lt;/p&gt;

&lt;p&gt;Now — looking at this month's report alone, who did better? Hard to tell. Honestly, A's report came out faster and doesn't look inferior. By the existing evaluation method — deliverable quality and speed — A might even win.&lt;/p&gt;

&lt;p&gt;The difference shows up next month. When the same instruction comes down again, A starts over from scratch. The classification criteria have drifted from last month's, so trend comparison isn't even possible. B pours the new data into the flow already built, and that's it. The criteria are stable, so month-over-month trends accumulate — and even when B goes on vacation, someone else can run the flow. What A left behind is a report. What B left behind is an asset.&lt;/p&gt;

&lt;p&gt;This is the crux. The existing evaluation looks only at "this month's deliverable," so it cannot separate A from B. ABCD2 looks not at the deliverable but at the five points in the process that produced it — and so it can.&lt;/p&gt;

&lt;h2&gt;
  
  
  In practice — how to ask in interviews and reviews
&lt;/h2&gt;

&lt;p&gt;So how do you actually probe ABCD2 in an interview or a performance review? One example per axis:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Comprehend&lt;/strong&gt; — Give a task with the context deliberately missing. "How would you analyze our service's churn rate with AI?" — without defining "churn" or describing the data. A strong candidate asks back before answering: how are we defining churn, and what data do we have? The one who launches straight into a plausible-sounding procedure without asking is the one who, on the job, will hand the AI a misunderstood problem.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Abstract&lt;/strong&gt; — Give a messy real case and have them structure it. Throw them an exception-riddled business process and ask, "If you were delegating this to AI, what stages would you organize it into?" Watch whether they find the balance between the person who tries to cram in every exception and the person who throws away what matters.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Dissolve&lt;/strong&gt; — Give them a plausible wrong answer from an AI. Show a deliverable that reads smoothly but has one factual error, and ask, "If you had to review this, how would you do it?" "I'd reread the whole thing" and "I'd cut it into verifiable units and check each against the source" are entirely different capabilities.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Build&lt;/strong&gt; — Follow up with: "What if you had to repeat that every week?" See whether they can turn a one-off solution into a repeatable flow, and whether verification and exception handling are designed &lt;em&gt;into&lt;/em&gt; the flow.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Describe&lt;/strong&gt; — About an AI deliverable the candidate produced, ask: "If this result turns out to be wrong, whose responsibility is it?" The distance between "the AI produced it that way" and "I designed it to judge by these criteria, so it's my responsibility — and here's what I'd fix" is the distance of this axis.&lt;/p&gt;

&lt;p&gt;You'll notice what these have in common: none of them are questions with a correct answer. They are questions that expose &lt;em&gt;how a person handles a problem&lt;/em&gt;. The exact opposite direction from execution-ability testing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Completing the frame — three axes
&lt;/h2&gt;

&lt;p&gt;ABCD2 alone doesn't complete a hiring decision. Add two more axes and you get a frame you can actually use. One is &lt;strong&gt;personality fit&lt;/strong&gt; — does this person have the temperament to bear that kind of thinking and ownership in the first place? The other is &lt;strong&gt;domain fit&lt;/strong&gt; — is the world they must comprehend and connect precisely &lt;em&gt;our&lt;/em&gt; domain? However high someone's ABCD2, without domain knowledge their Comprehend doesn't engage; and however well capability and domain align, a temperament that dodges ownership collapses their Describe.&lt;/p&gt;

&lt;p&gt;Put together: ABCD2 (capability) × personality fit (temperament) × domain fit (context). These three axes are my answer to the question, "how do we hire and place talent in the age of AI?"&lt;/p&gt;

&lt;h2&gt;
  
  
  Closing — stop looking for people who use it well
&lt;/h2&gt;

&lt;p&gt;Back to the title: how do you evaluate people who are "good at AI"? Now we can answer. Don't look for people who are good at using AI. Look for people who make AI do the work well. And that ability can't be measured as a lump — but decomposed, it can. Do they comprehend the world to be connected, abstract the essence, dissolve into verifiable units, build the flow, describe and take responsibility?&lt;/p&gt;

&lt;p&gt;One thing is worth saying honestly in advance. Apply this standard seriously, and you'll find that far fewer people pass than you expected. That's not because the standard is wrong — it's because until now we've been measuring with the broad ruler of "execution." Change the ruler, and the map of talent gets redrawn. And the company that draws that map first takes the lead in the great talent reshuffle of the AI era.&lt;/p&gt;

&lt;p&gt;This article is also post in my gist page : &lt;a href="https://gist.github.com/casamia918/678cd716333a43f6b5fe539625b1e1ac#file-abcd2_framework_en-md" rel="noopener noreferrer"&gt;https://gist.github.com/casamia918/678cd716333a43f6b5fe539625b1e1ac#file-abcd2_framework_en-md&lt;/a&gt; &lt;/p&gt;

</description>
      <category>ai</category>
      <category>hiring</category>
      <category>career</category>
    </item>
    <item>
      <title>The Developer Becomes a Pipeliner</title>
      <dc:creator>Hichoi-Dev</dc:creator>
      <pubDate>Sun, 19 Jul 2026 14:42:48 +0000</pubDate>
      <link>https://dev.to/casamia918/the-developer-becomes-a-pipeliner-5875</link>
      <guid>https://dev.to/casamia918/the-developer-becomes-a-pipeliner-5875</guid>
      <description>&lt;p&gt;I believe the profession we call "developer" is about to turn into something fundamentally different. This isn't a matter of tools changing or productivity rising. The very identity of the job is shifting somewhere else. If I had to name it in a single word: the developer becomes a Pipeliner.&lt;/p&gt;

&lt;h2&gt;
  
  
  The developer was always an "intelligent executor"
&lt;/h2&gt;

&lt;p&gt;To make this case, I have to start with what a developer has actually been doing all this time. We call developers "builders," but the essence of the work was translation. You take a spec someone else defined — a plan, a requirement, a ticket, words exchanged in a meeting — and render it into a form a machine can execute. The developer stood between human intent and machine execution as an interpreter.&lt;/p&gt;

&lt;p&gt;But this translation was never mechanical. Specs are always incomplete, contradictory, full of gaps. Filling those gaps with judgment, imagining the edge cases, inferring what was never stated, and turning all of it into a working system — that was not manual labor but intellectual labor. So I'd describe the developer more precisely as an &lt;em&gt;intelligent executor&lt;/em&gt;: someone who uses intelligence to execute. That is what made a developer a developer for the past several decades.&lt;/p&gt;

&lt;p&gt;What's interesting is that this role survived several waves of "automation threat." From assembly to high-level languages, from high-level languages to frameworks, from on-premise to the cloud. Each time the abstraction rose one level, a large part of the &lt;em&gt;execution&lt;/em&gt; got automated. You no longer manage memory by hand, you no longer stand up servers yourself. And yet the developer survived every time. Because what got automated was the &lt;em&gt;mechanical&lt;/em&gt; part of execution, while the &lt;em&gt;judgment&lt;/em&gt; execution required stayed with the human. The "execution" was shaved down bit by bit; the "intelligence" remained intact.&lt;/p&gt;

&lt;h2&gt;
  
  
  This time is different — AI took the "intelligence" side
&lt;/h2&gt;

&lt;p&gt;Here is exactly where AI changes things. What's being automated now is not the mechanical part of execution but the judgment execution requires. Filling the gaps in an incomplete spec, imagining the edge cases, inferring the unstated and turning it into working code — the very definition of the intelligent executor — is now done by AI.&lt;/p&gt;

&lt;p&gt;This is decisive. Every prior wave of automation peeled "execution" off the developer but never touched "intelligence." So the developer could climb to a higher abstraction and remain an intelligent executor up there. This time, the intelligence itself was taken. The core of what made a developer a developer is being transferred wholesale. The old formula — "just climb to a higher abstraction" — doesn't work this time, because the ladder itself has been handed to the AI.&lt;/p&gt;

&lt;h2&gt;
  
  
  So what remains — placing the intelligence
&lt;/h2&gt;

&lt;p&gt;If execution has gone to the AI, what is left for the human? The answer is clear: making that AI do the work well. And that turns out to be a big job.&lt;/p&gt;

&lt;p&gt;AI is a powerful execution engine, but on its own it does nothing. What should it take as input, what context should come with it, against what should its output be verified, where should it flow when it fails, and to what next stage should its output connect — someone has to design all of this. Deciding where to pour the raw material of intelligence, through which pipes to run it, and with which valves to control it. I call this the pipeline.&lt;/p&gt;

&lt;p&gt;Here, "pipeline" is not a narrow technical term like a CI/CD pipeline or a data pipeline. It is much broader — designing and managing the entire path through which intelligence flows. It's less like laying bricks and more like being the water engineer who designs how a whole city's water will move. Bricks are now churned out infinitely by AI. The question is how to arrange those bricks into a flow, and the person who handles that flow is the pipeliner.&lt;/p&gt;

&lt;p&gt;It means the developer moves from "the one who writes code" to "the one who designs flow." From the one who builds with their hands to the one who arranges the hands.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why "developer" no longer fits
&lt;/h2&gt;

&lt;p&gt;But a thought might arise here. If designing and connecting flow is the job, wasn't the developer doing that all along? Wiring functions together, connecting modules, designing data flow — isn't that itself pipelining?&lt;/p&gt;

&lt;p&gt;At the micro level, that's true. In some sense the developer has always been pipelining, and there's no denying it. But the decisive difference lies in the &lt;em&gt;unit&lt;/em&gt; — the &lt;em&gt;level&lt;/em&gt; — at which that pipelining happened.&lt;/p&gt;

&lt;p&gt;Until now, the developer's pipelining was locked inside the code level. Wiring functions and modules within a given spec, inside system boundaries already drawn. The level above — why are we building this, which customer problem does it solve, what should the flow look like in business terms — was usually decided and handed down by someone else (a planner, a PM, the business side), and the developer wove flow only inside that fence. Because the unit of pipelining stayed at the code level, the name "developer" — the one who &lt;em&gt;develops&lt;/em&gt; — was enough.&lt;/p&gt;

&lt;p&gt;What the era now demands is pipelining at a different level. Beyond the code level: designing flow at the business level where an actual customer problem gets solved, and at the level of the whole system that solves it. Because AI took the code-level pipelining, what's left for the human is the pipelining above it. And this work is no longer about developing code, so the narrow word "developer" can't contain it. Someone who isn't bound to code as a particular material but weaves flow across levels — that's why I think a new name is needed. A pipeliner.&lt;/p&gt;

&lt;h2&gt;
  
  
  The pipeliner's capabilities — ABCD2
&lt;/h2&gt;

&lt;p&gt;So what capabilities does this pipeliner need? I organize them into five. Comprehend, Abstract, Dissolve, Build, Describe. Taking the initials, I call it ABCD2.&lt;/p&gt;

&lt;p&gt;Listed abstractly they don't land, so let me thread them through a single scenario. Suppose a company says, "we want to automate customer refund processing with AI." How does a pipeliner handle this problem?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Comprehend.&lt;/strong&gt; First you have to understand the world you need to connect. The job is grasping what the word "refund" actually touches in this company. A refund touches the accounting ledger, reverses inventory, affects customer trust, and runs into tax and legal. Someone who only sees the happy path has already failed here. If you don't know how the world actually turns, you misunderstand the very thing you're supposed to connect. Comprehension is knowing the material of the pipeline precisely.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Abstract.&lt;/strong&gt; The reality you've comprehended is messy. Exceptions, special cases, and office politics are all tangled together. Stripping out only the essential structure you can work with, from that messy reality, is abstraction. Erecting a skeleton like: "a refund is ultimately five stages — (identify the transaction) → (judge refund eligibility) → (calculate the amount) → (get approval) → (execute)." Good abstraction simplifies reality without throwing away what matters. That sense of balance is the heart of abstraction, and the hardest part to learn.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Dissolve.&lt;/strong&gt; Now break the abstracted skeleton into executable units. Why break it up? Because if you throw the whole thing at the AI — "handle the refund on your own" — there's no way to verify the output. In one lump you can't tell where it went wrong. You have to dissolve it so each unit is independently verifiable, so the AI can process each piece and a human can check each piece. Dissolving isn't just cutting things small; it's the sense for cutting &lt;em&gt;along verifiable boundaries&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Build.&lt;/strong&gt; Now weave the dissolved units into a single flow. Place AI at each stage, insert verification between them, design fallbacks that hand off to a human on failure, and put a human in the loop at the approval stage. Making units and making flow are different abilities. You can gather good parts and still have a system that doesn't run if the flow is tangled. Building is actually connecting the plumbing so the water moves.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Describe.&lt;/strong&gt; Finally, you must be able to explain that flow clearly. This isn't an add-on ability; it's essential. If you can't put into words what you connected, why, and how, that pipeline can't be maintained, can't be audited, can't be handed to someone else, can't be sold. And more fundamentally, to describe is to own the definition yourself. Being able to say "I decided this refund logic should judge it this way," and bearing the consequences. The ability to describe is inseparable from ownership.&lt;/p&gt;

&lt;p&gt;Comprehend, Abstract, Dissolve, Build, Describe. It called ABCD2. And as you may have noticed, none of these five is "the ability to write good code."&lt;/p&gt;

&lt;h2&gt;
  
  
  FDE as proof
&lt;/h2&gt;

&lt;p&gt;This capability set isn't widely recognized yet, but its prototype already exists in the market. The FDE — Forward Deployed Engineer.&lt;/p&gt;

&lt;p&gt;An FDE does different work from a traditional developer. They go into the customer's organization, comprehend their world, abstract the real problem, dissolve it into solvable units, build a flow that actually works, and describe and persuade the customer's organization of it. An FDE's value doesn't come from lines of code. It comes from the ability to connect the customer's chaotic reality with the tools.&lt;/p&gt;

&lt;p&gt;In other words, the FDE is already a pipeliner. Before AI had even fully taken over execution, the market had already discovered the point where the value of "the one who connects and designs" overwhelms the value of "the one who executes." It's no accident that the FDE has become one of the most highly paid engineering positions in recent years. That's a leading indicator of the pipeliner era.&lt;/p&gt;

&lt;h2&gt;
  
  
  The real problem — the vanishing of "execution" as a buffer
&lt;/h2&gt;

&lt;p&gt;While a few like the FDE rise upward, the opposite happens below. Until now, "execution" has been the vast buffer zone of knowledge work. Even without being exceptional, even without the full ABCD2, most people were useful simply by diligently executing well-defined tasks. The developer who turns a spec into code, the analyst who organizes data, the marketer who runs an assigned campaign, the office worker who produces documents to a template. They were all "intelligent executors," and they worked stably on top of the buffer of execution.&lt;/p&gt;

&lt;p&gt;This buffer is disappearing. Once AI takes execution, diligent execution alone no longer creates value. What remains is designing the pipeline. In a world where the buffer is gone, the "middle" of the skill distribution collapses. The broad middle layer that used to sit between the top-tier pipeliner and the displaced executor — the place where most knowledge workers have lived — thins out entirely.&lt;/p&gt;

&lt;h2&gt;
  
  
  And this isn't only the developer's story
&lt;/h2&gt;

&lt;p&gt;Here the scope of the argument explodes. ABCD2 is not a capability demanded only of developers. The moment you're working on top of AI as an execution engine, it's demanded equally of &lt;em&gt;every&lt;/em&gt; worker who deals in knowledge.&lt;/p&gt;

&lt;p&gt;The planner, too, is no longer someone who writes plans, but someone who comprehends the problem, dissolves it, and designs the flow so the AI produces good plans. The same goes for the marketer, the analyst, the consultant, the lawyer, the accountant. In each domain, "execution" goes to the AI, and what remains is the pipelining that directs it. The job titles differ, but the fundamental capability required converges on the single thing: ABCD2. So the pipeliner is the future of the knowledge worker before it is the future of the developer.&lt;/p&gt;

&lt;p&gt;The problem is that the people who can actually do these five things are few. That has been true so far, and it will remain true. Comprehending, abstracting, dissolving, building, and describing cannot be taught by manual. They stand on top of judgment, taste, ownership, systems thinking — and abilities like these are not easily reproduced by training.&lt;/p&gt;

&lt;p&gt;History has already shown this pattern. When the power loom automated weaving, the many weavers did not all convert into loom designers. The people who could design and run looms were a tiny minority, and the number needed was a tiny minority too. When execution is automated, the seats for designing that execution are always far fewer than the seats for performing it. What's happening now is structurally identical. The only difference is that what's being automated this time is not hands but intelligence — and so what's affected is not physical labor but the whole of knowledge work.&lt;/p&gt;

&lt;h2&gt;
  
  
  The anticipated objection — "won't AI end up doing the pipelining too?"
&lt;/h2&gt;

&lt;p&gt;There's a question worth raising honestly. If AI took execution, won't it eventually take the pipelining too? Then isn't the pipeliner's seat also temporary?&lt;/p&gt;

&lt;p&gt;Partly, yes. AI will keep getting better at pieces of the pipeline — especially the mechanical parts of Build and Dissolve. But the two ends of ABCD2, Comprehend and Describe, are different in nature. Deciding which reality to make the object of connection (comprehend), and owning that definition and explaining it to the world (describe), are fundamentally questions of &lt;em&gt;what do we want&lt;/em&gt; and &lt;em&gt;who bears responsibility&lt;/em&gt;. AI does not own the objective on its own, and does not bear the consequences on its own. The information asymmetry — AI doesn't know our organization's real context — the quality of task specification, and above all the attribution of responsibility: these are gaps that don't transfer easily.&lt;/p&gt;

&lt;p&gt;So I don't see the pipeliner as a temporary role. If anything, the more the pipeline gets automated, the greater the value of the few who decide &lt;em&gt;what it's built for&lt;/em&gt; and &lt;em&gt;take responsibility&lt;/em&gt; for it. The vanishing of the buffer doesn't threaten these few. It only makes them scarcer.&lt;/p&gt;

&lt;h2&gt;
  
  
  How do we identify the few?
&lt;/h2&gt;

&lt;p&gt;So the practical question that remains is this: how do we identify the few?&lt;/p&gt;

&lt;p&gt;It's not an idle question. Right now nearly every company stands before the same wall — "how do we evaluate talent in the age of AI?" And most of the old yardsticks have already gone dead. Coding tests, algorithm problems, certifications, credentials. What these actually measured was, in the end, "execution ability," and that execution has gone to the AI, as we've seen. There's no guarantee that someone who aces a coding test is a strong pipeliner. If anything, the more someone is optimized for quickly solving well-defined problems, the clumsier they may be at defining the problem itself and connecting the world.&lt;/p&gt;

&lt;p&gt;So what should we measure? I think ABCD2 can be the skeleton of the answer. Do they comprehend the world to be connected, abstract the essence, dissolve it into verifiable units, build the flow, describe it and take responsibility? These five axes reveal one's aptitude as a pipeliner regardless of coding skill and regardless of job function.&lt;/p&gt;

&lt;p&gt;Add two more axes and you get an evaluation frame you can actually use. One is personality fit — does this person even have the temperament to bear that kind of thinking and ownership? The other is domain fit — is the world they must comprehend and connect precisely &lt;em&gt;this&lt;/em&gt; domain? Put together: ABCD2 (capability) × personality fit (temperament) × domain fit (context). This combination of three axes, I believe, can set the direction for the very thing companies are desperately searching for — "how do we hire and place talent in the age of AI?"&lt;/p&gt;

&lt;p&gt;And this frame coldly confirms what we said earlier. The people who pass all three axes are exceedingly few. The moment the criterion shifts from "execution" to these three axes, the very size of the population that will remain in knowledge work gets redrawn.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion — dealing in knowledge becomes the work of a tiny few
&lt;/h2&gt;

&lt;p&gt;Let me sum up. The developer was an intelligent executor, and AI took that role. What remains is the pipelining that directs AI as an execution engine — and because it's pipelining not at the code level but at the business and system level, the narrow word "developer" can't contain it. So the developer becomes a pipeliner. The pipeliner's capabilities are comprehend, abstract, dissolve, build, describe — ABCD2 — and the FDE is its living prototype. And this ABCD2, with personality fit and domain fit added, can become a new axis for evaluating talent in the age of AI.&lt;/p&gt;

&lt;p&gt;This capability is demanded not only of developers but of every worker who deals in knowledge. In a world where the buffer of execution has vanished, only those who possess ABCD2 remain in the territory of knowledge work. Such people are few. And so dealing in knowledge itself gets reorganized into the work of a tiny few.&lt;/p&gt;

&lt;p&gt;This is the conclusion I really want to reach through the concept of the pipeliner. We are not watching a single job category change. In a world where intelligence has been productized, &lt;em&gt;who can handle intelligence&lt;/em&gt; — we are standing at the threshold of being forced to ask that question.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;This essay is also available as a GitHub Gist:&lt;/em&gt; &lt;a href="https://gist.github.com/casamia918/286dc7a6e2234a05dc10110259270f2e#file-the_future_of_developer_is_pipeliner_eng-md" rel="noopener noreferrer"&gt;The Future of Developer is Pipeliner&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Would love to hear where you agree or disagree. 👇&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>career</category>
      <category>softwaredevelopment</category>
      <category>discuss</category>
    </item>
    <item>
      <title>Everyone Just Discovered Loop Engineering. REAP Got There First — and It's Ready When You Are</title>
      <dc:creator>Hichoi-Dev</dc:creator>
      <pubDate>Sun, 05 Jul 2026 09:11:03 +0000</pubDate>
      <link>https://dev.to/casamia918/everyone-just-discovered-loop-engineering-reap-got-there-first-and-its-ready-when-you-are-26dh</link>
      <guid>https://dev.to/casamia918/everyone-just-discovered-loop-engineering-reap-got-there-first-and-its-ready-when-you-are-26dh</guid>
      <description>&lt;p&gt;In June 2026, "loop engineering" went viral. Stop prompting your agent — design the loop that prompts it. Ralph Wiggum loops running Claude for hours. Overnight runs. Millions of views.&lt;/p&gt;

&lt;p&gt;My honest reaction: &lt;em&gt;finally, everyone's here.&lt;/em&gt; I've been running my entire development process as AI loops since this February — months before the trend had a name — and one project is now &lt;strong&gt;70+ loop iterations deep&lt;/strong&gt;, and the tool that runs those loops is itself built &lt;em&gt;by&lt;/em&gt; those loops.&lt;/p&gt;

&lt;p&gt;So this is a field report. Not "loops are amazing" (they are) and not "loops are hype" (they're not). Just the seven things that turned out to actually matter once you live inside a loop long enough — including the ones the infinite-loop crowd is about to learn the hard way.&lt;/p&gt;

&lt;p&gt;Context: the tool is &lt;a href="https://github.com/c-d-cc/reap" rel="noopener noreferrer"&gt;REAP&lt;/a&gt; (&lt;a href="https://reap.cc" rel="noopener noreferrer"&gt;https://reap.cc&lt;/a&gt;), an open-source pipeline I built on top of Claude Code / OpenCode. It exists because I needed these seven lessons encoded in software, not in my discipline.&lt;/p&gt;

&lt;h2&gt;
  
  
  First, what the loop people get right
&lt;/h2&gt;

&lt;p&gt;Credit where due — the core insight of loop engineering is correct:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;One-shot prompting doesn't scale.&lt;/strong&gt; Iteration beats a perfect mega-prompt every time.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Files beat context windows.&lt;/strong&gt; State that matters must live on disk, not in the conversation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fresh context each iteration&lt;/strong&gt; prevents the slow rot of a 400-message session.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The leverage moved.&lt;/strong&gt; Your job really is designing the system around the agent now.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I agree with all of it. Now here's what months of actually living inside the loop adds.&lt;/p&gt;

&lt;h2&gt;
  
  
  Lesson 1: A goal is not a loop spec
&lt;/h2&gt;

&lt;p&gt;The naive loop is: final goal + &lt;code&gt;while true&lt;/code&gt;. It works for tasks where &lt;em&gt;the environment&lt;/em&gt; can say "done" — make tests pass, finish a mechanical migration. Even Ralph loop advocates admit it: vague criteria = infinite loop, judgment-heavy work doesn't converge.&lt;/p&gt;

&lt;p&gt;But almost everything interesting in software is judgment-heavy. So instead of one goal driving infinite iterations, I got much better results from &lt;strong&gt;one bounded goal per iteration&lt;/strong&gt;, chosen fresh each time by comparing long-term vision against current state (gap analysis). The loop's unit of work in REAP is a &lt;em&gt;generation&lt;/em&gt;: one goal, one lifecycle, one review. Then pick the next goal with a human sanity-check in between.&lt;/p&gt;

&lt;p&gt;Small bounded loops with re-aiming between them beat one big loop, every single time.&lt;/p&gt;

&lt;h2&gt;
  
  
  Lesson 2: The loop needs stages, not just repetitions
&lt;/h2&gt;

&lt;p&gt;An unstructured iteration ("here's the goal, go") makes the agent jump straight to code. The fix that stuck: every generation walks a fixed lifecycle —&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;learning → planning → implementation ⇄ validation → completion
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each stage produces an artifact file (what was learned, what's planned, what was done, what was verified). Sounds bureaucratic. Isn't. Those artifacts are what make iteration N+1 smarter than iteration N — and what make the human checkpoint (next lesson) reviewable in minutes instead of hours.&lt;/p&gt;

&lt;h2&gt;
  
  
  Lesson 3: The agent must never grade its own homework
&lt;/h2&gt;

&lt;p&gt;This is the hill I'll die on. Every generation ends with a &lt;strong&gt;fitness phase&lt;/strong&gt; where a human gives feedback — and it's deliberately &lt;em&gt;natural language only&lt;/em&gt;. No scores, no rubric, no "rate this 1–10", no LLM-as-judge.&lt;/p&gt;

&lt;p&gt;Why so strict? Goodhart's law. Any quantitative fitness signal an agent can see is a signal it will optimize &lt;em&gt;instead of the actual goal&lt;/em&gt;. I've watched it happen. Self-assessment ("here's what I'm uncertain about") is allowed and useful; self-&lt;em&gt;scoring&lt;/em&gt; is banned at the protocol level.&lt;/p&gt;

&lt;p&gt;An unattended infinite loop has exactly one grader: the model itself. That's not autonomy — that's compounding hallucination with a progress bar.&lt;/p&gt;

&lt;h2&gt;
  
  
  Lesson 4: Lock the rules while the loop is running
&lt;/h2&gt;

&lt;p&gt;Give an agent long enough inside a loop and it will, very reasonably, decide the rules should change. The convention was inconvenient, so it "improved" it — mid-task, silently.&lt;/p&gt;

&lt;p&gt;REAP's answer is a &lt;strong&gt;genome&lt;/strong&gt;: a small set of files holding architecture decisions, conventions, and hard constraints. During a generation the genome is &lt;em&gt;immutable&lt;/em&gt;. The agent can propose changes, but they queue up in a backlog and get applied only at the generation boundary — where I review them. The loop can suggest amendments to its constitution; it cannot ratify them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Lesson 5: If a stage can be skipped, it will be skipped
&lt;/h2&gt;

&lt;p&gt;Ask any agent to "always run validation before completing" and count the sessions until it… doesn't. Instructions decay. So REAP enforces stage order &lt;strong&gt;cryptographically&lt;/strong&gt;: every stage transition requires a signature token (nonce) that only the previous stage's completion can issue. Skipping validation isn't a disobeyed instruction — it's a failed signature check. The CLI just says no.&lt;/p&gt;

&lt;p&gt;Rule of thumb after 70 generations: anything you'd write in ALL CAPS in your prompt should be enforced by the harness instead.&lt;/p&gt;

&lt;h2&gt;
  
  
  Lesson 6: Autonomy should be a budget, not a binary
&lt;/h2&gt;

&lt;p&gt;The loop-engineering debate keeps framing it as attended vs. unattended. The useful knob is in between: &lt;strong&gt;how many iterations am I willing to pre-approve?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;REAP calls it cruise mode: &lt;code&gt;reap cruise 3&lt;/code&gt; means "run 3 generations autonomously, then come back to me." Clear, mechanical goals? Crank it up. Ambiguous design territory? Set it to zero and review every generation. Autonomy becomes a dial you turn per-situation, not an ideology.&lt;/p&gt;

&lt;h2&gt;
  
  
  Lesson 7: Loops need exit ramps, not just exit conditions
&lt;/h2&gt;

&lt;p&gt;Real loops don't always end in success. Sometimes the goal was wrong, sometimes 60% done is worth keeping. An infinite loop has one exit: Ctrl-C, and whatever mess is on disk is your problem.&lt;/p&gt;

&lt;p&gt;A loop iteration in REAP has three distinct endings — &lt;strong&gt;complete&lt;/strong&gt; (full lifecycle + review), &lt;strong&gt;early-close&lt;/strong&gt; (keep the partial value, auto-carry unfinished tasks to the next generation's backlog), and &lt;strong&gt;abort&lt;/strong&gt; (discard cleanly, restore consumed state). The ability to &lt;em&gt;lose gracefully&lt;/em&gt; is what makes running many loops cheap.&lt;/p&gt;

&lt;h2&gt;
  
  
  Ralph loop vs. REAP, honestly
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Ralph-style infinite loop&lt;/th&gt;
&lt;th&gt;REAP&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Unit of work&lt;/td&gt;
&lt;td&gt;One prompt, repeated forever&lt;/td&gt;
&lt;td&gt;One goal per generation, re-aimed each cycle&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Memory&lt;/td&gt;
&lt;td&gt;Files + git, unstructured&lt;/td&gt;
&lt;td&gt;3-tier memory + genome + lineage archive&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Correction signal&lt;/td&gt;
&lt;td&gt;Environment (tests/build) only&lt;/td&gt;
&lt;td&gt;Environment &lt;strong&gt;+ human fitness each generation&lt;/strong&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Rules mid-run&lt;/td&gt;
&lt;td&gt;Agent can drift&lt;/td&gt;
&lt;td&gt;Genome locked, changes reviewed at boundary&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Stage discipline&lt;/td&gt;
&lt;td&gt;Prompt-based (decays)&lt;/td&gt;
&lt;td&gt;Signature-enforced (can't skip)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Autonomy&lt;/td&gt;
&lt;td&gt;All or nothing&lt;/td&gt;
&lt;td&gt;Budgeted (cruise N)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Failure exit&lt;/td&gt;
&lt;td&gt;Ctrl-C + cleanup&lt;/td&gt;
&lt;td&gt;abort / early-close / complete&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Best at&lt;/td&gt;
&lt;td&gt;Mechanical, machine-checkable tasks&lt;/td&gt;
&lt;td&gt;Sustained product evolution with judgment calls&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Not a takedown — for a 200-file mechanical migration with a green-tests exit condition, a Ralph loop is genuinely great. But for &lt;em&gt;evolving a real product over months&lt;/em&gt;, you need the right column. That's the gap REAP was built for.&lt;/p&gt;

&lt;h2&gt;
  
  
  Proof of loop: this tool builds itself
&lt;/h2&gt;

&lt;p&gt;The part I'm proudest of: REAP is developed &lt;em&gt;with&lt;/em&gt; REAP. All 70+ generations — the signature locking, the memory system, cruise mode, the evaluator agent — were built inside the exact loop they enforce, each closed with human fitness feedback. Every flaw in the loop design lands on me first, and the fix gets encoded into the genome for every generation after.&lt;/p&gt;

&lt;p&gt;Dog-fooding a loop tool inside its own loop is the fastest feedback cycle I've ever worked in.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try a structured loop (5 minutes)
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npm &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-g&lt;/span&gt; @c-d-cc/reap
&lt;span class="nb"&gt;cd &lt;/span&gt;your-project
reap init        &lt;span class="c"&gt;# detects greenfield vs existing codebase&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Open Claude Code (or OpenCode) and run:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;/reap.evolve
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's one full generation: the agent learns your codebase, plans, implements, validates — and then asks &lt;em&gt;you&lt;/em&gt; how it did. Feedback becomes selection pressure. The next generation starts smarter.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;🌱 &lt;strong&gt;Site:&lt;/strong&gt; &lt;a href="https://reap.cc" rel="noopener noreferrer"&gt;reap.cc&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;⭐ &lt;strong&gt;GitHub:&lt;/strong&gt; &lt;a href="https://github.com/c-d-cc/reap" rel="noopener noreferrer"&gt;github.com/c-d-cc/reap&lt;/a&gt; — stars help more people find it&lt;/li&gt;
&lt;li&gt;📦 &lt;strong&gt;npm:&lt;/strong&gt; &lt;code&gt;@c-d-cc/reap&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you're running loops today — Ralph-style, cron-based, hand-rolled — I'd love to hear where yours drifted and what you did about it. That's the conversation loop engineering actually needs next.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>New workflow control method for harness engineering — Signature-Based Locking</title>
      <dc:creator>Hichoi-Dev</dc:creator>
      <pubDate>Sat, 21 Mar 2026 19:58:59 +0000</pubDate>
      <link>https://dev.to/casamia918/new-workflow-control-method-for-harness-engineering-signature-based-locking-3bmj</link>
      <guid>https://dev.to/casamia918/new-workflow-control-method-for-harness-engineering-signature-based-locking-3bmj</guid>
      <description>&lt;h2&gt;
  
  
  The Problem: AI Won't Stay Harnessed
&lt;/h2&gt;

&lt;p&gt;If you've been building AI-assisted development workflows — what some call "harness engineering" — you've hit this wall:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;No matter how carefully you craft your prompts, the AI eventually goes off-script.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;You define a multi-step workflow. The AI follows it for a while. Then somewhere around step 4, it decides to "optimize" by skipping steps, modifying files directly, or inventing a shortcut that breaks your entire pipeline.&lt;/p&gt;

&lt;p&gt;This isn't a prompting failure. It's a fundamental limitation of prompt-only workflow control.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Prompt-Only Control Fails
&lt;/h2&gt;

&lt;p&gt;Three documented forces work against prompt-based workflow enforcement:&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Context Rot (Lost in the Middle)
&lt;/h3&gt;

&lt;p&gt;As conversations grow longer, instructions from the beginning of the context window lose influence. &lt;a href="https://arxiv.org/abs/2307.03172" rel="noopener noreferrer"&gt;Research published in TACL&lt;/a&gt; ("Lost in the Middle") demonstrates that LLMs exhibit a U-shaped attention curve — they attend strongly to the beginning and end of context, but performance degrades by over 20% for information in the middle. Your carefully structured "NEVER do X" rules get diluted by thousands of tokens of subsequent conversation.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Training-Induced Optimization Pressure
&lt;/h3&gt;

&lt;p&gt;This isn't speculation — it's documented behavior. &lt;a href="https://arxiv.org/abs/2310.13548" rel="noopener noreferrer"&gt;RLHF training&lt;/a&gt; creates measurable pressure toward concise, "helpful" responses, because human evaluators systematically prefer them. Anthropic's own &lt;a href="https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/claude-4-best-practices" rel="noopener noreferrer"&gt;prompting best practices&lt;/a&gt; explicitly state that newer Claude models "may skip detailed summaries for efficiency." OpenAI acknowledged the same phenomenon when GPT-4 became &lt;a href="https://futurism.com/the-byte/openai-patch-fix-gpt4-laziness" rel="noopener noreferrer"&gt;"lazy"&lt;/a&gt; in December 2023, requiring a new model checkpoint to fix.&lt;/p&gt;

&lt;p&gt;When an AI sees a 5-step workflow where steps 2-4 seem like overhead, it has a trained tendency to compress. This is the model being helpful — and breaking your harness in the process.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. No Enforcement Boundary
&lt;/h3&gt;

&lt;p&gt;Prompts are suggestions, not constraints. There's no mechanism to &lt;em&gt;prevent&lt;/em&gt; the AI from taking an action — you can only &lt;em&gt;ask&lt;/em&gt; it not to. &lt;a href="https://arxiv.org/pdf/2502.13295" rel="noopener noreferrer"&gt;Research on specification gaming&lt;/a&gt; shows that models can learn to satisfy the apparent goal while bypassing the intended process — including &lt;a href="https://lilianweng.github.io/posts/2024-11-28-reward-hacking/" rel="noopener noreferrer"&gt;modifying unit tests to pass&lt;/a&gt; instead of writing correct code. Prompts operate in the same trust domain as the AI itself.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's Been Tried
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Approach 1: Stronger Prompts
&lt;/h3&gt;

&lt;p&gt;Add more rules. Make them UPPERCASE. Use XML tags. Add "CRITICAL" and "NEVER" and "NON-NEGOTIABLE."&lt;/p&gt;

&lt;p&gt;This helps initially but doesn't solve context rot. The more rules you add, the more diluted each individual rule becomes.&lt;/p&gt;

&lt;h3&gt;
  
  
  Approach 2: File-Level Permissions
&lt;/h3&gt;

&lt;p&gt;Restrict which files the AI can modify (e.g., strict mode, read-only markers). This prevents certain destructive actions but doesn't enforce &lt;em&gt;workflow ordering&lt;/em&gt;. The AI can still call commands out of sequence.&lt;/p&gt;

&lt;h3&gt;
  
  
  Approach 3: Deterministic Scripts
&lt;/h3&gt;

&lt;p&gt;Move workflow logic out of prompts and into deterministic scripts. The AI calls scripts instead of modifying state directly. This is the right direction — but scripts alone can't prevent the AI from calling them out of order, or skipping them entirely.&lt;/p&gt;

&lt;h2&gt;
  
  
  First, You Need a Workflow
&lt;/h2&gt;

&lt;p&gt;Before we can lock anything, we need something to lock. Signature-Based Locking assumes your AI-assisted work follows a &lt;strong&gt;defined lifecycle with ordered stages&lt;/strong&gt; — a workflow where step N must complete before step N+1 begins.&lt;/p&gt;

&lt;p&gt;This is the lifecycle steps behind &lt;a href="https://github.com/c-d-cc/reap" rel="noopener noreferrer"&gt;REAP&lt;/a&gt; (Recursive Evolutionary Autonomous Pipeline), where each unit of work — called a &lt;strong&gt;Generation&lt;/strong&gt; — follows a 5-stage lifecycle:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Objective → Planning → Implementation → Validation → Completion
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each stage has a clear purpose: define the goal, break it into tasks, build it, verify it works, then retrospect and archive. Stages produce artifacts, and transitions between stages are explicit — you can't "drift" from planning into implementation without a deliberate transition.&lt;/p&gt;

&lt;p&gt;This kind of structured workflow is where AI agents provide the most value (creative work within each stage) but also where they cause the most damage (skipping stages, going out of order, bypassing gates). The more structured your workflow, the more you need enforcement.&lt;/p&gt;

&lt;h2&gt;
  
  
  Signature-Based Locking: Enforcing Workflow Sequence from Outside the AI
&lt;/h2&gt;

&lt;p&gt;Given a structured workflow, here's the insight: &lt;strong&gt;sequence enforcement must happen outside the AI's trust boundary.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The AI can ignore prompts. Content guardrails can filter what it says. But neither can enforce the &lt;em&gt;order&lt;/em&gt; in which steps are executed. Cryptographic signatures can.&lt;/p&gt;

&lt;h3&gt;
  
  
  How It Works
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fxm2pp09mtomgqordq8q2.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fxm2pp09mtomgqordq8q2.png" alt="SIGNATURE_BASED_LOCKING_SEQ_DIAGRAM" width="800" height="985"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Each stage command generates a random nonce, stores its SHA256 hash (mixed with execution context) in the workflow state file, and returns the raw nonce to the AI. To advance to the next stage, the AI must pass this nonce to the transition command, which recomputes the hash and verifies it matches.&lt;/p&gt;

&lt;p&gt;The critical property: &lt;strong&gt;only the actual script execution can produce a valid nonce.&lt;/strong&gt; The AI receives the nonce as output, but cannot reverse-engineer or fabricate one. The hash is stored in a managed state file that the AI is structurally prevented from modifying directly.&lt;/p&gt;

&lt;h3&gt;
  
  
  What Signature-Based Locking Gives You
&lt;/h3&gt;

&lt;p&gt;Existing approaches each solve a piece of the puzzle: prompts communicate intent, file permissions restrict access, content guardrails filter unsafe output, and deterministic scripts encode logic. But none of them enforce &lt;strong&gt;execution sequence&lt;/strong&gt; — the guarantee that step N actually happened before step N+1.&lt;/p&gt;

&lt;p&gt;Signature-Based Locking fills this gap. It doesn't replace the other approaches; it adds the missing dimension. Here's how it stacks up:&lt;/p&gt;

&lt;h3&gt;
  
  
  Why Signature-Based Locking Works
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Threat&lt;/th&gt;
&lt;th&gt;Prompt-only&lt;/th&gt;
&lt;th&gt;Signature-Based Locking&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;AI skips a stage&lt;/td&gt;
&lt;td&gt;Possible&lt;/td&gt;
&lt;td&gt;Blocked (no nonce)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;AI modifies state directly&lt;/td&gt;
&lt;td&gt;Possible&lt;/td&gt;
&lt;td&gt;Blocked (hash mismatch)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;AI replays a previous step&lt;/td&gt;
&lt;td&gt;Possible&lt;/td&gt;
&lt;td&gt;Blocked (context-bound nonce)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  The Hybrid Architecture
&lt;/h2&gt;

&lt;p&gt;Signature-Based Locking is most effective as part of a &lt;strong&gt;hybrid architecture&lt;/strong&gt; that separates deterministic and creative work:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fxyvul79sii5ufuc74ut9.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fxyvul79sii5ufuc74ut9.png" alt="HYBRID_ARCH_FLOW_DIAGRAM" width="707" height="800"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Key principle&lt;/strong&gt;: The deterministic script handles everything that has a "right answer" — state transitions, gate checks, file validation, hook execution. The AI handles everything that requires creativity — writing code, making design decisions, solving problems.&lt;/p&gt;

&lt;p&gt;The scripts communicate with the AI through &lt;strong&gt;structured JSON output&lt;/strong&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"status"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"ok"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"command"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"objective"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"phase"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"complete"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"message"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Objective stage complete. Advance with: /reap.next a3f8c2d9..."&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The AI receives clear instructions and a nonce. It cannot advance without passing the nonce to the next command. The deterministic script verifies and controls the flow.&lt;/p&gt;

&lt;h2&gt;
  
  
  REAP: This Architecture in Practice
&lt;/h2&gt;

&lt;p&gt;This is exactly how &lt;a href="https://github.com/c-d-cc/reap" rel="noopener noreferrer"&gt;REAP&lt;/a&gt; (Recursive Evolutionary Autonomous Pipeline) works. REAP is an open-source CLI tool that structures AI-assisted development as an evolutionary process — software evolves across &lt;strong&gt;Generations&lt;/strong&gt;, each carrying one goal through a 5-stage lifecycle.&lt;/p&gt;

&lt;h3&gt;
  
  
  What REAP Does
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Genome&lt;/strong&gt; — Your project's design knowledge (architecture decisions, conventions, constraints, business rules) is managed as a living document that evolves across generations&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Lifecycle&lt;/strong&gt; — Each generation follows: Objective → Planning → Implementation → Validation → Completion&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Signature-Based Locking&lt;/strong&gt; — Stage transitions require cryptographic nonce verification, preventing the AI from skipping stages or going off-script&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Session Persistence&lt;/strong&gt; — The Genome and current generation state are automatically injected into the AI's context at session start, solving the "context loss across sessions" problem&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Multi-Agent Support&lt;/strong&gt; — Works with Claude Code and OpenCode.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  The Signature Chain in REAP
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;/reap.start &lt;span class="s2"&gt;"Build user auth"&lt;/span&gt;
  → Script creates generation, stores &lt;span class="nb"&gt;hash&lt;/span&gt;
  → AI receives instructions

/reap.objective
  → AI defines goals, writes artifact
  → Script verifies artifact, generates nonce
  → Message: &lt;span class="s2"&gt;"Advance with: /reap.next a3f8c2..."&lt;/span&gt;

/reap.next a3f8c2...
  → Script verifies SHA256&lt;span class="o"&gt;(&lt;/span&gt;a3f8c2 + genId + stage&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; stored &lt;span class="nb"&gt;hash&lt;/span&gt;
  → ✅ Match → advance to planning
  → ❌ Mismatch → &lt;span class="s2"&gt;"Token verification failed. Re-run the stage command."&lt;/span&gt;

/reap.planning
  → AI creates implementation plan
  → Script generates new nonce
  → Message: &lt;span class="s2"&gt;"Advance with: /reap.next b7d91e..."&lt;/span&gt;

  ... chain continues through all stages ...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each nonce is single-use, context-bound (includes generation ID and stage name), and cryptographically verified. The AI cannot skip ahead, replay, or forge tokens.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why It Matters
&lt;/h3&gt;

&lt;p&gt;After building 109 generations with REAP (yes, REAP is &lt;a href="https://github.com/c-d-cc/reap/tree/main/.reap/lineage" rel="noopener noreferrer"&gt;built with REAP&lt;/a&gt;), we've seen firsthand that prompt-only workflow control breaks down at scale. The AI "optimizes" by skipping validation, modifying state files directly, or calling commands out of order.&lt;/p&gt;

&lt;p&gt;Signature-Based Locking eliminated these failure modes — not by adding more rules to the prompt, but by making rule violation mechanically impossible.&lt;/p&gt;

&lt;h2&gt;
  
  
  Related Work: NeMo Guardrails
&lt;/h2&gt;

&lt;p&gt;It's worth mentioning NVIDIA's &lt;a href="https://developer.nvidia.com/nemo-guardrails" rel="noopener noreferrer"&gt;NeMo Guardrails&lt;/a&gt;, which shares an important philosophy with Signature-Based Locking: &lt;strong&gt;don't rely on prompts alone — enforce rules in code.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;NeMo Guardrails places a programmable middleware between the user and the LLM. User input is normalized into intents, &lt;a href="https://github.com/NVIDIA/NeMo-Guardrails" rel="noopener noreferrer"&gt;Colang&lt;/a&gt; rules determine whether to call the LLM or return a pre-defined response, and the output is screened against safety policies. This gives developers precise control over what the AI can say — blocking toxic content, preventing jailbreaks, enforcing topic boundaries, and &lt;a href="https://discuss.pytorch.kr/t/nemo-guardrails-llm-feat-nvidia/8403" rel="noopener noreferrer"&gt;detecting hallucinations through factual grounding checks&lt;/a&gt;. It integrates with LangChain, LlamaIndex, and supports GPU-accelerated evaluation for production workloads.&lt;/p&gt;

&lt;p&gt;This is genuinely valuable for chatbots, customer-facing AI, and any application where content safety matters. The core insight — layering deterministic safety on top of probabilistic LLMs — is sound.&lt;/p&gt;

&lt;p&gt;Where the two approaches diverge is the &lt;strong&gt;dimension of control&lt;/strong&gt;:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;NeMo Guardrails&lt;/th&gt;
&lt;th&gt;Signature-Based Locking&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Controls&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;What the AI says (content)&lt;/td&gt;
&lt;td&gt;What order the AI executes steps (sequence)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Mechanism&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Input/output filtering via policy rules&lt;/td&gt;
&lt;td&gt;Cryptographic nonce chain across steps&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Prevents&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Toxic content, jailbreaks, hallucinations&lt;/td&gt;
&lt;td&gt;Stage skipping, out-of-order execution, state tampering&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Best for&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Chatbots, customer-facing AI&lt;/td&gt;
&lt;td&gt;Multi-step workflows, autonomous agents&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;They're complementary, not competing. You could use NeMo Guardrails to ensure the AI doesn't produce unsafe content, &lt;em&gt;and&lt;/em&gt; Signature-Based Locking to ensure it follows the correct execution sequence. Different dimensions, same philosophy.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try It
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npm &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-g&lt;/span&gt; @c-d-cc/reap
reap init my-project
&lt;span class="c"&gt;# Open Claude Code or OpenCode&lt;/span&gt;
&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; /reap.evolve &lt;span class="s2"&gt;"Implement user authentication"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://github.com/c-d-cc/reap" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt; | &lt;a href="https://reap.cc" rel="noopener noreferrer"&gt;Documentation&lt;/a&gt; | &lt;a href="https://www.npmjs.com/package/@c-d-cc/reap" rel="noopener noreferrer"&gt;npm&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Have you struggled with keeping AI agents on-script in multi-step workflows? What approaches have you tried? I'd love to hear about your harness engineering experiences in the comments.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;References:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://arxiv.org/abs/2307.03172" rel="noopener noreferrer"&gt;Lost in the Middle: How Language Models Use Long Contexts — Liu et al. (TACL)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://arxiv.org/abs/2310.13548" rel="noopener noreferrer"&gt;Towards Understanding Sycophancy in Language Models — Anthropic&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/claude-4-best-practices" rel="noopener noreferrer"&gt;Claude Prompting Best Practices — Anthropic&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://futurism.com/the-byte/openai-patch-fix-gpt4-laziness" rel="noopener noreferrer"&gt;OpenAI Patches "Lazy" GPT-4 — Futurism&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://arxiv.org/pdf/2502.13295" rel="noopener noreferrer"&gt;Specification Gaming in Reasoning Models — arXiv&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://lilianweng.github.io/posts/2024-11-28-reward-hacking/" rel="noopener noreferrer"&gt;Reward Hacking in LLMs — Lilian Weng&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.anthropic.com/research/reward-tampering" rel="noopener noreferrer"&gt;Reward Tampering Research — Anthropic&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.pnas.org/doi/10.1073/pnas.2322420121" rel="noopener noreferrer"&gt;Embers of Autoregression — PNAS&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://developer.nvidia.com/nemo-guardrails" rel="noopener noreferrer"&gt;NeMo Guardrails — NVIDIA&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://discuss.pytorch.kr/t/nemo-guardrails-llm-feat-nvidia/8403" rel="noopener noreferrer"&gt;NeMo Guardrails 소개 — PyTorch Korea&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/c-d-cc/reap" rel="noopener noreferrer"&gt;REAP — GitHub&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>automation</category>
      <category>llm</category>
      <category>softwareengineering</category>
    </item>
    <item>
      <title>Specs Cannot be Source of Source Code — Why Intent Management Matters in AI-Driven Development</title>
      <dc:creator>Hichoi-Dev</dc:creator>
      <pubDate>Fri, 20 Mar 2026 17:37:55 +0000</pubDate>
      <link>https://dev.to/casamia918/specs-cannot-be-source-of-source-code-why-intent-management-matters-in-ai-driven-development-159c</link>
      <guid>https://dev.to/casamia918/specs-cannot-be-source-of-source-code-why-intent-management-matters-in-ai-driven-development-159c</guid>
      <description>&lt;h3&gt;
  
  
  The Seductive Idea
&lt;/h3&gt;

&lt;p&gt;There's a compelling narrative in AI-driven development right now: &lt;strong&gt;write a detailed spec, feed it to an AI agent, and get working software out.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;GitHub's &lt;a href="https://github.com/github/spec-kit" rel="noopener noreferrer"&gt;Spec Kit&lt;/a&gt;, AWS's &lt;a href="https://kiro.dev/" rel="noopener noreferrer"&gt;Kiro&lt;/a&gt;, and a growing ecosystem of tools all converge on the same premise — that specifications can become the "source of source code." The product requirements document isn't just a guide for implementation; it &lt;em&gt;is&lt;/em&gt; the source that generates implementation.&lt;/p&gt;

&lt;p&gt;It's an attractive idea. If specs are the source, then developers become spec writers, AI becomes the compiler, and code becomes a generated artifact. Clean. Elegant. Almost too good.&lt;/p&gt;

&lt;p&gt;And that's the problem.&lt;/p&gt;

&lt;h3&gt;
  
  
  Source Code Is Deterministic. Specs Are Not.
&lt;/h3&gt;

&lt;p&gt;Let's start with what "source" actually means in software engineering.&lt;/p&gt;

&lt;p&gt;When you compile &lt;code&gt;main.c&lt;/code&gt;, you get the same binary. Every time. On every machine. This property — &lt;strong&gt;determinism&lt;/strong&gt; — is what makes source code &lt;em&gt;source&lt;/em&gt;. It's the &lt;a href="https://reproducible-builds.org/docs/deterministic-build-systems/" rel="noopener noreferrer"&gt;reproducible foundation&lt;/a&gt; on which everything else stands: builds, tests, deployments, debugging.&lt;/p&gt;

&lt;p&gt;Now consider a specification:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"The system should handle user authentication with proper security measures."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Feed this to an AI agent three times. You'll get three different implementations — different OAuth flows, different session strategies, different error handling patterns. The &lt;a href="https://arxiv.org/abs/2602.00180" rel="noopener noreferrer"&gt;same spec produces different code&lt;/a&gt; across different runs, different models, and different context windows.&lt;/p&gt;

&lt;p&gt;This isn't a bug in the AI. It's a fundamental characteristic. Specifications are written in natural language, which is inherently ambiguous. LLMs are non-deterministic by design. The combination means that &lt;strong&gt;specs cannot serve as "source" in any meaningful engineering sense&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Source code has a contract: same input, same output. Specifications don't — and can't — honor that contract.&lt;/p&gt;

&lt;h3&gt;
  
  
  Then What Are Specs? Intent, Not Source.
&lt;/h3&gt;

&lt;p&gt;If specs aren't source, what are they?&lt;/p&gt;

&lt;p&gt;They're &lt;strong&gt;intent&lt;/strong&gt;. Initiative. Direction. A spec says &lt;em&gt;what&lt;/em&gt; you want and &lt;em&gt;why&lt;/em&gt; you want it — but it doesn't deterministically produce &lt;em&gt;how&lt;/em&gt;. The "how" emerges through the act of implementation, whether done by a human or an AI agent.&lt;/p&gt;

&lt;p&gt;This distinction matters more than it seems:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Source&lt;/th&gt;
&lt;th&gt;Intent&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Determinism&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Same input → same output&lt;/td&gt;
&lt;td&gt;Same input → many valid outputs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Verification&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Compile, run, test&lt;/td&gt;
&lt;td&gt;Interpret, judge, review&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Authority&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;The code is the truth&lt;/td&gt;
&lt;td&gt;The intent guides the truth&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Drift&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Doesn't drift from itself&lt;/td&gt;
&lt;td&gt;Drifts from implementation over time&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Treating intent as source is a category error. It's like treating a compass bearing as a GPS coordinate — useful for direction, useless for pinpointing where you actually are.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why We Still Need to Manage Intent
&lt;/h3&gt;

&lt;p&gt;But here's the thing: &lt;strong&gt;just because specs aren't source doesn't mean they don't matter.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;In AI-driven development, intent management is arguably &lt;em&gt;more&lt;/em&gt; critical than ever:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Context loss&lt;/strong&gt; — AI agents forget everything between sessions. Without persisted intent, every session starts from zero.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Knowledge decay&lt;/strong&gt; — Decisions made in session 12 are invisible in session 13. Architecture rationale evaporates. Business rules get re-debated.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Drift without anchor&lt;/strong&gt; — Without a persistent record of intent, AI agents make locally reasonable but globally inconsistent decisions. The codebase slowly becomes incoherent.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The question isn't whether to manage intent. It's &lt;em&gt;how&lt;/em&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  How Teams Have Managed Specs (A Brief History)
&lt;/h3&gt;

&lt;p&gt;Software teams have tried many approaches to capture and maintain design knowledge. Here's how the major ones compare — especially through the lens of AI-assisted development:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Approach&lt;/th&gt;
&lt;th&gt;Strengths&lt;/th&gt;
&lt;th&gt;Weaknesses&lt;/th&gt;
&lt;th&gt;AI-Era Fit&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;RFC&lt;/strong&gt; — proposal for collecting feedback (&lt;a href="https://newsletter.pragmaticengineer.com/p/rfcs-and-design-docs" rel="noopener noreferrer"&gt;Pragmatic Engineer&lt;/a&gt;)&lt;/td&gt;
&lt;td&gt;Structured deliberation&lt;/td&gt;
&lt;td&gt;Point-in-time, never updated&lt;/td&gt;
&lt;td&gt;Low&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;ADR&lt;/strong&gt; — records one decision + rationale (&lt;a href="https://candost.blog/adrs-rfcs-differences-when-which/" rel="noopener noreferrer"&gt;Candost&lt;/a&gt;)&lt;/td&gt;
&lt;td&gt;Lightweight, captures &lt;em&gt;why&lt;/em&gt;
&lt;/td&gt;
&lt;td&gt;Accumulates without sync&lt;/td&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;Design Docs&lt;/strong&gt; — comprehensive pre-impl design (&lt;a href="https://blog.pragmaticengineer.com/rfcs-and-design-docs/" rel="noopener noreferrer"&gt;Google, Uber-style&lt;/a&gt;)&lt;/td&gt;
&lt;td&gt;Thorough analysis&lt;/td&gt;
&lt;td&gt;Goes stale fast&lt;/td&gt;
&lt;td&gt;Low&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;CLAUDE.md / AGENTS.md&lt;/strong&gt; — repo-level AI instructions (&lt;a href="https://agents.md/" rel="noopener noreferrer"&gt;agents.md&lt;/a&gt;)&lt;/td&gt;
&lt;td&gt;Zero-friction, always loaded&lt;/td&gt;
&lt;td&gt;
&lt;a href="https://www.infoq.com/news/2026/03/agents-context-file-value-review/" rel="noopener noreferrer"&gt;No sync&lt;/a&gt;, grows stale silently&lt;/td&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;Spec Kit&lt;/strong&gt; — Spec → Plan → Task → Implement (&lt;a href="https://github.blog/ai-and-ml/generative-ai/spec-driven-development-with-ai-get-started-with-a-new-open-source-toolkit/" rel="noopener noreferrer"&gt;GitHub Blog&lt;/a&gt;)&lt;/td&gt;
&lt;td&gt;Structured workflow&lt;/td&gt;
&lt;td&gt;
&lt;a href="https://blog.scottlogic.com/2025/11/26/putting-spec-kit-through-its-paces-radical-idea-or-reinvented-waterfall.html" rel="noopener noreferrer"&gt;One-shot&lt;/a&gt;, no cross-session continuity&lt;/td&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;Kiro&lt;/strong&gt; — IDE with built-in spec workflow (&lt;a href="https://kiro.dev/blog/kiro-and-the-future-of-software-development/" rel="noopener noreferrer"&gt;kiro.dev&lt;/a&gt;)&lt;/td&gt;
&lt;td&gt;Integrated experience&lt;/td&gt;
&lt;td&gt;
&lt;a href="https://martinfowler.com/articles/exploring-gen-ai/sdd-3-tools.html" rel="noopener noreferrer"&gt;Static specs, manual updates&lt;/a&gt;, IDE-locked&lt;/td&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Every approach above shares a common failure mode: &lt;strong&gt;they treat specification as a one-time event, not a continuous process.&lt;/strong&gt; You write the RFC, make the decision, and move on. You create the design doc, build the feature, and the doc rots. You set up CLAUDE.md on day one, and by week three it describes a project that no longer exists.&lt;/p&gt;

&lt;p&gt;Drew Breunig captured this perfectly with the &lt;a href="https://www.dbreunig.com/2026/03/04/the-spec-driven-development-triangle.html" rel="noopener noreferrer"&gt;Spec-Driven Development Triangle&lt;/a&gt; — specs, code, and tests form a triangle that must stay in sync, but keeping them in sync is where everyone fails.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Sync Problem Is a Workflow Problem
&lt;/h3&gt;

&lt;p&gt;Here's the insight that most tools miss: &lt;strong&gt;spec drift isn't a documentation problem. It's a workflow problem.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;You can't solve it by writing better specs. You can't solve it by adding a linter that checks specs against code. You can't solve it with a pre-commit hook that nags you to update the docs.&lt;/p&gt;

&lt;p&gt;You solve it by making knowledge maintenance &lt;strong&gt;an inseparable part of the development workflow itself&lt;/strong&gt; — not something you do after the "real work," but part of the work.&lt;/p&gt;

&lt;p&gt;This requires three things:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;A knowledge base&lt;/strong&gt; that's structured enough for AI to reference, but lightweight enough for humans to maintain&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A sync mechanism&lt;/strong&gt; that's embedded in the development lifecycle, not bolted on as an afterthought&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;An iterative workflow&lt;/strong&gt; that revisits and evolves knowledge across sessions, not just within a single feature&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Most tools get one or two of these. Almost none get all three.&lt;/p&gt;

&lt;h3&gt;
  
  
  REAP's Answer: Genome + Sync + Recursive Workflow
&lt;/h3&gt;

&lt;p&gt;This is the problem &lt;a href="https://github.com/c-d-cc/reap" rel="noopener noreferrer"&gt;REAP&lt;/a&gt; was built to solve. Not by treating specs as source code, but by building a &lt;strong&gt;recursive workflow where knowledge evolves alongside the code it describes&lt;/strong&gt;.&lt;/p&gt;

&lt;h4&gt;
  
  
  The Genome: A Living Knowledge Base
&lt;/h4&gt;

&lt;p&gt;REAP maintains a "Genome" — a structured collection of project knowledge stored in &lt;code&gt;.reap/genome/&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;.reap/genome/
  principles.md      # Architecture decisions (ADR-style, with rationale)
  conventions.md      # Development rules and enforced standards
  constraints.md      # Technical choices and validation commands
  domain/             # Business rules that can't be derived from code
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The Genome isn't a spec. It doesn't try to describe &lt;em&gt;what to build&lt;/em&gt;. It captures &lt;strong&gt;what you've learned&lt;/strong&gt; — architecture principles, business rules, constraints, conventions. It's the accumulated knowledge that makes your project &lt;em&gt;your project&lt;/em&gt;, not a generic codebase.&lt;/p&gt;

&lt;p&gt;Every time an AI agent starts a session in a REAP project, the Genome is automatically injected into its context. The agent doesn't start from zero — it starts with your project's institutional knowledge.&lt;/p&gt;

&lt;h4&gt;
  
  
  Sync Through the Lifecycle, Not After It
&lt;/h4&gt;

&lt;p&gt;Here's where REAP diverges from every tool listed above. Knowledge sync isn't a separate activity — it's &lt;strong&gt;built into the development lifecycle&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Each "Generation" (a unit of work) follows a five-stage cycle:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Objective → Planning → Implementation → Validation → Completion
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;During &lt;strong&gt;Implementation&lt;/strong&gt;, when you discover something that contradicts the Genome — a business rule that changed, an architectural assumption that proved wrong — you don't stop to update docs. You log it as a backlog item and keep building.&lt;/p&gt;

&lt;p&gt;During &lt;strong&gt;Completion&lt;/strong&gt;, those discoveries are reviewed and the Genome is updated. Knowledge evolution happens as a natural part of finishing work, not as a separate maintenance chore that everyone skips.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;This is the critical difference.&lt;/strong&gt; The Genome stays in sync with reality because updating it is part of the workflow, not something you do "when you have time" (which means never).&lt;/p&gt;

&lt;h4&gt;
  
  
  Recursive, Not One-Shot
&lt;/h4&gt;

&lt;p&gt;But the most important differentiator isn't the Genome or the sync — it's that the &lt;strong&gt;workflow is recursive&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Spec Kit gives you: Specify → Plan → Task → Implement. Done. Start over from scratch for the next feature.&lt;/p&gt;

&lt;p&gt;REAP gives you an &lt;strong&gt;endless chain of generations&lt;/strong&gt;, where each generation inherits the knowledge from all previous ones:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Gen 1: Build auth → learns "we use JWT" → Genome updated
Gen 2: Build API → starts knowing "we use JWT" → learns "rate limiting needed" → Genome updated
Gen 3: Build dashboard → starts knowing both → builds on accumulated knowledge
...
Gen N: Genome reflects N generations of accumulated learning
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each generation archives its artifacts in a &lt;strong&gt;Lineage&lt;/strong&gt; — a complete history of what was decided, what was built, and what was learned. The Genome is a living summary; the Lineage is the full record.&lt;/p&gt;

&lt;p&gt;This recursive structure means:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;No cold starts&lt;/strong&gt; — Every generation begins with the full context of everything that came before&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No knowledge loss&lt;/strong&gt; — Decisions made in generation 5 are still accessible in generation 50&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Natural evolution&lt;/strong&gt; — The Genome grows more accurate over time, not less — the opposite of traditional specs&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  The Right Mental Model
&lt;/h3&gt;

&lt;p&gt;Here's how to think about it:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Traditional&lt;/th&gt;
&lt;th&gt;SDD&lt;/th&gt;
&lt;th&gt;REAP&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Code is...&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;The only truth&lt;/td&gt;
&lt;td&gt;A generated artifact&lt;/td&gt;
&lt;td&gt;The truth, always&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Spec is...&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Pre-work that rots&lt;/td&gt;
&lt;td&gt;Source of truth&lt;/td&gt;
&lt;td&gt;Per-generation Objective (scoped, disposable)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Knowledge is...&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;In people's heads&lt;/td&gt;
&lt;td&gt;In spec documents&lt;/td&gt;
&lt;td&gt;In an evolving Genome&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Workflow is...&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Ad hoc&lt;/td&gt;
&lt;td&gt;One-shot pipeline&lt;/td&gt;
&lt;td&gt;Recursive generations&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Sync happens...&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Never&lt;/td&gt;
&lt;td&gt;Manually&lt;/td&gt;
&lt;td&gt;Built into each generation's completion&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Code remains the source of truth. The Genome doesn't replace it — it &lt;strong&gt;complements&lt;/strong&gt; it by capturing the &lt;em&gt;intent, rationale, and constraints&lt;/em&gt; that code alone can't express. And the recursive workflow ensures the two stay in sync, generation after generation.&lt;/p&gt;

&lt;h3&gt;
  
  
  Try It
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npm &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-g&lt;/span&gt; @c-d-cc/reap
reap init my-project

&lt;span class="c"&gt;# In Claude Code or OpenCode:&lt;/span&gt;
&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; /reap.start
&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; /reap.evolve &lt;span class="s2"&gt;"Implement user authentication"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;REAP is open source, MIT licensed, and supports Claude Code and OpenCode today.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/c-d-cc/reap" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt; | &lt;a href="https://reap.cc" rel="noopener noreferrer"&gt;Documentation&lt;/a&gt; | &lt;a href="https://www.npmjs.com/package/@c-d-cc/reap" rel="noopener noreferrer"&gt;npm&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Specs can't be source code. But the intent behind them — the decisions, the constraints, the hard-won lessons — that's worth managing. The question is whether your workflow makes that management automatic or optional. Because optional means it won't happen.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;References:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://reproducible-builds.org/docs/deterministic-build-systems/" rel="noopener noreferrer"&gt;Reproducible Builds — reproducible-builds.org&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://arxiv.org/abs/2602.00180" rel="noopener noreferrer"&gt;Spec-Driven Development: From Code to Contract — arXiv&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://newsletter.pragmaticengineer.com/p/rfcs-and-design-docs" rel="noopener noreferrer"&gt;Engineering Planning with RFCs, Design Documents and ADRs — Pragmatic Engineer&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://agents.md/" rel="noopener noreferrer"&gt;AGENTS.md — Open Standard for AI Agents&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.dbreunig.com/2026/03/04/the-spec-driven-development-triangle.html" rel="noopener noreferrer"&gt;The Spec-Driven Development Triangle — Drew Breunig&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.blog/ai-and-ml/generative-ai/spec-driven-development-with-ai-get-started-with-a-new-open-source-toolkit/" rel="noopener noreferrer"&gt;Spec-Driven Development with AI — GitHub Blog&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://martinfowler.com/articles/exploring-gen-ai/sdd-3-tools.html" rel="noopener noreferrer"&gt;Exploring SDD Tools — Martin Fowler&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://blog.scottlogic.com/2025/11/26/putting-spec-kit-through-its-paces-radical-idea-or-reinvented-waterfall.html" rel="noopener noreferrer"&gt;Putting Spec Kit Through Its Paces — Scott Logic&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://kiro.dev/blog/kiro-and-the-future-of-software-development/" rel="noopener noreferrer"&gt;Kiro and the Future of Software Development — kiro.dev&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.infoq.com/news/2026/03/agents-context-file-value-review/" rel="noopener noreferrer"&gt;AGENTS.md Value Reassessment — InfoQ&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
    </item>
    <item>
      <title>Why Spec-Driven Development Fails— And a Better Way to Structure AI Development</title>
      <dc:creator>Hichoi-Dev</dc:creator>
      <pubDate>Wed, 18 Mar 2026 21:14:10 +0000</pubDate>
      <link>https://dev.to/casamia918/why-spec-driven-development-fails-and-what-we-can-learn-from-it-2pec</link>
      <guid>https://dev.to/casamia918/why-spec-driven-development-fails-and-what-we-can-learn-from-it-2pec</guid>
      <description>&lt;h2&gt;
  
  
  SDD: The Right Problem, Wrong Solution
&lt;/h2&gt;

&lt;p&gt;Spec-Driven Development (SDD) is the idea that detailed specifications — written upfront — can guide AI agents to produce working software. GitHub's &lt;a href="https://github.com/github/spec-kit" rel="noopener noreferrer"&gt;Spec Kit&lt;/a&gt; is a representative example, formalizing this into a workflow: &lt;strong&gt;Specify → Plan → Task → Implement&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;SDD recognized a real problem: "prompt and pray" doesn't scale. Beyond toy projects, you need a way to communicate intent to AI that goes beyond "build me an auth system." The core insight — that &lt;strong&gt;structure matters&lt;/strong&gt; — is valid.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Core Problem: Specs Are Non-Deterministic
&lt;/h2&gt;

&lt;p&gt;The fundamental flaw: SDD treats specifications as authoritative sources of truth, but LLMs exhibit non-deterministic behavior. The same specification produces different implementations across different runs—varying architectural choices, data structures, and error handling. As the analysis notes, "Because of the non-deterministic nature of this technology, there will always remain a very non-negligible probability that it does things that we don't want." This means specifications cannot serve as reliable sources of truth the way source code does.&lt;/p&gt;

&lt;h2&gt;
  
  
  SDD Is Waterfall in Disguise
&lt;/h2&gt;

&lt;p&gt;SDD essentially recreates Waterfall methodology:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Big Design Up Front with exhaustive specifications&lt;/li&gt;
&lt;li&gt;Sequential phases completing before the next begins&lt;/li&gt;
&lt;li&gt;Assumption that thorough planning eliminates execution uncertainty&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Real-world testing revealed inefficiency: one hands-on evaluation required &lt;strong&gt;33 minutes and 2,577 lines of markdown&lt;/strong&gt; to produce 689 lines of code, compared to 8 minutes using iterative prompting—approximately 10x slower with no quality improvement.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Specifications Drift
&lt;/h2&gt;

&lt;p&gt;Specifications and code inevitably diverge because:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;AI makes unanticipated architectural choices&lt;/li&gt;
&lt;li&gt;Each iteration accumulates undocumented decisions&lt;/li&gt;
&lt;li&gt;Specs become post-hoc documentation rather than guides&lt;/li&gt;
&lt;li&gt;Developers spend time reading lengthy markdown instead of solving problems&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The Real Question
&lt;/h2&gt;

&lt;p&gt;Rather than "exhaustive upfront specifications," the answer aligns with decades of software engineering wisdom: &lt;strong&gt;iterative development with accumulated learning&lt;/strong&gt;—essentially Agile methodology adapted for AI collaboration.&lt;/p&gt;

&lt;h2&gt;
  
  
  A Different Approach
&lt;/h2&gt;

&lt;p&gt;This is what motivated me to build &lt;a href="https://github.com/c-d-cc/reap" rel="noopener noreferrer"&gt;REAP&lt;/a&gt; (Recursive Evolutionary Autonomous Pipeline). Rather than treating development as a spec-to-code translation, REAP structures AI-assisted development as an &lt;strong&gt;evolutionary process&lt;/strong&gt; — closer to how experienced developers actually work.&lt;/p&gt;

&lt;h3&gt;
  
  
  How REAP Works
&lt;/h3&gt;

&lt;p&gt;Development happens in &lt;strong&gt;Generations&lt;/strong&gt;. Each generation carries one focused goal through a 5-stage lifecycle:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Objective → Planning → Implementation → Validation → Completion
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This isn't just a linear pipeline. Each stage has gates, and stages can &lt;strong&gt;regress&lt;/strong&gt; — if validation fails, you loop back to implementation with the failure context preserved. This mirrors the real-world "build → test → fix → test again" cycle that SDD's sequential model ignores.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Genome: Knowledge That Evolves
&lt;/h3&gt;

&lt;p&gt;Where SDD puts specifications at the center, REAP puts a &lt;strong&gt;Genome&lt;/strong&gt; at the center — a living record stored in &lt;code&gt;.reap/genome/&lt;/code&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;principles.md&lt;/code&gt;&lt;/strong&gt; — Architecture decisions with rationale (ADR-style)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;conventions.md&lt;/code&gt;&lt;/strong&gt; — Development rules and enforced standards&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;constraints.md&lt;/code&gt;&lt;/strong&gt; — Technical choices and validation commands&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;domain/&lt;/code&gt;&lt;/strong&gt; — Business rules that can't be derived from code&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The Genome isn't written once and forgotten. It &lt;strong&gt;evolves across generations&lt;/strong&gt;. When you discover something during implementation that contradicts the Genome, you log it as a backlog item. At the end of each generation, discoveries are reviewed and the Genome is updated. Over time, the Genome becomes an increasingly accurate map of your project — not a spec that drifts from reality.&lt;/p&gt;

&lt;h3&gt;
  
  
  What Makes It Different from SDD
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;SDD&lt;/th&gt;
&lt;th&gt;REAP&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Source of truth&lt;/td&gt;
&lt;td&gt;Specification document&lt;/td&gt;
&lt;td&gt;Evolved Genome + source code&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Planning scope&lt;/td&gt;
&lt;td&gt;Entire project upfront&lt;/td&gt;
&lt;td&gt;One generation at a time&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;When plans break&lt;/td&gt;
&lt;td&gt;Spec drift → update spec → regenerate&lt;/td&gt;
&lt;td&gt;Discovery → backlog → evolve Genome&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Validation&lt;/td&gt;
&lt;td&gt;Spec compliance&lt;/td&gt;
&lt;td&gt;Actual tests, type checks, builds&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Knowledge persistence&lt;/td&gt;
&lt;td&gt;Specs (static)&lt;/td&gt;
&lt;td&gt;Genome (evolving) + Lineage (history)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Context for AI&lt;/td&gt;
&lt;td&gt;Spec document&lt;/td&gt;
&lt;td&gt;Genome + generation state (auto-injected)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  Context That Persists
&lt;/h3&gt;

&lt;p&gt;Every time you start an AI session in a REAP project, the &lt;strong&gt;SessionStart hook&lt;/strong&gt; automatically injects the Genome, current generation state, and workflow rules into the AI's context. The AI doesn't start from zero — it starts with your project's accumulated knowledge.&lt;/p&gt;

&lt;p&gt;This solves SDD's "spec drift" problem at the root. The Genome stays in sync with reality because it's &lt;strong&gt;updated as part of the development process&lt;/strong&gt;, not maintained as a separate artifact.&lt;/p&gt;

&lt;h3&gt;
  
  
  Try It
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npm &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-g&lt;/span&gt; &lt;span class="s2"&gt;"@c-d-cc/reap"&lt;/span&gt;
reap init my-project
&lt;span class="c"&gt;# Open Claude Code or OpenCode&lt;/span&gt;
&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; /reap.start
&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; /reap.evolve &lt;span class="s2"&gt;"Implement user authentication"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;REAP supports multiple AI agents — Claude Code and OpenCode today, with an extensible adapter system for adding more.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/c-d-cc/reap" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt; | &lt;a href="https://reap.cc" rel="noopener noreferrer"&gt;Documentation&lt;/a&gt; | &lt;a href="https://www.npmjs.com/package/@c-d-cc/reap" rel="noopener noreferrer"&gt;npm&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;What's your experience with spec-driven development? Have you found structure that works for AI-assisted development? I'd love to hear in the comments.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;References:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://github.com/github/spec-kit" rel="noopener noreferrer"&gt;GitHub Spec Kit&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.blog/ai-and-ml/generative-ai/spec-driven-development-with-ai-get-started-with-a-new-open-source-toolkit/" rel="noopener noreferrer"&gt;Spec-Driven Development with AI — GitHub Blog&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://blog.scottlogic.com/2025/11/26/putting-spec-kit-through-its-paces-radical-idea-or-reinvented-waterfall.html" rel="noopener noreferrer"&gt;Putting Spec Kit Through Its Paces — Scott Logic&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://marmelab.com/blog/2025/11/12/spec-driven-development-waterfall-strikes-back.html" rel="noopener noreferrer"&gt;SDD: The Waterfall Strikes Back — Marmelab&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://martinfowler.com/articles/exploring-gen-ai/sdd-3-tools.html" rel="noopener noreferrer"&gt;Exploring SDD Tools — Martin Fowler&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://developer.microsoft.com/blog/spec-driven-development-spec-kit" rel="noopener noreferrer"&gt;Diving Into SDD With Spec Kit — Microsoft Developer Blog&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://vibecoding.app/blog/spec-kit-review" rel="noopener noreferrer"&gt;Spec Kit Review — Vibecoding.app&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.thoughtworks.com/en-us/insights/blog/agile-engineering-practices/spec-driven-development-unpacking-2025-new-engineering-practices" rel="noopener noreferrer"&gt;SDD: Unpacking 2025's Key Practice — Thoughtworks&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>productivity</category>
      <category>softwaredevelopment</category>
      <category>softwareengineering</category>
    </item>
    <item>
      <title>I built a dev tool that "evolves" code with AI — REAP</title>
      <dc:creator>Hichoi-Dev</dc:creator>
      <pubDate>Wed, 18 Mar 2026 02:57:43 +0000</pubDate>
      <link>https://dev.to/casamia918/i-built-a-dev-tool-that-evolves-code-with-ai-reap-k17</link>
      <guid>https://dev.to/casamia918/i-built-a-dev-tool-that-evolves-code-with-ai-reap-k17</guid>
      <description>&lt;h2&gt;
  
  
  The Problem
&lt;/h2&gt;

&lt;p&gt;If you've been building with AI agents (like Claude Code), you've probably encountered these problems:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Context loss&lt;/strong&gt; — Start a new session and your context is gone. You end up clinging to long sessions just to avoid losing everything the AI has learned.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Stale documentation&lt;/strong&gt; — You try to persist knowledge in READMEs and CLAUDE.md files, but they quietly go stale as the project moves forward.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AI going rogue&lt;/strong&gt; — Sometimes the AI just ignores your carefully crafted docs and does its own thing anyway.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;We're all stuck at the same bottleneck — the context window just isn't enough for long-running projects.&lt;/p&gt;

&lt;p&gt;I tried existing tools like spec-kit and superpower — they're decent for one-off feature work, but didn't quite fit for sustained, long-term development.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Built
&lt;/h2&gt;

&lt;p&gt;So I built &lt;strong&gt;REAP&lt;/strong&gt; (Recursive Evolutionary Autonomous Pipeline) — an open-source CLI tool inspired by generational evolution in biology.&lt;/p&gt;

&lt;p&gt;The idea: &lt;strong&gt;AI and humans evolve software across generations.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Genome (Design &amp;amp; Knowledge)
  → Evolution (Generational Progress)
    → Civilization (Source Code)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  How It Works
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Genome
&lt;/h3&gt;

&lt;p&gt;Your project's design knowledge is managed as a "Genome" — architecture decisions, business rules, conventions, and constraints.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;.reap/genome/
├── principles.md      # Architecture principles
├── domain/            # Business rules
├── conventions.md     # Development conventions
└── constraints.md     # Technical constraints
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Life Cycle
&lt;/h3&gt;

&lt;p&gt;Each generation follows a five-stage lifecycle:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Objective → Planning → Implementation → Validation → Completion
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Objective&lt;/strong&gt; — Define goal, requirements, and acceptance criteria&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Planning&lt;/strong&gt; — Break down tasks, choose approach&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Implementation&lt;/strong&gt; — Build with AI + human collaboration&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Validation&lt;/strong&gt; — Run tests, verify completion&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Completion&lt;/strong&gt; — Retrospective + apply Genome changes + archive&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Evolution
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;When a generation completes, it gets archived in the lineage, and the next generation picks up new goals.&lt;/li&gt;
&lt;li&gt;Lessons learned within a generation get folded back into the Genome.&lt;/li&gt;
&lt;li&gt;Through this iterative pipeline, your source code (the "Civilization") keeps evolving.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Quick Start
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Install&lt;/span&gt;
npm &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-g&lt;/span&gt; @c-d-cc/reap

&lt;span class="c"&gt;# Initialize&lt;/span&gt;
reap init my-project

&lt;span class="c"&gt;# Run a full generation in Claude Code&lt;/span&gt;
claude
&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; /reap.evolve &lt;span class="s2"&gt;"Implement user authentication"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;/reap.evolve&lt;/code&gt; runs the entire generation lifecycle — from Objective through Completion — interactively with you.&lt;/p&gt;

&lt;h2&gt;
  
  
  Links
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;GitHub&lt;/strong&gt;: &lt;a href="https://github.com/c-d-cc/reap" rel="noopener noreferrer"&gt;github.com/c-d-cc/reap&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Docs&lt;/strong&gt;: &lt;a href="https://reap.cc" rel="noopener noreferrer"&gt;reap.cc&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;MIT licensed. Contributions and feedback are welcome!&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>opensource</category>
    </item>
  </channel>
</rss>
