<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Cleber de Lima</title>
    <description>The latest articles on DEV Community by Cleber de Lima (@cleberdelima).</description>
    <link>https://dev.to/cleberdelima</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3607194%2Fb81b1058-0dcb-45d3-84ad-67e471a945a1.png</url>
      <title>DEV Community: Cleber de Lima</title>
      <link>https://dev.to/cleberdelima</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/cleberdelima"/>
    <language>en</language>
    <item>
      <title>88% of companies use AI. Only 39% can point to any bottom-line impact.
Adoption you can buy. Impact you have to build.
This new article is the full map: seven stages, each building an asset the next depends on. 
#AIEgineering #AI #SDLC #AIDLC</title>
      <dc:creator>Cleber de Lima</dc:creator>
      <pubDate>Mon, 17 Aug 2026 09:54:34 +0000</pubDate>
      <link>https://dev.to/cleberdelima/88-of-companies-use-ai-only-39-can-point-to-any-bottom-line-impact-adoption-you-can-buy-3ke1</link>
      <guid>https://dev.to/cleberdelima/88-of-companies-use-ai-only-39-can-point-to-any-bottom-line-impact-adoption-you-can-buy-3ke1</guid>
      <description>&lt;div class="ltag__link--embedded"&gt;
  &lt;div class="crayons-story "&gt;
  &lt;a href="https://dev.to/cleberdelima/the-ai-operating-model-the-full-roadmap-from-adoption-to-impact-4aco" class="crayons-story__hidden-navigation-link"&gt;The AI Operating Model: The Full Roadmap from Adoption to Impact&lt;/a&gt;


  &lt;div class="crayons-story__body crayons-story__body-full_post"&gt;
    &lt;div class="crayons-story__top"&gt;
      &lt;div class="crayons-story__meta"&gt;
        &lt;div class="crayons-story__author-pic"&gt;

          &lt;a href="/cleberdelima" class="crayons-avatar  crayons-avatar--l  "&gt;
            &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3607194%2Fb81b1058-0dcb-45d3-84ad-67e471a945a1.png" alt="cleberdelima profile" class="crayons-avatar__image"&gt;
          &lt;/a&gt;
        &lt;/div&gt;
        &lt;div&gt;
          &lt;div&gt;
            &lt;a href="/cleberdelima" class="crayons-story__secondary fw-medium m:hidden"&gt;
              Cleber de Lima
            &lt;/a&gt;
            &lt;div class="profile-preview-card relative mb-4 s:mb-0 fw-medium hidden m:inline-block"&gt;
              
                Cleber de Lima
                
                
              
              &lt;div id="story-author-preview-content-4416727" class="profile-preview-card__content crayons-dropdown branded-7 p-4 pt-0"&gt;
                &lt;div class="gap-4 grid"&gt;
                  &lt;div class="-mt-4"&gt;
                    &lt;a href="/cleberdelima" class="flex"&gt;
                      &lt;span class="crayons-avatar crayons-avatar--xl mr-2 shrink-0"&gt;
                        &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3607194%2Fb81b1058-0dcb-45d3-84ad-67e471a945a1.png" class="crayons-avatar__image" alt=""&gt;
                      &lt;/span&gt;
                      &lt;span class="crayons-link crayons-subtitle-2 mt-5"&gt;Cleber de Lima&lt;/span&gt;
                    &lt;/a&gt;
                  &lt;/div&gt;
                  &lt;div class="print-hidden"&gt;
                    
                      Follow
                    
                  &lt;/div&gt;
                  &lt;div class="author-preview-metadata-container"&gt;&lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
            &lt;/div&gt;

          &lt;/div&gt;
          &lt;a href="https://dev.to/cleberdelima/the-ai-operating-model-the-full-roadmap-from-adoption-to-impact-4aco" class="crayons-story__tertiary fs-xs"&gt;&lt;time&gt;Aug 17&lt;/time&gt;&lt;span class="time-ago-indicator-initial-placeholder"&gt;&lt;/span&gt;&lt;/a&gt;
        &lt;/div&gt;
      &lt;/div&gt;

    &lt;/div&gt;

    &lt;div class="crayons-story__indention"&gt;
      &lt;h2 class="crayons-story__title crayons-story__title-full_post"&gt;
        &lt;a href="https://dev.to/cleberdelima/the-ai-operating-model-the-full-roadmap-from-adoption-to-impact-4aco" id="article-link-4416727"&gt;
          The AI Operating Model: The Full Roadmap from Adoption to Impact
        &lt;/a&gt;
      &lt;/h2&gt;
        &lt;div class="crayons-story__tags"&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/ai"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;ai&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/softwareengineering"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;softwareengineering&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/softwaredevelopment"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;softwaredevelopment&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/agents"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;agents&lt;/a&gt;
        &lt;/div&gt;
      &lt;div class="crayons-story__bottom"&gt;
        &lt;div class="crayons-story__details"&gt;
            &lt;a href="https://dev.to/cleberdelima/the-ai-operating-model-the-full-roadmap-from-adoption-to-impact-4aco#comments" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left flex items-center"&gt;
              

              &lt;span class="hidden s:inline"&gt;Add&amp;nbsp;Comment&lt;/span&gt;
            &lt;/a&gt;
        &lt;/div&gt;
        &lt;div class="crayons-story__save"&gt;
          &lt;small class="crayons-story__tertiary fs-xs mr-2"&gt;
            45 min read
          &lt;/small&gt;
        &lt;/div&gt;
      &lt;/div&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;/div&gt;


</description>
    </item>
    <item>
      <title>The AI Operating Model: The Full Roadmap from Adoption to Impact</title>
      <dc:creator>Cleber de Lima</dc:creator>
      <pubDate>Mon, 17 Aug 2026 09:44:11 +0000</pubDate>
      <link>https://dev.to/cleberdelima/the-ai-operating-model-the-full-roadmap-from-adoption-to-impact-4aco</link>
      <guid>https://dev.to/cleberdelima/the-ai-operating-model-the-full-roadmap-from-adoption-to-impact-4aco</guid>
      <description>&lt;p&gt;Your AI rollout looks finished. Licenses are deployed, the adoption charts point up, most merged work touches an agent somewhere, and every team can show you something impressive it built last month. Then someone on the board asks the only question that matters: what did all of this change in the business? The room goes quiet, because the honest answer is smaller than the dashboards suggest.&lt;/p&gt;

&lt;p&gt;You are not alone in that room. &lt;a href="https://www.deloitte.com/us/en/insights/topics/talent/ai-adoption-to-ai-adaptation.html" rel="noopener noreferrer"&gt;Deloitte's research on AI adoption&lt;/a&gt; finds that 84 percent of organizations have not redesigned jobs or workflows around AI, and that fewer than 60 percent of workers who have access to AI actually use it in their daily work. &lt;a href="https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai" rel="noopener noreferrer"&gt;McKinsey's State of AI survey&lt;/a&gt; shows 88 percent of companies using AI in at least one function, while only 39 percent can attribute any bottom-line impact to it, and most of those say the impact is below five percent. Near-universal usage, near-zero transformation. That gap is not a technology gap. It is an operating model gap.&lt;/p&gt;

&lt;p&gt;This article is the map for closing it: the full roadmap from adoption, which you can buy, to impact, which you have to build. For the past twelve months I have been writing about the individual aspects of this transformation, one article at a time: machine-ready specs, testing, continuous flow, loops, evals, token economics, the harness. Each of those pieces answered one question and deliberately left the bigger one open: in which order do these capabilities come together, and what does the whole journey look like when one organization runs it end to end? This is that article, which is why it is longer than my usual ones. The decision it asks of you is simple to state. Will you run AI adoption as a staged transformation, where each stage builds an asset the next stage depends on, or will you keep running it as a rolling tool deployment and hope that depth shows up on its own?&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;My experience is blunt on this one: choose the latter, and depth will not show up by itself.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Two terms carry the whole piece, so let me define them before anything depends on them. &lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The operating model&lt;/strong&gt; is the full system that turns intent into running software: who decides what gets built, how work flows from an idea to production, where quality is checked, what gets measured, and who is accountable for what. &lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;AI transformation is not a tool deployment sitting on top of that system. It is a redesign of the system itself. &lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;Maturity&lt;/strong&gt; is the second term, and it hides two questions that most scorecards collapse into one: are teams using AI well in their everyday work, and is the organization capable of sustaining and scaling what those teams do? You can score high on the first and fail the second, and actually this is where most companies are today.&lt;/p&gt;

&lt;p&gt;Accenture and Carnegie Mellon's Software Engineering Institute published a &lt;a href="https://www.sei.cmu.edu/news/sei-and-accenture-release-ai-adoption-maturity-model-to-help-organizations-scale-ai-with-predictable-outcomes/" rel="noopener noreferrer"&gt;maturity model in June 2026&lt;/a&gt; built from a review of more than 100 existing frameworks, interviews with executives, and a survey of nearly 600 practitioners. The sentence from SEI's Ipek Ozkaya that should hang in your program office: "True AI maturity is not measured by how much AI an organization deploys, but by its ability to build trustworthy and resilient capabilities, rigorous engineering practices, and governance approaches aligned with business outcomes." That is this article compressed into a scorecard.&lt;/p&gt;

&lt;p&gt;The moment that taught me the difference arrived on one slide of my transformation program at Betsson. The adoption telemetry was everything a CTO could ask for: nearly every engineer active weekly, agent features in regular use, a large share of merged work carrying AI assistance. Then we scored the same teams against a five-level maturity ladder, level one meaning ad-hoc individual experimentation and level five meaning self-improving systems inside fixed guardrails, and most teams landed in the bottom two levels. Both numbers were true. The telemetry measured opened doors; the ladder measured changed work. &lt;/p&gt;

&lt;p&gt;The proof that the gap was real, and closable, came from one particular team that decided to really lean into the transformation. Same people, completely different approach and ways of working: intent turned into machine-readable specs, machine gates did the first pass of verification instead of tired human eyes, humans approved at defined checkpoints, and a small, focused multidisciplinary squad owned the outcome from definition to production. &lt;/p&gt;

&lt;p&gt;That team's delivery cycle went from around 80 days to 5. Code generation was a small part of the story. The breakthrough was structural, and I have been telling it on stage ever since, because it is the whole argument of this article.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Rebuilding the processes and the collaboration around the new capabilities AI makes possible is the quantum leap, and it is what brings the biggest payoff.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Sit back and enjoy the reading: what follows is that journey, from breadth without depth to a whole organization running the way that one team did. &lt;/p&gt;

&lt;h2&gt;
  
  
  The map at a glance
&lt;/h2&gt;

&lt;p&gt;Before going into the details, let's define the shape. &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Stage 0, Prepare:&lt;/strong&gt; the tools, the models, and the connectors are evaluated and decided once; budgets and limits are set inside the tools; and the governance baseline and first harness standards are written before the crowd arrives. The asset is a decided stack with rules.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Stage 1, Adopt:&lt;/strong&gt; the first training waves run with chosen domains and chosen people, and their learnings feed the harness. The asset is proof, plus a trained first cohort. &lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Stage 2, Structure:&lt;/strong&gt; the harness evolves, the company knowledge base starts to compound, and specifications become machine-ready. The asset is an environment where agents work reliably at falling cost. &lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Stage 3, Scale:&lt;/strong&gt; the proven practice rolls out wave after wave to every team, carried by FDEs and harness engineers, and every new team onboards straight into the harness. The asset is the practice at full coverage.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Stage 4, Verify:&lt;/strong&gt; evals, systematic tests for AI behavior, and gates take over the first pass of quality. The asset is trust you can measure. &lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Stage 5, Automate:&lt;/strong&gt; loops run work unattended inside the gates, and a managed model portfolio serves them at the right cost. The asset is delivery capacity that no longer scales with attention. &lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Stage 6, Reorganize:&lt;/strong&gt; the ways of working change shape, product and engineering merge around the new speed, and the same playbook extends beyond engineering. The asset is an operating model that compounds. &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Read together, the seven stages are the distance between the two numbers in the opening: adoption is what the early stages secure, and measured impact is what the last stage finally delivers.&lt;/p&gt;

&lt;h2&gt;
  
  
  Stage 0: Prepare. The asset is a decided stack, not a catalog
&lt;/h2&gt;

&lt;p&gt;There is a stage before the rollout most people call the start, and it is the cheapest insurance in this whole journey. Before the first license lands in a builder's hands, four decisions need a first version: which tools the organization runs, what they are allowed to cost, which rules keep you compliant, and which working standards every team will follow.&lt;/p&gt;

&lt;p&gt;The first is the stack decision, and it covers three lists, not one. &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Tools:&lt;/strong&gt; run a short, honest evaluation of the coding assistants, agentic IDEs, and no-code agent builders against your own work, not against vendor demos, with security and cost vetting built into the process. &lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Models:&lt;/strong&gt; standardize which models are approved for which kind of work, one default per job plus vetted alternatives. This is standardization, not routing; the machinery that dispatches each request to the right model belongs to the harness and the portfolio in later stages, but the menu it dispatches from is decided here, because a menu every engineer composes alone is how spend and quality become impossible to compare. &lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Connectors:&lt;/strong&gt; evaluate the external MCP servers and third-party tools your agents will call, and treat them as the supply chain they are, because a connector receives credentials and reaches into your systems, so nothing gets on the approved list without a security review.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The output across all three is one approved list with three statuses, standard, pilot, and retired, a pilot gate for new entries, and a clear answer to the question every engineer will otherwise answer alone: which tool, which model, which connector, for which job. One boundary keeps this stage honest: it governs what you adopt from the market, not what you build. The MCPs and tools you build yourself belong to the harness work of stage two, and the economics of running models on your own infrastructure at scale belong to the portfolio decisions of stage five.&lt;/p&gt;

&lt;p&gt;The second is the money, and it must be set up inside the tools on day one, not discovered on the first invoice. The billing model of this stack changed under everyone's feet: coding assistants moved from flat seats to metered, usage-based pricing, which means an agentic workflow can spend in an afternoon what a seat used to cost in a month. &lt;a href="https://thenextweb.com/news/microsoft-claude-code-retreat-ai-cost" rel="noopener noreferrer"&gt;Uber's CTO described burning the entire planned annual AI coding budget in four months&lt;/a&gt;, and Microsoft answered the same math by pulling agentic tool access from thousands of its own engineers. So allocate budgets per team and per seat before the rollout, and use the admin planes the tools now ship: default models set to the standard tier, model allow-lists matching your approved menu, spend alerts at the levels where you want a conversation to happen, and hard caps only as the backstop against runaways. Two lessons from operating this. Defaults beat caps: most people never come close to their limits, so the real spend lever is which model the tool reaches for first, not how hard you cap the ceiling. And budget by persona, not by average: a small group of power users will legitimately spend many times the median because they are running the most agentic work, and a cap sized for the average punishes exactly the people getting the most from the stack; give them alerts and a review, not a wall. The organization-wide chokepoint, one gateway with wallets and circuit breakers, comes in stage five; stage zero sets the limits where the tools already provide them.&lt;/p&gt;

&lt;p&gt;The third is the governance baseline, and if you operate in Europe this is not optional pre-work, it is law with dates attached. The &lt;a href="https://digital-strategy.ec.europa.eu/en/policies/regulatory-framework-ai" rel="noopener noreferrer"&gt;EU AI Act&lt;/a&gt; is in force and phasing in: its prohibitions and its AI literacy obligation have applied since February 2025, which means organizations deploying AI must already ensure the people using it are adequately trained, and further obligations keep arriving through 2027. Three pieces need a first version before rollout. An acceptable-use policy that tells every builder what may and may not go into these tools: which code, which data, which customer information, under which vendor terms, with retention and training-on-your-data clauses checked per tool on the approved list. A risk classification of your intended uses against the AI Act's categories, so you know which are minimal risk, which carry transparency duties, and which would cross into high-risk territory and need a different governance lane before they ship. And named accountability: who approves a new AI use case, who owns incidents involving AI output, and where the record of those decisions lives. Thin, again, is fine: a two-page policy every team has actually read beats a governance framework nobody has finished. And notice the convenient overlap: the literacy the law requires is exactly the training stage one delivers, so compliance and capability are the same budget line here, not competing ones.&lt;/p&gt;

&lt;p&gt;The fourth is the first harness definition. Stage two will build the harness in depth, but its skeleton must exist before adoption scales: the instruction-file template every repository starts from, the naming and state conventions, the guardrail baseline naming the actions no agent may ever take, and the default autonomy level a team starts at. Thin is fine; written is what matters. A one-page standard adopted by the first ten teams beats a perfect framework announced after fifty teams have invented their own.&lt;/p&gt;

&lt;p&gt;The failure mode of skipping this stage is tool sprawl, and unlike the other failure modes it is almost invisible while it happens. Every team picks its own assistant, negotiates its own licenses, writes its own rules or none, and the organization discovers months later that it runs a dozen overlapping tools with no shared standard, ungoverned spend scattered across cost centers, and a security review backlog nobody planned. The money goes into duplicated licenses and, worse, into the consolidation program you will eventually have to run to undo it all, at many times the cost of deciding early.&lt;/p&gt;

&lt;p&gt;You are ready to advance the moment the approved list, the budgets, the policy, and the one-page standard exist. This stage is measured in weeks, and it costs decision effort, not budget.&lt;/p&gt;

&lt;p&gt;The asset of this stage, made concrete:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Artifacts:&lt;/strong&gt; the approved list covering tools, models, and external connectors, each with standard, pilot, and retired status and a pilot gate for new entries; the model standard, one default per kind of work plus vetted alternatives; the security review record per approved connector; the budget allocation per team and per seat, configured in each tool's admin plane with standard-tier defaults, allow-lists, spend alerts, and hard caps as the backstop; the acceptable-use policy with per-tool data boundaries; the AI Act risk classification of intended uses, with named accountability for approving new ones; the instruction-file template; the guardrail baseline naming the actions no agent may take; the default autonomy level per domain; the one-page harness standard every onboarded team starts from.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Metrics:&lt;/strong&gt; tools, models, and connectors per job category, which should be one standard plus at most one pilot each; share of teams on the standard stack; spend per seat and per team against budget, visible from the first week; share of spend landing on the standard model tier; time from a new entry appearing to an approve-or-retire decision; AI spend outside the approved list and unvetted connectors in use, both of which should trend to zero.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Impact:&lt;/strong&gt; every team that onboards in stage one starts inside the same rules, on the same models, through vetted connectors, with spend visible and bounded from the first week, so the harness work of stage two lands on one foundation instead of fifty improvised ones, finance never meets this program through a surprise invoice, and the consolidation program nobody budgets for never has to exist.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Stage 1: Adopt. The asset is people, not licenses
&lt;/h2&gt;

&lt;p&gt;With the stack decided, everything starts with tools in hands, and this stage is where most organizations both start and quietly stop. The work is real: roll out the stack the previous stage decided, clear the remaining security and legal steps, and put it in front of every builder. But the deliverable of stage one is not deployed software. It is changed behavior, and behavior does not change by email announcement.&lt;/p&gt;

&lt;p&gt;Unless your team is quite small, what works is running adoption in waves, with a repeating change cycle inside each wave, not a one-shot rollout. And the waves come in two distinct movements that ask for different disciplines, the first one here in stage one, the second in stage three once the harness is ready: &lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The first waves exist to prove and to learn, the scale-out exists to industrialize what they proved. Confuse the two and you either scale what was never proven or keep re-proving what already works. &lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Sequence the waves deliberately: start with the domains that are most ready and where value is densest, take them through the full cycle, and let their results pull the next wave in. Inside each wave the cycle is the same four moves:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Art of Possible:&lt;/strong&gt; Show the domain what is possible with demos built on its own work, not generic ones. You will be surprised by the reaction of some people when they see a good result they did not believe was possible.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Training:&lt;/strong&gt; Train in role-based cohorts measured in hours of practice, not a lunch-and-learn, so every wave graduates a group that worked with the tools on its real backlog. &lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Champions Network:&lt;/strong&gt; Grow champions inside the teams, embedded users who scale adoption peer to peer, because engineers believe colleagues before they believe a central program. &lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Share the Success:&lt;/strong&gt; Collect the wave's success stories and feed them forward as the marketing for the next one. &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Then keep the waves coming, because the tools change monthly and most feature adoption happens long after the first rollout. Run a few waves and one pattern becomes hard to ignore: the same license produces a power user or a dormant seat depending almost entirely on the enablement around it, which is why hours of practice predict usage far better than tool choice does.&lt;/p&gt;

&lt;p&gt;Who goes in the first wave matters as much as what the wave teaches, so compose it deliberately.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Do not just open enrolment, and never roll out everything to everyone at the same time. This is change management, not scheduling.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Open the enrolment to a selected audience only, and let them opt in. The idea here is to ensure the right people are on the first waves, but only the ones who really want to be there. Fill most of the seats with the people already leaning in: the ones experimenting on their own, asking for licenses, showing colleagues what they built. Give preference to your respected architects and senior developers who are willing to commit to an enterprise-level approach rather than a personal toolkit, because they are the cornerstones: when the people others already seek for advice adopt a new technology or way of working seriously, the rest of the organization tends to follow without being pushed. &lt;/p&gt;

&lt;p&gt;Here is a tricky one: reserve a few seats, a few, not many, for your sharpest skeptics and open detractors. Their concerns are usually rational, the case I made in a previous &lt;a href="https://dev.to/cleberdelima/smart-engineers-rational-resistance-and-real-ai-adoption-5eo8"&gt;article&lt;/a&gt;, and a skeptic who converts inside the first wave becomes more persuasive than any champion, while one who stays unconvinced hands you the failure catalogue for free, before it costs anything.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;What you should never do is mandate the first wave onto the indifferent: forced adopters produce the dormant seats your dashboards will later mistake for a tooling problem.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is the whole job of the first waves: not coverage, but proof, converted skeptics, visible wins, and the first honest version of what to teach everyone else.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;And do not fund the waves alone.&lt;/strong&gt; The vendors on your approved list want your adoption to succeed at least as much as you do, because their renewal depends on it, so negotiate enablement into the contracts while stage zero is still deciding the stack: training credits, instructor-led workshops, office hours with their engineers, certification seats, early-access programs. Partners can carry the surge too: bringing outside trainers into the first waves is cheaper and faster than building an internal academy before you know what to teach, and by the time the later waves arrive, your own champions can take over the room. The training budget shrinks considerably when the people selling you the tools co-fund the adoption of them.&lt;/p&gt;

&lt;p&gt;Scaling to everyone is deliberately not this stage's job. The first waves exist to prove, to learn, and to feed the harness; carrying the practice to the whole organization is stage three, and it only starts once the harness is ready to receive teams.&lt;/p&gt;

&lt;p&gt;Expect the dip. Productivity goes down before it goes up, because people are learning a new way to work while still being paid to deliver: &lt;a href="https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/" rel="noopener noreferrer"&gt;METR measured&lt;/a&gt; experienced developers 19 percent slower with AI tools while believing they were 24 percent faster. People need time to adjust and adapt, and that shows up as a cost on cycle time that is recovered as the practice matures. In my experience the payback comes in sprints, not in months or quarters, so insist, and keep the focus on improving and compounding.&lt;/p&gt;

&lt;p&gt;This stage fails in two directions, and both burn money. Stopping here is the license illusion, the failure from the opening paragraphs. Deloitte's researchers draw the distinction as adoption versus adaptation: adoption tells you someone opened the door, adaptation tells you they changed how they work. An organization that stops at stage one keeps paying a tool bill that grows with headcount while the value stays anecdotal, sits in a J-curve dip it never climbs out of, and reads dashboards that look like progress. Skipping it is worse: harnesses, gates, and agents rolled out to people who never built the judgment to use them become infrastructure for users who do not show up, and you will meet that cost again at every later stage, because a skipped asset does not disappear, it comes back with interest. You are ready to advance when the first waves have delivered their proof, cornerstones committed, skeptics converted or their objections documented, graduates still active a month later, and when your training has moved past prompting tricks into judgment: what to delegate, what to verify, when to stop trusting the output, how to optimize token usage and generate repeatable workflows.&lt;/p&gt;

&lt;p&gt;The asset of this stage, made concrete:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Artifacts:&lt;/strong&gt; the first-wave plan, rosters composed for proof: volunteers, cornerstone architects and developers, a few skeptics; a role-based training curriculum with a learning path per discipline; vendor and partner enablement commitments written into the contracts; a named champion in every team; a success-story library tagged by domain; an AI literacy module inside week-one onboarding; adoption dashboards that count people changing how they work, not seats assigned.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Metrics:&lt;/strong&gt; first waves completed, with their learnings documented and handed to the harness team; weekly active users as a share of all builders, not of licenses; hours of structured practice per person; the share of trained people still active thirty days after training; the number of domains with a live champion.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Impact:&lt;/strong&gt; the J-curve dip gets shorter and shallower, usage holds after the novelty fades, and every later stage inherits users instead of seats, and the harness team starts stage two with a real backlog instead of theory. If thirty-day retention of trained users is high, this asset exists; if usage craters after the launch push, it does not, whatever the license count says.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Stage 2: Structure. The harness is the multiplier
&lt;/h2&gt;

&lt;p&gt;Here is the reframe that separates organizations that scale from those that plateau: agent results are mostly not a model property. They are a property of the environment you run the model in. That environment has a name, the harness, and building it is the central engineering investment of the whole journey. I made this case in &lt;a href="https://dev.to/cleberdelima/loop-engineering-stop-prompting-your-agents-and-design-the-system-that-does-428p"&gt;Loop Engineering&lt;/a&gt; and it deserves its full form here: the harness is everything around the model, the instructions, the tools it may call, the environment it runs in, the state it keeps between sessions, and the feedback that tells it whether its work passed. A prompt file is not a harness. A harness is those five subsystems working together, and the feedback subsystem, the explicit commands that verify work, returns more than any other investment.&lt;/p&gt;

&lt;p&gt;The evidence for how much this layer matters is now quantified. &lt;a href="https://arxiv.org/abs/2605.23950" rel="noopener noreferrer"&gt;A position paper on harness-induced versus model-induced variance&lt;/a&gt; found in its variance decomposition on SWE-bench tasks that swapping the harness moved success rates several times more than swapping the model, with harness effects around seven to eight times larger than model effects, enough to reverse model rankings in most of the configurations tested. &lt;a href="https://arxiv.org/abs/2607.06906" rel="noopener noreferrer"&gt;A separate industry study, The Harness Effect&lt;/a&gt;, held tasks and models constant and changed only the orchestration layer: cost per task fell 41 percent, latency fell 44 percent, and token use fell 38 percent at the same output quality, though note it is authored by a vendor evaluating its own harness, so read the direction, not the decimals. The practical rule I give every team: when agent output disappoints, do not reach for a bigger model first. Check the harness. One well-written instruction file routinely outperforms a model upgrade, at a tiny fraction of the cost.&lt;/p&gt;

&lt;p&gt;Be clear about the starting point and the destination, because they are far apart. The harness you inherit from stage zero is a skeleton: a set of instructions, the approved MCP tools, and the guardrails. Stage two is where that skeleton grows into an enterprise enabler, and it grows in three directions. First, codified workflows: the recurring jobs of each discipline, turning a ticket into a spec, generating and reviewing code, design to code, infrastructure changes, incident analysis, each captured as a written, versioned workflow that encodes how your best people do the work, so every team inherits it instead of reinventing it. Second, a model router with local and smaller models serving the non-frontier work that fills most of a delivery day, and frontier models reserved for the moments where judgment is dense; most tasks are not frontier problems, and paying frontier prices for them is pure waste. Third, telemetry and analytics on every agent run: tokens, cost, latency, acceptance, so cost and accuracy stop being impressions and become curves you can steer. This is the machinery behind the numbers above, and it is why the harness is the rare investment that cuts cost and raises accuracy at the same time.&lt;/p&gt;

&lt;p&gt;None of this grows by itself, so staff it. Harness engineering is a central role, not a side duty: a small team of harness engineers, under the named owner, that treats the instruction standards, the hooks, the workflows, the router, and the telemetry as one product with a roadmap, and hardens what the first waves discover into the standard everyone else inherits. Make it a deliberate career path too: harness engineers make natural FDEs when the scale-out comes, and returning FDEs make the best harness engineers, because they have seen where the harness meets real teams and real deadlines.&lt;/p&gt;

&lt;p&gt;One component of the harness deserves its own attention because executives rarely hear about it: hooks, small deterministic scripts that fire at fixed points in the agent's lifecycle, before a tool call, after a file write, at session start. Hooks are how you make the probabilistic system behave predictably at the moments that matter. A hook can compress a ten-thousand-line test log into the hundred lines that matter before the model ever reads it, saving tokens and confusion. A hook can block a risky action, a database migration, a production credential, before it executes rather than after. A hook can index every artifact the agent produces so the next session starts informed. Guardrails written in prose are wishes; hooks are enforcement.&lt;/p&gt;

&lt;p&gt;One component deserves to be treated as a first-class company asset rather than a feature of the harness: the knowledge base. This is the shared brain every agent and every workflow draws on: standards, architectural decisions, incident history, domain meaning, the reasons behind the rules. Keep it unified but federated by domain, each area feeding and curating its own slice, because it compounds: the more it is fed, the better every workflow that reads it becomes, which makes it the one part of the stack that gains value with age instead of losing it. It must be engineered with precision, not volume, because as I argued in &lt;a href="https://dev.to/cleberdelima/your-coding-agents-are-drowning-in-context-you-pay-twice-in-tokens-and-in-precision-1pp7"&gt;Your Coding Agents Are Drowning in Context&lt;/a&gt;, badly curated context makes you pay twice, in tokens and in precision. Give it a named owner and domain stewards, measure it on coverage and retrieval quality, and defend it in the budget like the asset it is, because context is the moat no competitor can copy.&lt;/p&gt;

&lt;p&gt;Put those two ownership decisions together and you get the strategic line of this whole stage, the one I drew in &lt;a href="https://dev.to/cleberdelima/token-economics-why-your-ai-bill-is-a-capital-decision-not-a-cost-to-cut-2ch"&gt;Token Economics&lt;/a&gt;:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Rent the intelligence, own the memory and the harness.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Models will keep changing names, prices, and owners, and you should let them: intelligence is the part you buy in a falling market. The harness and the knowledge base are the parts you keep, and they are the only parts that compound.&lt;/p&gt;

&lt;p&gt;Stage two is also where the specification discipline from &lt;a href="https://dev.to/cleberdelima/from-user-stories-to-machine-ready-specs-why-your-requirements-process-is-breaking-down-in-the-age-3h96"&gt;From User Stories to Machine-Ready Specs&lt;/a&gt; becomes infrastructure rather than advice. User stories worked because human developers filled the gaps with implicit knowledge. Agents have no implicit knowledge; vague input produces wrong output at scale. Specs with explicit inputs, outputs, constraints, and acceptance criteria become versioned artifacts living next to the code, refined through pull requests by product managers and architects together. This is the first visible change in ways of working: product people start writing for two audiences, humans and machines, and the quality of their writing starts to bound the quality of the software.&lt;/p&gt;

&lt;p&gt;In practice this stage does not wait for stage one to complete; it starts almost together with it. The harness team should be standing by the time the first wave begins, because those waves are its laboratory: every friction an early team hits, every workaround a cornerstone invents, every objection a skeptic raises is raw material the harness team hardens into the standard. Keep that loop running, wave learnings in, harness versions out, until the harness is reliable enough to receive teams that were never hand-picked. That reliability, not a date on a plan, is what opens stage three. One rule about timing matters more than any other here: structure starts per team, not per company. The moment a team comes out of onboarding it should enter the harness work, while other teams are still in stage one. If structure waits for the whole organization to finish adopting, the vacuum fills itself: every team invents its own instruction files, its own conventions, its own defaults, and what should have been one harness becomes fifty private ones that later have to be reconciled at painful cost. &lt;/p&gt;

&lt;p&gt;The frustration arrives on schedule too, because teams working without structure get inconsistent results and rising bills, conclude the technology is overhyped, and disengage exactly when the discipline that would have fixed both was within reach. Stage zero's one-page standard is what makes this parallelism safe: teams can start early because they all start from the same page.&lt;/p&gt;

&lt;p&gt;The failure mode of skipping this stage deserves a name: the model tax. It is paying frontier-model prices to compensate for an environment you never built: noisy context inflating every token bill, endless tool churn as teams shop for the model that will finally fix what the environment keeps breaking, and every upgrade producing a new round of surprises because nothing around the model was stable. The money leaks twice, once in inflated inference costs and once in the engineer-weeks spent re-tuning after each release, and neither leak ever carries the harness's name in the budget, which is why this tax goes unnoticed for years. You are ready to advance when the harness has a named owner and a roadmap like any product, when a new team can onboard agents in days because the environment tells them how, and when your specs are versioned and machine-readable in at least the pilot domains.&lt;/p&gt;

&lt;p&gt;The asset of this stage, made concrete:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Artifacts:&lt;/strong&gt; an instruction file per repository, around a hundred lines that route to deeper docs on demand; hooks at the risky and noisy points of the agent lifecycle; reproducible environments with lockfiles; state files that survive sessions; explicit verification commands per repo; a versioned specs directory living beside the source; the codified workflow library per discipline; the model router configuration with local models serving the non-frontier work; telemetry and analytics on every agent run; the central harness engineering team, under a named owner, running it all as one product with a roadmap; and the knowledge base, unified, federated by domain, with a named owner and domain stewards.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Metrics:&lt;/strong&gt; days for a new team to onboard agents; tokens and cost per completed task; first-pass acceptance rate of agent output; clarification round-trips per feature; share of tasks served by local or smaller models without quality loss; knowledge-base coverage and retrieval quality; and an ablation score, the measured drop in success when you remove one harness subsystem at a time.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Impact:&lt;/strong&gt; the studies above put the range on the table: double-digit reductions in cost, latency, and tokens at constant quality from orchestration changes alone, and success-rate swings larger than any model swap. This is the stage where the two curves cross: cost per task falling while accuracy rises, with a growing share of work leaving the frontier tier and a knowledge base that makes every quarter's harness better than the last. If every model upgrade still produces a round of surprises, the asset is not there yet.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Stage 3: Scale. The asset is the practice, carried to everyone
&lt;/h2&gt;

&lt;p&gt;With proof in hand from the first waves and a harness hardened from their learnings, scaling stops being a leap of faith and becomes an industrial operation. Still, it is a different discipline, and the moment to start is when proof stops being the question.&lt;/p&gt;

&lt;p&gt;Be clear about what does not change: every scale-out wave still runs the same four moves as the first ones, its own demos on the domain's real work, its trained cohorts, its champions, its success stories feeding the next wave. What changes is the layer you add on top of that cycle.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The landing zone&lt;/strong&gt; is different now: every scale-out wave onboards into the harness from day one, with the workflows, the routing, and the guardrails already in place. That landing is why scale is its own stage, and why it comes after the harness, never before. Scaling on raw tools is one of the most expensive mistakes on this map, because it makes every team pay for adoption twice: they learn a raw-tool way of working now, and have to re-adapt to the harness when it finally arrives. &lt;/p&gt;

&lt;p&gt;This is where you borrow the strongest staffing pattern the AI industry itself uses: &lt;strong&gt;the forward deployed engineer&lt;/strong&gt;, FDE for short. Palantir coined the role years ago. Instead of shipping software and hoping customers would figure it out, it embedded engineers with the customer, inside the customer's real problems, and let what they learned flow back into the product. The frontier labs have adopted the same model for exactly the problem you face at this stage: OpenAI deploys its hardest enterprise products through forward deployed engineers rather than self-serve, with the lessons of every deployment feeding back into the product, and Anthropic is building forward-deployed teams to carry its models into enterprises. Read that carefully. &lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The companies with the best models on earth do not believe the model sells or installs itself. They believe adoption is an embedding problem, and they staff it accordingly.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The internal version draws from two pools. Take the architects and senior developers who came out of the first waves committed, together with the harness engineers who hardened those waves' learnings into the standard, and rotate them into the next domains as embedded FDEs: &lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;em&gt;Not trainers who deliver a course and leave, but builders who sit inside the receiving team for a wave, take that team's real backlog, and construct the first AI-assisted workflows with the team, on its code, its constraints, its deadlines.&lt;/em&gt;&lt;/strong&gt; &lt;/p&gt;

&lt;p&gt;The difference against classroom training (which still happens, as part of the four-move cycle every wave runs) is the difference this whole stage runs on: a course transfers knowledge, an embedded engineer transfers a working practice, and the practice survives because the team watched it being built on their own work, not on a demo repository. &lt;/p&gt;

&lt;p&gt;Give the rotation prime status: &lt;em&gt;a named role, a defined tour of one or two waves, and recognition that this is senior, key work, because it is&lt;/em&gt;. &lt;/p&gt;

&lt;p&gt;The people you want will not volunteer for something that looks like a side duty.&lt;/p&gt;

&lt;p&gt;What separates an FDE program from internal consulting is one thing: the return path. Everything an FDE builds in the field, the instruction files, the workflows, the answers to hard objections, flows back into the central harness and knowledge base, so every embedding makes the next one shorter and cheaper. Without the feedback loop and the two-way communication, FDEs are expensive consultants leaving local snowflakes behind; with it, FDEs become how the harness learns. And they are also a seed: an AI-fluent engineer embedded in a domain is exactly the shape stage six will scale beyond engineering, and your first FDEs are where those pods will come from.&lt;/p&gt;

&lt;p&gt;The failure mode of stalling before this stage is the pilot island: proof stranded in a handful of teams while the rest of the organization keeps working the old way. It is a quiet failure, because everything inside the island looks like success, the pilots deliver, the demos impress, the case studies write themselves. The money is what never arrives: the program has paid for the platform, the harness, and the first waves, but captures value only where the island's edge happens to fall, and when the board reads the pilot's numbers next to the organization's flat ones, it concludes the technology overpromised when the truth is that the practice was never distributed. &lt;/p&gt;

&lt;p&gt;You are ready to advance when every delivery team has been through a wave, when onboarding into the harness is simply how a new team starts, and when the practice holds its shape in teams nobody hand-picked.&lt;/p&gt;

&lt;p&gt;The asset of this stage, made concrete:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Artifacts:&lt;/strong&gt; the scale-out wave calendar covering every delivery team; the FDE rotation, first-wave cornerstones and harness engineers on named tours with a defined return path for what they build; per-domain onboarding kits generated from the harness; the champions network as a standing structure, not a launch-phase one.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Metrics:&lt;/strong&gt; coverage, the share of teams through a wave and working inside the harness; time and cost per wave, falling as embeddings compound; workflows and instruction files returned to the platform per embedding; variance of adoption depth across teams, shrinking as the practice standardizes; thirty-day retention holding even though rosters are no longer hand-picked.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Impact:&lt;/strong&gt; the proof of the first waves becomes the norm of the organization: every team lands on the same harness with a working practice from day one, enablement cost per wave falls while the harness gets richer, and the value case stops resting on pilots because the whole delivery organization runs the practice the pilots proved.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Stage 4: Verify. Trust becomes measurable, or autonomy stays a demo
&lt;/h2&gt;

&lt;p&gt;Every stage so far still assumes a human reads everything the machine produces. That assumption is the ceiling on the whole investment, and this stage is where you lift it, carefully. The instrument is the eval: a systematic test that measures how well an AI system performs on your specific tasks, made of inputs and success criteria encoded in grading logic. If your teams can already recite unit tests and CI, the translation is direct: evals are the tests, and the &lt;a href="https://dev.to/cleberdelima/evals-are-your-new-ci-the-acceptance-layer-for-work-your-team-didnt-write-20k7"&gt;eval gate is the CI&lt;/a&gt;, for behavior instead of code. A small, well-designed private eval tells you more about which model and which configuration to use than any public benchmark, because public benchmarks are saturated and contaminated, and your workload is not on them.&lt;/p&gt;

&lt;p&gt;Build the estate from what you already have. Real merged pull requests that fixed real issues become tasks. Review comments become learned rules. Incidents become permanent regression scenarios, which is how the operations loop closes: every production failure becomes a check that the same failure cannot ship twice. Grade with three families of graders, deterministic code checks where right and wrong are objective, model-based judges with calibrated rubrics where nuance is needed, and periodic human spot checks to keep the judges honest. Then wire the result where it can act: into the pipeline, so that a score drop blocks a deploy the same way a failing test does. The cautionary tale for why this must be automated comes from a published post-mortem in my research: a team shipped a support agent after three days of careful manual testing, and it quoted a deprecated refund policy in production for five days. No crash, no hallucination drama, just a plausible wrong answer that manual testing missed and an eval suite would have caught.&lt;/p&gt;

&lt;p&gt;Evals are also how you survive the &lt;a href="https://dev.to/cleberdelima/the-eval-gate-upgrading-models-without-breaking-your-agents-979"&gt;upgrade treadmill&lt;/a&gt;. Models now ship monthly, and each new one interacts differently with your standing instructions, a decay I described in &lt;a href="https://dev.to/cleberdelima/instruction-debt-your-prompts-are-aging-like-code-1li7"&gt;Instruction Debt&lt;/a&gt;. Without an eval gate, every model upgrade is a leap of faith; with one, it is a canary run and a diff. The same applies inside a model family: effort settings and automatic fallbacks mean the configuration serving your request may not be the one you selected, so log what actually served every request and test the configuration, not the brand name.&lt;/p&gt;

&lt;p&gt;This stage is where the autonomy ladder starts to climb, one earned rung at a time: suggest-only, then edit-with-review where a human approves diffs, then execute-in-sandbox, then autonomous-within-guardrails, where an agent may open pull requests but never merge to protected branches and never touch production credentials. The rung you are on should be a written policy per domain, not an accident of which team moves fastest. And the split that makes any of it safe: the agent that writes is never the agent that checks. A separate checker, ideally on a different model, with one standing instruction the writer never sees: treat the change as wrong until the spec and the existing tests prove otherwise, and never edit a test to make it pass.&lt;/p&gt;

&lt;p&gt;The failure mode of skipping this stage is the review wall, and it shows up in the delivery statistics. The &lt;a href="https://cloud.google.com/resources/content/2025-dora-ai-assisted-software-development-report" rel="noopener noreferrer"&gt;2025 DORA report&lt;/a&gt;, drawing on around five thousand practitioners, found AI adoption now improves throughput but still increases delivery instability: more speed, less stability, exactly what you would predict when generation scales and verification does not. Industry benchmark data from LinearB's 2026 report, which I could not link directly and cite from my research notes, shows AI-authored pull requests merging at less than half the rate of human ones and idling several times longer before a reviewer opens them: generation scaled, trust did not. The money goes into your most expensive people: senior engineers spending their days reading machine output line by line, pull requests waiting in queues while the agents that produced them sit idle, speed purchased at agent prices and delivered at the pace of human eyes. That is &lt;a href="https://dev.to/cleberdelima/the-velocity-trap-why-your-ai-productivity-gains-are-an-illusion-o6o"&gt;The Velocity Trap&lt;/a&gt; at organizational scale. You are ready to advance when eval scores gate deploys in the pipeline, when your checker agents reject real work every week and people fix the work rather than the checker, and when a model upgrade is a routine canary run rather than a project.&lt;/p&gt;

&lt;p&gt;The asset of this stage, made concrete:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Artifacts:&lt;/strong&gt; an eval suite of graded tasks built from your own merged PRs and incidents; written grader rubrics with anchored scoring examples; the CI configuration where a score drop blocks the deploy; a model-upgrade canary pipeline; the written autonomy ladder per domain; the checker agent's standing instruction, versioned like code.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Metrics:&lt;/strong&gt; eval pass rate per model and configuration; checker rejection rate, which must stay above zero every week; escaped-defect rate in AI-assisted lanes against your human-reviewed baseline; agreement between judge scores and human spot checks; days to certify a new model.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Impact:&lt;/strong&gt; verification stops consuming your senior engineers' reading hours and runs at machine speed instead; the instability the DORA data warns about gets pulled back toward your baseline while throughput keeps its gains; and a model release becomes an afternoon of canary runs instead of a quarter of anxiety. The review queue, the place where the gains of stages two and three go to die, stops being the bottleneck.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Stage 5: Automate. Loops do the work, the portfolio serves it
&lt;/h2&gt;

&lt;p&gt;Only now, with a harness that makes agents reliable and gates that make them trustworthy, does it pay to carefully reduce human intervention. This is loop engineering, the discipline I covered in depth in &lt;a href="https://dev.to/cleberdelima/loop-engineering-stop-prompting-your-agents-and-design-the-system-that-does-428p"&gt;Loop Engineering&lt;/a&gt;: stop being the person who prompts the agent, design the system that does it instead. A loop is a goal, a way to check progress against it, and a rule for when to stop. The build order is strict, and it is the clearest example of why this whole roadmap is sequenced: get one manual run reliable first, capture what made it work as a written skill, wrap the skill in a loop with a gate and a stop condition, and only then put it on a schedule. Scheduling something you never made reliable by hand is how loops blow up overnight. Notice that the order repeats the stages: the skill assumes a harness, the gate assumes evals, the schedule assumes both.&lt;/p&gt;

&lt;p&gt;Not everything deserves a loop. The filter is four conditions that must all hold: the task repeats at least weekly, something can automatically reject bad output, the agent can do the work end to end, and done is objective rather than a matter of taste. Miss one, and a good prompt is still the better tool. Where the conditions hold, the compounding is real: at the far end, Stripe's engineers have described an internal pipeline that merges more than a thousand machine-written pull requests a week through deterministic and LLM gates, a reported figure I keep on hand not as a target but as proof that the ceiling is nowhere near where most teams assume. The metric that keeps loops honest is cost per accepted change, logged weekly: below roughly half of proposals accepted, a loop costs more than it returns, and the fix is almost always the gate, not the prompt.&lt;/p&gt;

&lt;p&gt;Automation at this scale changes what you buy, which is why the model portfolio belongs in this stage. Loops consume tokens without fatigue, so the bill stops being a per-seat license and becomes metered infrastructure, the capital decision I described in &lt;a href="https://dev.to/cleberdelima/token-economics-why-your-ai-bill-is-a-capital-decision-not-a-cost-to-cut-2ch"&gt;Token Economics&lt;/a&gt;. Three moves keep it governed. &lt;/p&gt;

&lt;p&gt;First, put an LLM gateway in front of everything before the volume arrives: one chokepoint that meters spend with hard budgets, trips a circuit breaker when a loop iterates without progress, caches repeated queries, and logs which model actually served each request.&lt;/p&gt;

&lt;p&gt;Second, scale the routing the harness began: the model router from stage two becomes governed policy at the gateway, frontier models where the task is judgment-heavy, local and cheaper models where it is not, effort settings tuned per task rather than maximum by default. The routing policy is a living document your eval estate validates, because a cheaper model that fails the evals is not cheaper.&lt;/p&gt;

&lt;p&gt;Third, hold optionality: keep a self-hosted open-weight tier in the portfolio, but enter it with honest numbers. Independent cost analyses in my research notes put the break-even for self-hosting a large model at around eleven billion tokens a month at sustained utilization, figures you should re-run against your own vendor quotes, and at low utilization self-hosting costs more than premium APIs. For most organizations it is a late-stage move justified by volume, jurisdiction, or control rather than sticker price.&lt;/p&gt;

&lt;p&gt;Control deserves one concrete story, because it is the least obvious reason to hold portfolio optionality. When &lt;a href="https://huggingface.co/blog/security-incident-july-2026" rel="noopener noreferrer"&gt;Hugging Face investigated the breach of its own infrastructure&lt;/a&gt;, an attack that ran on an autonomous agent executing over 17,000 recorded actions, the hosted frontier models it reached for first refused much of the forensic work: their safety guardrails could not tell an incident responder from an attacker. The forensics ran instead on a self-hosted open-weight model on Hugging Face's own hardware, and &lt;a href="https://huggingface.co/blog/jeffboudier/open-model-cyber-defense" rel="noopener noreferrer"&gt;their engineers' published advice&lt;/a&gt; is now mine too: have a capable model you can run on your own infrastructure vetted and ready before an incident, because the day you need it is the wrong day to start.&lt;/p&gt;

&lt;p&gt;The failure mode of this stage run too early is automated chaos: loops amplifying whatever they sit on, unverified output merging at machine speed, and a token bill growing faster than the value it buys. AI multiplies everything, and if the foundations are wrong it multiplies chaos. You will see it first in the one metric this stage lives by: cost per accepted change rising while the acceptance rate sinks below the point where the loop returns less than it burns, which is the machine telling you the gates beneath it were never real. You are ready to advance when at least one loop has run for a quarter with flat or falling cost per accepted change, when the gateway gives finance a number it trusts, and when your routing policy survives contact with your evals.&lt;/p&gt;

&lt;p&gt;The asset of this stage, made concrete:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Artifacts:&lt;/strong&gt; a loop spec per automated workflow, its goal, its per-pass rule, its stop condition; the skill files that encode how the work is done; a state board that outlives every session; the gateway configuration with budgets, circuit breakers, and caching; a routing policy the evals validate; one vetted self-hostable model with a tested deployment recipe on the shelf.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Metrics:&lt;/strong&gt; cost per accepted change, weekly, per loop; acceptance rate, with roughly half as the floor below which a loop loses money; tokens per accepted change; cache hit rate at the gateway; spend against budget per team; the share of traffic served by each tier of the portfolio.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Impact:&lt;/strong&gt; delivery capacity stops scaling with human attention. Work is discovered, executed, and verified while nobody watches, at a unit cost finance can see and cap, and the same gateway that meters the spend is the kill switch when a loop misbehaves.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Stage 6: Reorganize. The operating model catches up with the machines
&lt;/h2&gt;

&lt;p&gt;Everything to this point can be built inside the engineering function. The last stage cannot, because the constraint has moved again. When generation is fast and verification is automated, the bottleneck becomes decision latency: how long intent waits between the person who holds it and the system that can execute it. Squeezing that latency means changing the organization, not the tooling, and this is where the series' longest-running arguments converge. The fixed-cadence ceremonies I questioned in &lt;a href="https://dev.to/cleberdelima/the-end-of-agile-when-the-assumptions-beneath-your-methodology-collapse-3g12"&gt;The End of Agile&lt;/a&gt; and the compressed delivery chain I described in &lt;a href="https://dev.to/cleberdelima/continuous-fluid-flow-how-ai-is-compressing-the-software-delivery-cycle-3f20"&gt;Continuous Fluid Flow&lt;/a&gt; stop being provocations and become the operating instructions.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The delivery unit that works at this stage is small and temporary: two or three people spanning product, design, and engineering, plus multiple agents, plus the harness, owning one outcome end to end from definition through production. &lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;What that squad runs is not a sprint. It is the AI-DLC cycle, three phases turning as one continuous loop: Inception, Construction, Operations. The names sound familiar; the content is not.&lt;/p&gt;

&lt;p&gt;Inception is where quality is born, and it is a working session, not a document trail. The squad sits together, product, design, engineering, with the AI in the room proposing: it decomposes the intent into units of work, drafts the stories and acceptance criteria, surfaces the edge cases and risks, asks the clarifying questions nobody wrote down. The humans correct, decide, and remove ambiguity, and the guardrails and success metrics are set here, before a line of code exists. The output is a machine-ready specification, versioned like code, plus the checks that will judge the result. A few focused hours of this compress what used to be weeks of sequential refinement, and every ambiguity removed here is cost avoided everywhere downstream, because an agent executes a vague spec at exactly the speed it executes a precise one.&lt;/p&gt;

&lt;p&gt;Construction is fast precisely because Inception was thorough. Agents generate code and tests inside the harness, the gates from stage four do the first pass of verification, and the humans validate live: architecture, trade-offs, the things a spec cannot fully carry. The unit of progress shrinks from the two-week sprint to cycles measured in hours or days, and something structural happens to the oldest split in engineering: building and reviewing merge into one act, because the senior engineer is no longer typing the code and then waiting for a colleague to read it, the senior engineer is verifying while the agents produce.&lt;/p&gt;

&lt;p&gt;Operations is where the loop closes instead of ending. Deployment is AI-assisted, observability watches at machine granularity, and agents run the first pass of incident analysis. But the point is what flows backward: incidents become guardrails, so the same failure cannot ship twice; learnings become specs, so the next Inception starts smarter; production telemetry and user behavior feed the backlog, so the loop restarts with better information than it began. Close the week the way I argued in &lt;a href="https://dev.to/cleberdelima/continuous-fluid-flow-how-ai-is-compressing-the-software-delivery-cycle-3f20"&gt;Continuous Fluid Flow&lt;/a&gt;: a stretch of early-life support where the squad watches what it shipped, then a compounding session that turns the cycle's learnings into reusable assets, new skills, better guardrails, improved specs, instead of into a list of opinions. Paradoxically, all this automation needs more synchronous human collaboration, not less: a few hours of the right people deciding together replaces weeks of asynchronous alignment, because the machine can execute at whatever speed the humans can decide.&lt;/p&gt;

&lt;p&gt;Now hold your ceremony calendar against that loop and watch what survives. The daily standup dies, because status lives on the state boards both humans and agents write to, and telemetry reports better than people do; the meetings that remain exist to decide, not to inform. Sprint planning becomes the Inception workshop, scheduled when an intent is ready rather than when the calendar says so. Story-point estimation goes with it, replaced by measured cycle-time distributions, the collapse I argued in &lt;a href="https://dev.to/cleberdelima/the-end-of-agile-when-the-assumptions-beneath-your-methodology-collapse-3g12"&gt;The End of Agile&lt;/a&gt;. The sprint review becomes outcome review at the gates: shipped increments and eval results, not slideware demos. And the retrospective becomes that compounding session, the one ceremony that gains weight in this world, because it is the only one whose output is an asset rather than a feeling. The test for any surviving meeting is simple: does it make a decision or produce an asset? If it does neither, it is queue time wearing a calendar invite.&lt;/p&gt;

&lt;p&gt;Product management changes the most, which is why product belongs inside this transformation from stage two onward, not as a late guest. When code stops being the constraint, knowing what to build becomes the expensive skill. The product role shifts from managing a backlog of stories toward curating intent: writing specifications precise enough to be contracts, deciding what is worth building at all, and validating outcomes rather than implementation. &lt;/p&gt;

&lt;p&gt;The strongest external evidence that pairing domain judgment with engineering fluency is the unlock comes from &lt;a href="https://x.com/praveenTweets/status/2074605343439810922" rel="noopener noreferrer"&gt;Uber's agentic pods&lt;/a&gt;: the company paired its most AI-proficient engineers with domain experts from business functions, gave each pod two weeks, and ran sixteen pods across sixteen functions in two months. Capital allocation analysis went from 15 hours to 30 minutes; financial pacing reports from two days to ten minutes; marketing quality assurance from two weeks to under an hour. The lesson Uber drew is the one I keep repeating: the biggest wins came from redesigning the whole workflow, not from automating a single task inside the old one. And that is also the proof that stage six is not an engineering stage at all: the same playbook, harness, gates, pods, runs in finance, operations, and marketing, which is where the second wave of value lives.&lt;/p&gt;

&lt;p&gt;The same blurring happens on the operations side, and it is just as deep. Operations stops being the downstream department that inherits what delivery throws over the wall, because in the AI-DLC cycle it is a phase of the same loop the squad owns. The SRE sensibility sits inside the squad; the toil that made "you build it, you run it" an empty slogan is now carried by agents watching telemetry, triaging incidents, and drafting the first root-cause analysis; and operational signals stop being complaints and become premium product input, feeding the next Inception directly. &lt;/p&gt;

&lt;p&gt;Quality moves the same way: the QA role stops being a manual pass at the end and becomes quality engineering, the discipline that designs and supervises the eval estate, the argument I made in &lt;a href="https://dev.to/cleberdelima/testing-reinvented-why-test-coverage-is-the-wrong-metric-31l3"&gt;Testing Reinvented&lt;/a&gt; carried to its organizational conclusion. Define, build, run, and learn stop being four departments' jobs and become four phases of one team's loop.&lt;/p&gt;

&lt;p&gt;Put all of it together and you can see the shape of a future product development organization.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Small, end-to-end squads formed around outcomes.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;They can be composed of two or three people plus an agent fleet, dissolving and reforming as intents change. A platform group running the shared machinery, the harness, the eval estate, the gateway, the knowledge base, as internal products with roadmaps. Product people embedded in squads as curators of intent and architects of specification. Quality engineers who own evals the way SREs own reliability. And far fewer pure coordination roles, because most of the layers in today's org chart exist to move information between people who do not sit together, and in this model the specs, the state boards, and the telemetry carry that information instead. &lt;/p&gt;

&lt;p&gt;The chart flattens, and the people who remain are the ones whose judgment cannot be automated. The early AI-native companies already show this silhouette, tiny teams with revenue per employee at multiples of the traditional software average, and while an enterprise will never copy a fifty-person startup, the direction of the silhouette is the same. &lt;/p&gt;

&lt;p&gt;Be honest about the path, though: you do not get there by announcing a reorganization. You get there team by team, the way this whole journey has moved, letting the squads that master the loop become the template the next ones copy.&lt;/p&gt;

&lt;p&gt;Stage six is also where you decide, deliberately, how far autonomy goes. There are now organizations running with no human in the code path at all: &lt;a href="https://factory.strongdm.ai/" rel="noopener noreferrer"&gt;StrongDM's software factory&lt;/a&gt; ships production security software under a charter that code must not be written or reviewed by humans, with three engineers writing specifications and evaluating outcomes against held-out scenarios the coding agents never see. I do not offer that as your target; for regulated domains it may never be permitted, and it is a bet most boards should not take today. I offer it as the far end of a dial you should set on purpose, domain by domain, in writing. A regulated core capped at supervised autonomy is a defensible policy. An unwritten "wherever the tools take us" is not a policy at all.&lt;/p&gt;

&lt;p&gt;The failure mode here is the faster typist: keeping the old organization while running new machinery inside it. AI grafted onto handoff-heavy, ceremony-paced delivery produces faster typing inside the same queue, and the queue still sets the speed: intent still waits days for a prioritization meeting, approvals still travel the same chain, and agents that could build in hours wait on decisions that take weeks. The money goes to capacity you bought and cannot use, the most expensive waste in this whole journey, because it looks like success on every dashboard except cycle time. The exit condition of stage six is the impact the whole journey was for: output rising faster than headcount, with people making the decisions and machines carrying the volume, measured honestly rather than assumed.&lt;/p&gt;

&lt;p&gt;The asset of this stage, made concrete:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Artifacts:&lt;/strong&gt; squad charters for the two-to-three-person pods and their end-to-end ownership; the AI-DLC cycle definition, Inception, Construction, Operations, as the squad's operating manual; the Inception workshop format and its spec templates; the ceremony map, what died, what replaced it, and what each replacement produces; the value dashboard where every initiative carries a baseline, an owner, and a KPI; the two-lens maturity heatmap; the autonomy ceiling register per domain; the kill-review calendar.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Metrics:&lt;/strong&gt; cycle time from idea to production; decision latency, how long intent waits for the person who can decide; the share of squads running the full cycle end to end, from Inception through Operations; output to headcount, delivery per engineer or revenue per employee; the share of initiatives with signed-off baselines; measured value against the fully loaded cost of the program.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Impact:&lt;/strong&gt; this is where the two numbers from the opening finally meet. Cycle-time compression of the 80-days-to-5 shape stops being one team's story and becomes the operating norm, workflow redesigns of the kind Uber's pods delivered turn days of work into minutes outside engineering too, and the board conversation changes from adoption percentages to a value line finance has signed.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Why the order is the product
&lt;/h2&gt;

&lt;p&gt;Read the stages backward and the dependency chain is visible. Reorganizing around AI speed presumes loops that deliver unattended. Loops presume gates that catch bad work automatically. Gates presume evals worth trusting. Evals presume a harness stable enough that results mean something. The scale-out presumes that harness carried to every team, and the harness presumes people who use the tools well. The people presume a stack somebody decided. Every arrow in that chain is a place I have watched organizations fail by skipping: loops on top of no evals produce confident garbage on a schedule; evals on top of no harness measure noise; harnesses handed to untrained teams become tools nobody uses, with a maintenance bill attached. Read the seven failure modes back in one line, tool sprawl, the license illusion, the model tax, the pilot island, the review wall, automated chaos, the faster typist, and they share one cause: a stage skipped. Boris Cherny, who created Claude Code, mapped &lt;a href="https://x.com/bcherny/status/2077929379661844559" rel="noopener noreferrer"&gt;the same territory from the practitioner side&lt;/a&gt; as five steps from gated access to AI-native operation, with the sharp observation that the step most teams botch is scaling agent count before the verification loop has earned trust. His ladder and mine are the same lesson at different altitudes: the constraint at each step is never tokens. It is bottlenecks and guardrails, found and built in order.&lt;/p&gt;

&lt;p&gt;The stages are not a calendar. A strong organization runs them overlapped, piloting stage-five loops in one domain while stage three is still rolling out in another, and it advances team by team rather than waiting for the whole company to clear a stage together: the onboarded teams move to structure while the rest are still adopting. What the order forbids is not overlap but inversion: no team should run a later-stage practice on top of an earlier-stage gap. That single rule is most of the governance you need.&lt;/p&gt;

&lt;h2&gt;
  
  
  Running it as a strategy, not a project
&lt;/h2&gt;

&lt;p&gt;The staged roadmap is the delivery half. The other half is how you run it from the top, and here is the shape of what this work has taught me, in principles rather than particulars.&lt;/p&gt;

&lt;p&gt;Run it as a named program with an executive sponsor, not a collection of initiatives inside the technology function. The numbers this transformation moves, revenue per employee, cost of delivery, customer experience, belong to the business, so accountability for them must sit where those numbers live, with the platform and capability work beneath. &lt;/p&gt;

&lt;p&gt;Give the program a guiding principle with teeth: every AI initiative earns its place with measurable business value, and no initiative launches without a baseline and the KPIs it is expected to move. &lt;/p&gt;

&lt;p&gt;Set targets the honest way around: measure baselines first, commit to directional ambitions after, and hold one commitment that needs no baseline at all, that the program must at minimum return its own cost in measured value. Then enforce the uncomfortable half of the discipline: when a use case does not move its KPI for a full quarter, it gets a mandatory review with three exits, kill it, pivot it once, or write down exactly why it stays. Zombie initiatives are how AI programs lose the board's trust.&lt;/p&gt;

&lt;p&gt;Measure maturity on both axes you met at the start of this article, quarterly for team adoption, semi-annually for organizational capability, and let the scores drive where the enablement effort goes next. Scoring teams on a five-level ladder sounds bureaucratic until you see what it changes: adoption stops being an opinion, lagging areas become visible while they are still cheap to fix, and the board conversation shifts from anecdotes to a trajectory. And write down what you will not do, in the architecture rather than in a policy document: the actions agents may never take autonomously, the domains where a human gate is permanent, the thresholds where automation stops and judgment begins. Those commitments are not brakes on the strategy. They are why you will still have a license to operate when a competitor's ungoverned agent makes the news.&lt;/p&gt;

&lt;h2&gt;
  
  
  The playbook: seven moves to start on Monday
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. Place yourself on the map with numbers, not opinions.&lt;/strong&gt; What: score your organization against the seven stages, and your teams against a five-level maturity ladder per work category, documentation, feature development, quality, operations. Why: every failed transformation I have examined misplaced itself at the start, usually one stage too generous. How: run the scoring workshop with three pilot teams first; the &lt;a href="https://www.sei.cmu.edu/library/ai-adoption-maturity-model/" rel="noopener noreferrer"&gt;SEI and Accenture model&lt;/a&gt; gives you an externally benchmarkable rubric to adapt. Pitfall: letting teams self-report; scores inflate a full level. Signal: one honest heatmap the whole leadership team accepts without re-arguing it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Stand up the adoption engine as a permanent cycle.&lt;/strong&gt; What: domain-tailored demos, role-based training in hours not minutes, embedded champions, success stories feeding the next domain. Why: behavior change is 80 percent of the outcome and it decays without reinforcement. How: pick the two most willing domains, run the full cycle, publish the wins internally, repeat quarterly; make AI literacy part of week-one onboarding. Pitfall: training everyone on prompting and no one on judgment. Signal: weekly active use rising without mandates.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Appoint a harness owner and ship version one.&lt;/strong&gt; What: one accountable owner, product not project, for the instruction files, tools, hooks, state, and verification commands your agents run inside. Why: the harness moves outcomes more than the model does, and unowned harnesses rot. How: start with a hundred-line instruction file that routes to deeper docs, explicit verification commands per repo, and two hooks, one that compresses noisy logs before the model reads them, one that blocks the actions you never want an agent to take. Pitfall: writing a manual when the agent needs a map. Signal: a new team onboards agents in days, and when output disappoints, people check the harness before blaming the model.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Build the eval estate and wire it into CI.&lt;/strong&gt; What: a private suite of graded tasks built from your own merged PRs, review comments, and incidents, gating deploys and model upgrades. Why: it converts trust from a feeling into a number, and it is the asset that makes every later stage safe. How: start with twenty tasks in one domain, grade with deterministic checks plus one calibrated LLM judge, block the pipeline when scores drop, and run every model upgrade as a canary against the suite. Pitfall: building evals once and letting them go stale; they are a living artifact. Signal: a checker rejects real work every week, and people fix the work.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5. Put the gateway in before the loops.&lt;/strong&gt; What: one chokepoint for all model traffic: budgets, circuit breakers, caching, full logs including which model actually served each request. Why: metered costs arrive with automation, and retrofitting governance onto a running fleet is painful and expensive. How: deploy a gateway with hard per-team budgets and a circuit breaker that halts any loop iterating without repo change; add a routing policy validated by your evals. Pitfall: treating self-hosting as a cost play at low volume; it only pays at sustained scale or for control reasons. Signal: finance trusts the number, and a runaway loop stops itself before anyone notices.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;6. Pilot one loop and one pod, in parallel.&lt;/strong&gt; What: one unattended loop on verification-friendly work, nightly CI triage, dependency updates, documentation drift, and one two-week pod pairing your most AI-fluent engineer with a domain expert outside engineering. Why: the loop proves stage five mechanics; the pod proves the playbook travels beyond engineering, and each generates the next round of internal believers. How: loop follows the strict build order, manual run, skill, gate, schedule, and reports cost per accepted change weekly; pod follows Uber's shape, understand the work first, build alongside the person who does it, validate with others who do the same job. Pitfall: piloting where failure is expensive. Signal: the loop's acceptance rate holds above half; the pod ships something the domain expert actually uses after the pod ends.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;7. Make value the gate for everything.&lt;/strong&gt; What: a value dashboard where every initiative carries a baseline, an owner, and the KPI it claims to move, reviewed on a fixed cadence with the kill rule in force. Why: it is the difference between a portfolio and a pile. How: baseline before launch, validate with holdouts where the money is material, and count only what finance signs off. Pitfall: counting activity, tokens consumed, agents built, PRs assisted, as value. Signal: at least one initiative killed or pivoted per quarter; paradoxically, that is the sign the discipline is real.&lt;/p&gt;

&lt;h2&gt;
  
  
  The strategic takeaway
&lt;/h2&gt;

&lt;p&gt;What separates adoption from impact is everything the machines cannot bring with them, and that is exactly what compounds in this journey: the harness, the eval estate, the knowledge layer, the written skills, the redesigned ways of working. Every one of those assets makes the next stage cheaper and every model upgrade more valuable, which is why organizations that build them pull ahead quietly for a while, and then the gap is suddenly obvious. What plateaus is everything most programs celebrate: licenses, adoption charts, prompting skill, review capacity. Those scale with headcount and attention, and headcount and attention are exactly what this technology stops rewarding. The cost of the gap is not visible this quarter, because stage-one organizations and stage-five organizations both have impressive demos. It becomes visible when the compounding curves separate, and by then the distance is not a budget line. It is years of accumulated assets the lagging organization has to build from zero, against a competitor whose machines are already feeding their own improvement.&lt;/p&gt;

&lt;p&gt;And if you keep only one funding rule from all of this, keep this one: fund people and the harness first, verification second, automation third, and the heavy options last, then hold the whole program to a single quarterly gauge, measured value against fully loaded cost. Everything else in this article is detail on top of that sentence.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;So place your organization on the map, honestly. Which stage are you in, which stage does your board believe you are in, and what would it cost to make those two answers the same? &lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Take this map to your next leadership meeting and ask those questions out loud. The discussion that follows will tell you more about your real stage than any dashboard, and whatever it reveals, the next move is already on the map.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>softwareengineering</category>
      <category>softwaredevelopment</category>
      <category>agents</category>
    </item>
    <item>
      <title>The Eval Gate: Upgrading Models Without Breaking Your Agents</title>
      <dc:creator>Cleber de Lima</dc:creator>
      <pubDate>Tue, 11 Aug 2026 10:43:40 +0000</pubDate>
      <link>https://dev.to/cleberdelima/the-eval-gate-upgrading-models-without-breaking-your-agents-979</link>
      <guid>https://dev.to/cleberdelima/the-eval-gate-upgrading-models-without-breaking-your-agents-979</guid>
      <description>&lt;p&gt;Somewhere in your stack, a model already has a retirement date. Anthropic now runs &lt;a href="https://platform.claude.com/docs/en/about-claude/models/migration-guide" rel="noopener noreferrer"&gt;a fixed 60-day window from deprecation to retirement&lt;/a&gt;: Opus 4.1, deprecated June 5, 2026, retired August 5. OpenAI &lt;a href="https://help.openai.com/en/articles/20001051-retiring-gpt-4o-and-other-chatgpt-models" rel="noopener noreferrer"&gt;fully retired GPT-4o in April 2026&lt;/a&gt;. The ground under your production agents moves on the vendor's schedule, not yours.&lt;/p&gt;

&lt;p&gt;Most teams treat the swap as maintenance: flip the alias, watch the dashboards, move on. But the dashboards watch error rate, latency, and throughput, and a model change can leave all three flat while it rewires how your agents behave. The API returns HTTP 200, the responses read fine, and the regression ships anyway.&lt;/p&gt;

&lt;p&gt;So the question is not whether you will change models; the retirement clock has answered that. It is whether each change flows through a standing, evidence-driven gate, the way &lt;a href="https://dev.to/cleberdelima/evals-are-your-new-ci-the-acceptance-layer-for-work-your-team-didnt-write-20k7"&gt;code changes flow through CI&lt;/a&gt;, or keeps arriving as a bet.&lt;/p&gt;

&lt;p&gt;I learned to distrust the upgrade instinct inside our own AI-DLC harness at Betsson. When agent output disappointed, the reflex in the room was always the same: swap the model, something newer must do better. The swaps barely moved our metrics. What moved them was &lt;a href="https://dev.to/cleberdelima/loop-engineering-stop-prompting-your-agents-and-design-the-system-that-does-428p"&gt;re-engineering the loop around the model&lt;/a&gt;, the maker-checker split, the gates that could fail work automatically, the standing instruction our checker agents carried that writers never saw. That discipline, more than any model upgrade, is what let squads compress work and generate real gains. &lt;/p&gt;

&lt;p&gt;So I hold both beliefs at once: the model is rarely the fix, and yet the model underneath you will keep changing whether you are ready or not. This article is about being ready.&lt;/p&gt;

&lt;h2&gt;
  
  
  An Upgrade Is a Hypothesis, Not a Fact
&lt;/h2&gt;

&lt;p&gt;The evidence that could reset your intuition came from Microsoft's own developer advocacy team. In July 2026, Waldek Mastykarz &lt;a href="https://developer.microsoft.com/blog/not-all-model-upgrades-are-upgrades" rel="noopener noreferrer"&gt;ran 150 agent tasks&lt;/a&gt; across 15 scenarios comparing Claude Sonnet 4.6 with its newer, 33 percent cheaper successor, Sonnet 5. On architecture tasks, the older model matched or beat the newer one on quality in 8 of 9 comparable scenarios. On a code-upgrade task class, the newer model passed 100 percent of runs against the older model's 60, because it followed a versioned instruction the older model kept overriding. And the "cheaper" model consumed up to 47 times more tokens on identical prompts, landing at 3.7 times more expensive per run on one task class: $2.01 against $0.55.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The good thing is that he now knows about it, do you ??&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;When the hypothesis is wrong, the failure is silent. See what happend with the OpenAI's April 2025 GPT-4o update, shipped to ChatGPT's &lt;a href="https://openai.com/index/sycophancy-in-gpt-4o/" rel="noopener noreferrer"&gt;500 million weekly users&lt;/a&gt; after offline evals and A/B tests both looked good. &lt;/p&gt;

&lt;p&gt;A changed mix of reward signals, one of them built on thumbs-up data, had weakened the signal holding sycophancy in check, and the model began validating doubts, fueling anger, and applauding bad ideas while every dashboard stayed green. &lt;/p&gt;

&lt;p&gt;Three days of social-media backlash later, OpenAI rolled it back, a rollback that itself took 24 hours. &lt;a href="https://openai.com/index/expanding-on-sycophancy/" rel="noopener noreferrer"&gt;The postmortem&lt;/a&gt; is blunt: expert testers had felt the model was "slightly off," no deployment eval tracked sycophancy, and shipping on the metrics anyway was "the wrong call." As &lt;a href="https://tianpan.co/blog/2026-04-17-prompt-canaries-deployment-llm-production" rel="noopener noreferrer"&gt;Tian Pan&lt;/a&gt; put it, Twitter was the production alerting system. &lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Your observability stack watches for failures that announce themselves. Behavioral regressions do not.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The Gate: CI's Shape, Adapted for Non-Determinism
&lt;/h2&gt;

&lt;p&gt;In a previous article, &lt;a href="https://dev.to/cleberdelima/evals-are-your-new-ci-the-acceptance-layer-for-work-your-team-didnt-write-20k7"&gt;Evals Are Your New CI&lt;/a&gt;, I argued that evals are the acceptance layer for agent-produced work. This one is about the pipeline that runs that estate every time the model underneath your agents changes. It has four parts, none of them complicated: CI/CD's shape, adapted for a system whose regressions are behavioral.&lt;/p&gt;

&lt;h3&gt;
  
  
  1 - &lt;strong&gt;Crate a Deployment manifest.&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;A model upgrade is never a single-variable change. Behavior is the combined result of several layers that move on their own, and if you swap the model without recording the state of the rest, you lose the ability to say later which one actually changed. &lt;a href="https://www.buildmvpfast.com/blog/agent-versioning-rollback-production-ai-update-zero-downtime-2026" rel="noopener noreferrer"&gt;BuildMVPFast's test&lt;/a&gt; is the one I would put on a wall: if you cannot recreate exact behavior from a manifest, you do not have versioning, you have a label. The manifest is the cheapest artifact in the whole pipeline, a few lines of config, and it is what makes every later step work, because the gate does its job by diffing the candidate's manifest against the stable one.&lt;/p&gt;

&lt;p&gt;The manifest need to have five layers, because each one can rewire behavior on its own.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt; The orchestration and routing logic wrapped around the model, kept in ordinary Git like the rest of the service. Most teams already version this; the manifest's only job here is to record which commit was live for a given run.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Prompts.&lt;/strong&gt; Every template the agent uses, system prompts and few-shot examples included, kept in their own registry with their own release cycle rather than inlined in code where a "small copy tweak" ships unreviewed. Prompts are code, and when an agent starts misbehaving the fix is almost never to hot-swap the prompt live; that is cowboy coding, not versioning.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Model.&lt;/strong&gt; Pin the dated identifier, never the alias. &lt;code&gt;claude-sonnet-4-6-20250514&lt;/code&gt; and &lt;code&gt;claude-sonnet-4-6-20250620&lt;/code&gt; answer to the same friendly name and do not behave the same way, and an alias rebinds under you on the vendor's schedule, which is the precise failure this article exists to prevent. The dated string is the only version of "which model were we running" that survives an audit.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Tool contracts.&lt;/strong&gt; The schemas, endpoints, and response formats of every tool the agent can call. Teams forget this layer because it lives outside their repo: when a third-party API like Stripe quietly changes a field, your agent's behavior changed even though not one line of your code did. Hash the tool schemas so a contract drift shows up as a manifest diff instead of a production surprise.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Retrieval index.&lt;/strong&gt; The version, or at minimum the timestamp, of whatever corpus the agent retrieves from. A prompt validated against last month's index can fail against this month's drifted one with the model held perfectly constant. &lt;a href="https://tianpan.co/blog/2026-04-17-prompt-canaries-deployment-llm-production" rel="noopener noreferrer"&gt;Tian Pan's minimal manifest&lt;/a&gt; is four fields for exactly this reason: &lt;code&gt;prompt_version&lt;/code&gt;, &lt;code&gt;model&lt;/code&gt;, &lt;code&gt;rag_index&lt;/code&gt; timestamp, and &lt;code&gt;tool_schema_hash&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Kept together, the manifest is small enough to sit at the head of every eval run and every deploy:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;prompt_version&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;    &lt;span class="s"&gt;checkout-agent@v14&lt;/span&gt;
&lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;             &lt;span class="s"&gt;claude-sonnet-4-6-20250514&lt;/span&gt;
&lt;span class="na"&gt;tool_schema_hash&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;  &lt;span class="s"&gt;3f9a1c&lt;/span&gt;            &lt;span class="c1"&gt;# payments, inventory, shipping&lt;/span&gt;
&lt;span class="na"&gt;rag_index&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;         &lt;span class="s"&gt;catalog-2026-08-01T02:00Z&lt;/span&gt;
&lt;span class="na"&gt;code_sha&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;          &lt;span class="s"&gt;a17be92&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is the whole payoff. When the gate flags a regression, you are not re-auditing the system blindly; you diff the two manifests, find the one field that moved, and start there. This extends &lt;a href="https://dev.to/cleberdelima/instruction-debt-your-prompts-are-aging-like-code-1li7"&gt;Instruction Debt (O-35)&lt;/a&gt; one layer down: the manifest tells you which layer changed, so you prune the stale instructions in that layer instead of re-auditing everything. Pin nothing and every regression is a guessing game; pin all five and it is a diff.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;2- Run the candidate in shadow, then widen in stages.&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;Three vendors who do not cite each other, &lt;a href="https://futureagi.com/blog/agent-rollout-strategies-2026/" rel="noopener noreferrer"&gt;Future AGI&lt;/a&gt;, &lt;a href="https://www.buildmvpfast.com/blog/agent-versioning-rollback-production-ai-update-zero-downtime-2026" rel="noopener noreferrer"&gt;BuildMVPFast&lt;/a&gt;, and &lt;a href="https://www.deepinspect.ai/blog/ai-gateway-canary-deployment" rel="noopener noreferrer"&gt;DeepInspect&lt;/a&gt;, landed this year on the same funnel: shadow traffic first, then progressively wider live slices, each stage with pre-registered rollback triggers. Read them as three calibrations of one pattern, not one standard; the shape holds, and the specific numbers are theirs to defend and yours to tune.&lt;/p&gt;

&lt;p&gt;Start with how the candidate gets exposed at all. There are four routing patterns, and they trade cost against how much signal they buy:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Pattern&lt;/th&gt;
&lt;th&gt;Who actually sees the candidate&lt;/th&gt;
&lt;th&gt;Cost overhead&lt;/th&gt;
&lt;th&gt;What it answers&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Shadow&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;No one; production serves, the candidate is scored offline&lt;/td&gt;
&lt;td&gt;1x (full duplication)&lt;/td&gt;
&lt;td&gt;Does the candidate behave reasonably on the real distribution?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Mirror&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;No one; a sampled subset is duplicated&lt;/td&gt;
&lt;td&gt;The sample rate&lt;/td&gt;
&lt;td&gt;The same question, cost-bounded&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Canary&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;A stratified live user slice&lt;/td&gt;
&lt;td&gt;The slice size&lt;/td&gt;
&lt;td&gt;Is it at least as good with real users in the loop?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Race&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Whichever candidate wins the latency race for that request&lt;/td&gt;
&lt;td&gt;Nx (parallel fan-out)&lt;/td&gt;
&lt;td&gt;Can two candidates clear a hard latency SLO together?&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Sequence them shadow to mirror to canary to full, and keep race for the rare case where a strict latency SLO is what matters most. Shadow is the stage everyone skips and no one should: both versions get the same request, the user sees the old version's response, the candidate's is logged and diffed, at zero user risk. Run it long enough to gather side-by-side comparisons of tool-call patterns, response lengths, and latency distributions before a single user is exposed.&lt;/p&gt;

&lt;p&gt;Then walk the ladder. &lt;a href="https://www.buildmvpfast.com/blog/agent-versioning-rollback-production-ai-update-zero-downtime-2026" rel="noopener noreferrer"&gt;BuildMVPFast's&lt;/a&gt; calibration is a good default to copy and adjust:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Stage&lt;/th&gt;
&lt;th&gt;Traffic to candidate&lt;/th&gt;
&lt;th&gt;Hold for&lt;/th&gt;
&lt;th&gt;Promote when&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Shadow&lt;/td&gt;
&lt;td&gt;0% (mirrored, discarded)&lt;/td&gt;
&lt;td&gt;24h+&lt;/td&gt;
&lt;td&gt;No errors in shadow responses&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Canary&lt;/td&gt;
&lt;td&gt;5%&lt;/td&gt;
&lt;td&gt;4-6h&lt;/td&gt;
&lt;td&gt;Error rate &amp;lt;2%, p99 latency &amp;lt;8s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Expanded canary&lt;/td&gt;
&lt;td&gt;25%&lt;/td&gt;
&lt;td&gt;12-24h&lt;/td&gt;
&lt;td&gt;Tool-call success rate &amp;gt;95%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Majority&lt;/td&gt;
&lt;td&gt;50%&lt;/td&gt;
&lt;td&gt;24-48h&lt;/td&gt;
&lt;td&gt;User-satisfaction signals stable&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Full&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;td&gt;steady state&lt;/td&gt;
&lt;td&gt;Every prior gate held&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Future AGI arrives at a structurally identical ladder with its own thresholds, gating each step on a Welch's t-test (p &amp;gt; 0.05 against a seven-day baseline) instead of a flat error rate, the same idea, just with stronger statistics. Two vendors, two sets of numbers, one ladder; pick the calibration that fits your traffic and revisit it as you learn.&lt;/p&gt;

&lt;p&gt;Two wrinkles that generic canary tooling gets wrong on agents specifically. First, split traffic by identity, not by request. An agent loop making five calls in one session must not get the new model on call two and the old one on call four, or its planning context breaks mid-loop; route the whole identity to one version for the canary's duration. Second, hold each stage through a full traffic cycle, 24 to 72 hours, not whatever window is convenient; a two-hour Tuesday canary has never seen your weekend.&lt;/p&gt;

&lt;p&gt;Watch the right signals while it runs, because the failures that matter here leave no HTTP error. &lt;a href="https://tianpan.co/blog/2026-04-17-prompt-canaries-deployment-llm-production" rel="noopener noreferrer"&gt;Tian Pan's&lt;/a&gt; behavioral stack sorts into three kinds. &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Distribution-shift signals:&lt;/strong&gt; output-length percentiles (p50/p95), sentiment across a sample (the signal that would have caught the sycophancy case), and refusal rate, which spikes in either direction when a new model's safety tuning collides with your existing prompts. &lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Task-outcome signals:&lt;/strong&gt; session-abandonment rate, in-session re-query rate (the same question asked twice is a clean proxy for "the first answer did not help," and needs no explicit feedback), and the edit-to-accept ratio on any draft-then-human-edits workflow. &lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Semantic-drift signals:&lt;/strong&gt; embedding cosine similarity against a golden response set, and an LLM-judge score against a reference, at the cost of one extra inference per sampled request. &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;You do not run these on all traffic; full semantic scoring on everything is too expensive, and sampling 1 to 5% gives enough statistical power to catch a real shift within hours. Budget for the fact that these windows run longer than infra alarms: a latency regression shows in minutes, but a tone shift needs enough samples for confidence, which at 5% routing is closer to 12 to 24 hours.&lt;/p&gt;

&lt;p&gt;Finally, gate on the distribution, never the mean. Future AGI's sharpest warning: a canary held at 1% for forty minutes looked clean on mean Groundedness, 0.91, while the real regression sat in a single sub-route at 0.62; the gate fired green because it answered the wrong question. Mean-only gating is their number-one anti-pattern from incident postmortems, and it is worst when a few big tenants drive most of your traffic, where a blind 5% canary quietly sends most of the candidate volume to your highest-value tenants and the average looks fine while the segment paying the bills is failing. Stratify by tenant or tag, start the candidate on the lowest-failure-cost segment you have (an internal tag is ideal), and widen only after the rubric holds on the slices you actually care about.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;3 - Make the decision deterministic: PROMOTE, HOLD, or ROLLBACK.&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;The sharpest data point in this article's research is from a production agent fleet's release-gate case study: agreement between human reviewers reading agent output and an automated structural gate measured &lt;a href="https://vadim.blog/evidence-driven-release-gates-llm-sales-agents/" rel="noopener noreferrer"&gt;kappa 0.13, barely above chance&lt;/a&gt;, because latency violations and routing errors leave no trace in the response text a human reads. One system, one case study, so treat the direction as the lesson, not the number. &lt;/p&gt;

&lt;p&gt;The gate that works is a deterministic function over a window of eval verdicts, not a single run and not another model's opinion: any safety violation vetoes promotion regardless of success rate, regression against the prior version's baseline forces rollback, and thin evidence holds rather than guesses. Three states matter because they map to three different operator actions, and the common mistake, collapsing HOLD into ROLLBACK, trains teams to distrust the gate. For calibration, one team's published numbers, not a standard: a 0.80 success floor, rollback on a 0.10 drop against the prior version's baseline, a minimum of five verdicts before any promotion, and 2 rollback-grade builds caught across 38 evaluation runs.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;4- Plan the rollback as a checklist, not a revert.&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;A rollback that only flips the traffic percentage back leaves regressions ghost-serving for hours, because reverting the deploy does not touch the state the bad candidate left behind. &lt;a href="https://futureagi.com/blog/agent-rollout-strategies-2026/" rel="noopener noreferrer"&gt;Future AGI&lt;/a&gt; spells out the steps the deploy alone skips: set the candidate's traffic to 0, which is an instant kill switch needing no redeploy; invalidate the cache namespace tagged with the candidate version, because semantic caches key on prompt hash, not agent version, and will keep serving the bad output; flush any downstream store that snapshotted that output; and bump the rubric version, so traffic scored after the rollback is measured against the right floor instead of the candidate's. Four steps, none of which "revert the deploy" performs for you.&lt;/p&gt;

&lt;p&gt;Wire the triggers before any traffic flips, as config, not a dashboard someone half-watches. &lt;a href="https://www.buildmvpfast.com/blog/agent-versioning-rollback-production-ai-update-zero-downtime-2026" rel="noopener noreferrer"&gt;BuildMVPFast's&lt;/a&gt; default triggers are a sane starting set: error rate above 5% on a 2-minute rolling window, p99 latency above 10s over 5 minutes, tool-call failure rate at twice baseline over 5 minutes, response-format violations above 3% over 10 minutes, and a hallucination score past a per-use-case threshold over 15 minutes. Any one of them fires the revert on the next request, not after a thread forms in Slack.&lt;/p&gt;

&lt;p&gt;Rolling back is not automatically the right move. Across production incidents the rough split is 60% rollback, 40% forward-fix, and it shifts toward forward-fix as your testing and observability mature. Roll back when the root cause is unclear and the blast radius is growing, when state is being corrupted, when several metrics degrade at once, or when it is 3am and no one is awake to diagnose. Push a fix forward when the cause is obvious and the fix is small, when a rollback would itself break already-migrated state, or when the damage is contained to a narrow subset you can flag off. The one thing you never do is decide this live, mid-incident, with no rule agreed in advance.&lt;/p&gt;

&lt;p&gt;Agent state is what makes an agent rollback harder than a stateless one, so design for the version boundary before you cross it: forward-compatible schemas only, meaning new fields are nullable and additive and nothing is renamed or removed inside the rollback window; a schema_version tag on every persisted record so an older version can safely ignore fields it does not understand; and no deleting a field until at least two deployment cycles after you introduced it. For long-running or multi-day tasks there is no clean answer, only three honest ones: let in-flight tasks finish on the old version via sticky routing, checkpoint and resume on state you designed to be version-agnostic from day one, or fail gracefully and restart the task, telling the user. Sometimes the last is the right one.&lt;/p&gt;

&lt;p&gt;And test the rollback before you need it. The discipline worth stealing is a monthly drill that deploys a deliberately broken version to staging and confirms the triggers actually fire and the state survives the revert intact. A rollback path you have never tested is a hypothesis, the same trap as the upgrade itself.&lt;/p&gt;

&lt;p&gt;In practice, all of this lives at the LLM gateway, the same control point this series has argued should own routing and budgets, and the same gate answers the portfolio question in reverse: a downgrade candidate needs identical discipline, because false economy is just regression with better marketing. Microsoft's 3.7x-more-expensive "cheaper" model is what skipping that check costs.&lt;/p&gt;

&lt;h2&gt;
  
  
  Before You Build Any of It, Ask the Other Question
&lt;/h2&gt;

&lt;p&gt;The counterweight comes from our own experience more than from the research. When a new model regresses your agents, the gate's job is not only to catch it. It is to force the question the upgrade instinct skips: is this a model problem, or a harness problem the new model just exposed? Anthropic engineers have described an experiment where the same model produced a broken $9 outcome and a working $200 outcome depending entirely on the harness around it. At Betsson, the fixes that lasted were nearly always on our side of the API: a stale guardrail written for a weaker model, a gate asking a binary question where a rubric was needed, context fed wrong. A gate that only ever answers "which model" will approve expensive swaps that were never the bottleneck. Sometimes the right output of an upgrade evaluation is a harness fix and no upgrade at all.&lt;/p&gt;

&lt;p&gt;There is also a cheap first step before the machinery: on each release, read the vendor's prompting guide for the new model, or feed it to the model and have it propose prompt updates. It costs an afternoon and catches the class of breakage that needs no pipeline at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the Gate Costs, and When to Kill It
&lt;/h2&gt;

&lt;p&gt;Nobody has published a credible end-to-end cost figure for running an upgrade gate, so I will not pretend to have one. The visible line items: shadow traffic roughly doubles spend on the routes you mirror, and every rollout stage re-pays a prompt-cache warm-up, because caches are model-scoped. Fund first: gateway-level traffic sampling and a shadow lane for your single most valuable agent workflow, reusing the eval estate the previous article had you build. Fund later: fleet-wide coverage and automated rollback wiring. Run the gate advisory for its first quarter, logging verdicts without blocking, and measure agreement between its verdicts and your senior reviewers' judgment on the same candidates. The kill metric: if after a quarter the gate has never disagreed with the alias-flip decision you would have made anyway, either your workloads sit far from the jagged edge or the gate is measuring the wrong things. Find out which before you scale it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start, Stop, Continue
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Executives&lt;/strong&gt;&lt;br&gt;
Start: requiring a gate verdict, with evidence, before any production model change, upgrades and downgrades alike; asking who owns the gate by name; the eval engineer is now &lt;a href="https://jobsbyculture.com/blog/ai-evals-engineer-career-guide-2026" rel="noopener noreferrer"&gt;a real hiring market&lt;/a&gt;, where "describe the eval you would run before flipping the switch" is an actual interview question.&lt;br&gt;
Stop: treating vendor deprecation notices as IT housekeeping; they are 60-day countdowns on production behavior. Stop approving model swaps justified by benchmarks alone.&lt;br&gt;
Continue: holding the harness accountable before the model, and funding it that way.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Engineers&lt;/strong&gt;&lt;br&gt;
Start: pinning dated model identifiers and writing the four-line deployment manifest this week; diffing tool-call traces between stable and candidate, not just final outputs.&lt;br&gt;
Stop: splitting canary traffic at the request level under agent loops; collapsing HOLD into ROLLBACK; trusting a revert without invalidating caches.&lt;br&gt;
Continue: reading the model's release notes and prompting guide on day one; feeding every gate-caught regression back into the eval estate.&lt;/p&gt;

&lt;h2&gt;
  
  
  Strategic Takeaway
&lt;/h2&gt;

&lt;p&gt;The gate compounds. Every upgrade it evaluates improves the eval estate, sharpens the baselines, and makes the next model change cheaper to absorb, which matters because there will always be a next one: the retirement clocks never stop. Blind upgrading and defensive pinning both plateau, and they fail the same way, as an unmeasured bet that today's behavior survives tomorrow's model. The organizations that absorb model churn as routine, gated change will treat every vendor release as a free option to get better. Everyone else will treat it as a threat.&lt;/p&gt;

&lt;p&gt;Evals Are Your New CI closed with a test: if a better model shipped tomorrow, could you tell by the end of the day, with a number, whether your most important agent workflow got better or worse? Here is the harder version: a model you depend on gets a retirement date today, sixty days out. Is the migration a project with a war room, or a pipeline run with a verdict at the end? Tell me where this breaks in your world; if you have upgraded agents in production without a gate and it went fine, I want to hear that too. Send this to whoever owns your model roster and ask them which question their last upgrade answered: did it pass, or did it just not fail loudly yet?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>programming</category>
      <category>softwaredevelopment</category>
    </item>
    <item>
      <title>Your pipeline is green. Every check passed. And the agent that wrote the code wrote the tests.
Green means compiled. Not correct.
Evals are your new CI: the acceptance layer for work your team did not write.</title>
      <dc:creator>Cleber de Lima</dc:creator>
      <pubDate>Tue, 04 Aug 2026 15:50:00 +0000</pubDate>
      <link>https://dev.to/cleberdelima/your-pipeline-is-green-every-check-passed-and-the-agent-that-wrote-the-code-wrote-the-tests-1594</link>
      <guid>https://dev.to/cleberdelima/your-pipeline-is-green-every-check-passed-and-the-agent-that-wrote-the-code-wrote-the-tests-1594</guid>
      <description>&lt;div class="ltag__link--embedded"&gt;
  &lt;div class="crayons-story "&gt;
  &lt;a href="https://dev.to/cleberdelima/evals-are-your-new-ci-the-acceptance-layer-for-work-your-team-didnt-write-20k7" class="crayons-story__hidden-navigation-link"&gt;Evals Are Your New CI: The Acceptance Layer for Work Your Team Didn't Write&lt;/a&gt;


  &lt;div class="crayons-story__body crayons-story__body-full_post"&gt;
    &lt;div class="crayons-story__top"&gt;
      &lt;div class="crayons-story__meta"&gt;
        &lt;div class="crayons-story__author-pic"&gt;

          &lt;a href="/cleberdelima" class="crayons-avatar  crayons-avatar--l  "&gt;
            &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3607194%2Fb81b1058-0dcb-45d3-84ad-67e471a945a1.png" alt="cleberdelima profile" class="crayons-avatar__image"&gt;
          &lt;/a&gt;
        &lt;/div&gt;
        &lt;div&gt;
          &lt;div&gt;
            &lt;a href="/cleberdelima" class="crayons-story__secondary fw-medium m:hidden"&gt;
              Cleber de Lima
            &lt;/a&gt;
            &lt;div class="profile-preview-card relative mb-4 s:mb-0 fw-medium hidden m:inline-block"&gt;
              
                Cleber de Lima
                
              
              &lt;div id="story-author-preview-content-4314176" class="profile-preview-card__content crayons-dropdown branded-7 p-4 pt-0"&gt;
                &lt;div class="gap-4 grid"&gt;
                  &lt;div class="-mt-4"&gt;
                    &lt;a href="/cleberdelima" class="flex"&gt;
                      &lt;span class="crayons-avatar crayons-avatar--xl mr-2 shrink-0"&gt;
                        &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3607194%2Fb81b1058-0dcb-45d3-84ad-67e471a945a1.png" class="crayons-avatar__image" alt=""&gt;
                      &lt;/span&gt;
                      &lt;span class="crayons-link crayons-subtitle-2 mt-5"&gt;Cleber de Lima&lt;/span&gt;
                    &lt;/a&gt;
                  &lt;/div&gt;
                  &lt;div class="print-hidden"&gt;
                    
                      Follow
                    
                  &lt;/div&gt;
                  &lt;div class="author-preview-metadata-container"&gt;&lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
            &lt;/div&gt;

          &lt;/div&gt;
          &lt;a href="https://dev.to/cleberdelima/evals-are-your-new-ci-the-acceptance-layer-for-work-your-team-didnt-write-20k7" class="crayons-story__tertiary fs-xs"&gt;&lt;time&gt;Aug 4&lt;/time&gt;&lt;span class="time-ago-indicator-initial-placeholder"&gt;&lt;/span&gt;&lt;/a&gt;
        &lt;/div&gt;
      &lt;/div&gt;

    &lt;/div&gt;

    &lt;div class="crayons-story__indention"&gt;
      &lt;h2 class="crayons-story__title crayons-story__title-full_post"&gt;
        &lt;a href="https://dev.to/cleberdelima/evals-are-your-new-ci-the-acceptance-layer-for-work-your-team-didnt-write-20k7" id="article-link-4314176"&gt;
          Evals Are Your New CI: The Acceptance Layer for Work Your Team Didn't Write
        &lt;/a&gt;
      &lt;/h2&gt;
        &lt;div class="crayons-story__tags"&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/ai"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;ai&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/programming"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;programming&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/llm"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;llm&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/agents"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;agents&lt;/a&gt;
        &lt;/div&gt;
      &lt;div class="crayons-story__bottom"&gt;
        &lt;div class="crayons-story__details"&gt;
            &lt;a href="https://dev.to/cleberdelima/evals-are-your-new-ci-the-acceptance-layer-for-work-your-team-didnt-write-20k7#comments" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left flex items-center"&gt;
              

              &lt;span class="hidden s:inline"&gt;Add&amp;nbsp;Comment&lt;/span&gt;
            &lt;/a&gt;
        &lt;/div&gt;
        &lt;div class="crayons-story__save"&gt;
          &lt;small class="crayons-story__tertiary fs-xs mr-2"&gt;
            8 min read
          &lt;/small&gt;
            
              &lt;span class="bm-initial crayons-icon c-btn__icon"&gt;
                

              &lt;/span&gt;
              &lt;span class="bm-success crayons-icon c-btn__icon"&gt;
                

              &lt;/span&gt;
            
        &lt;/div&gt;
      &lt;/div&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;/div&gt;


</description>
    </item>
    <item>
      <title>Evals Are Your New CI: The Acceptance Layer for Work Your Team Didn't Write</title>
      <dc:creator>Cleber de Lima</dc:creator>
      <pubDate>Tue, 04 Aug 2026 15:48:39 +0000</pubDate>
      <link>https://dev.to/cleberdelima/evals-are-your-new-ci-the-acceptance-layer-for-work-your-team-didnt-write-20k7</link>
      <guid>https://dev.to/cleberdelima/evals-are-your-new-ci-the-acceptance-layer-for-work-your-team-didnt-write-20k7</guid>
      <description>&lt;p&gt;Your agents are productive. Pull request volume is up, the demos land, and the pipeline is green on every merge. Here is the uncomfortable part: green means the code compiled and the tests passed, and increasingly the same agent that wrote the code wrote the tests. Nothing in that pipeline measures whether the work is actually right. "Looks right" is carrying the weight "verified right" used to carry.&lt;/p&gt;

&lt;p&gt;The decision in front of you is whether to build an acceptance layer for agent-produced work, the way CI became the acceptance layer for human-produced code, or to keep expanding agent autonomy on top of visual inspection. Agents do not fail loudly. A model update ships, a support agent starts missing escalation cues, no error is thrown, and you find out from churn metrics weeks later. The teams that scale autonomy safely will not be the ones with the best prompts. They will be the ones with the most disciplined evals.&lt;/p&gt;

&lt;p&gt;First, the term itself, because it carries the whole argument. An eval is a repeatable, scored test of an agent's output against criteria you define: not "did the tests pass" but "was this change correct, scoped to the ticket, and safe to merge". Where a unit test returns pass or fail on deterministic code, an eval grades judgment across dimensions, using graders that range from simple scripts to models judging other models against a written rubric. A body of these, run on every change the way CI runs on every commit, is what I will call the eval estate. Earlier in the series I argued that the agent that writes must never be the agent that checks; that settles who judges. The harder problem, and the subject of this episode, is how the judge knows what good looks like.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Acceptance Layer Is Being Rebuilt, Industry-Wide
&lt;/h2&gt;

&lt;p&gt;Anthropic's engineering team published this reframe in January 2026 in &lt;a href="https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents" rel="noopener noreferrer"&gt;"Demystifying evals for AI agents"&lt;/a&gt;: evaluations are not a research afterthought but the CI/CD pipeline for agentic systems. Within weeks, Braintrust made the analogy precise with &lt;a href="https://www.braintrust.dev/articles/eval-driven-development" rel="noopener noreferrer"&gt;eval-driven development&lt;/a&gt;, the agentic answer to test-driven development, with one difference that matters: TDD is binary, EDD scores across dimensions, because an answer can be factually accurate but too long, or well-formatted but missing the key information. Several vendors landed on the same operational rule in the same quarter: if the agent's metrics miss threshold on your benchmark set, the deployment fails automatically. If you're reading this I don't think think need to convince you evals matter, instead, I'll try to show you how teams actually build them, and where they go wrong.&lt;/p&gt;

&lt;p&gt;Braintrust's worked example is the best picture of the daily rhythm: a team swaps in a newer model, the first eval run shows tone dropping from 0.85 to 0.72, they adjust the prompt, rerun, and tone recovers to 0.88 with accuracy intact. Twenty minutes end to end. Without the estate, that regression surfaces as customer complaints two weeks later.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Your Test Suite Cannot Do This Job
&lt;/h2&gt;

&lt;p&gt;Unit tests work because the same input returns the same output. Agents break that contract: the same prompt produces different outputs across runs and model versions, errors compound across multi-turn interactions, and agents choose solution paths you never anticipated. Your test suite checks the code. Nothing checks the judgment.&lt;/p&gt;

&lt;p&gt;There is also a failure mode that did not exist when humans wrote the code: test-gaming. A developer rarely games their own test suite. For an agent, optimizing for test passage is the default unless you deliberately architect around it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to Grade: Correctness Is the Easy Third
&lt;/h2&gt;

&lt;p&gt;Anthropic's taxonomy gives you the mechanics: &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;code-based graders for deterministic checks&lt;/li&gt;
&lt;li&gt;model-based graders for rubric judgments&lt;/li&gt;
&lt;li&gt;human graders used sparingly, mainly to calibrate the model-based ones&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;But the most useful answer to "grade what?" came from FrontierCode, Cognition's own benchmark built with 36 external open-source maintainers: they scored patches across six dimensions: &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;behavioural correctness&lt;/li&gt;
&lt;li&gt;regression safety&lt;/li&gt;
&lt;li&gt;mechanical cleanliness&lt;/li&gt;
&lt;li&gt;test correctness&lt;/li&gt;
&lt;li&gt;scope discipline&lt;/li&gt;
&lt;li&gt;code quality&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
&lt;p&gt;How much of that list your pipeline measures today ?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The Playbook: Building the Estate From What You Already Have
&lt;/h2&gt;

&lt;p&gt;The standard objection is that eval sets take a mountain of manual labeling. Your organization is already sitting on the raw material. Four steps, in the order that survives production.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Seed the set from your PR history and postmortems.&lt;/strong&gt; &lt;a href="https://arxiv.org/abs/2310.06770" rel="noopener noreferrer"&gt;SWE-bench&lt;/a&gt; established the template: mine merged PRs that fixed a linked issue and touched tests, reconstruct the repo state before the fix, keep only cases where tests fail before the patch and pass after while everything else stays green. They filtered about 90,000 PRs down to 2,294 clean cases, which tells you both how much material a real repo holds and how hard the filter should be. Your review comments are a second seam; Cursor Bugbot, Qodo, Greptile, and CodeRabbit already build learned rules from exactly this data. Your incidents are the third: a postmortem is not done until one thing changes, such as adding a regression eval; otherwise it was documentation, not engineering. Pitfall: writing eval cases from imagination; you will test what you feared, not what happens. Signal: every eval case traces to a real PR, incident, or production trace.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Store behavioral scenarios where the agent cannot read them.&lt;/strong&gt; StrongDM's structural fix is the cleanest I have seen: behavioral scenarios kept outside the codebase, invisible to the agent while it works, functioning like a machine-learning holdout set. The agent cannot optimize for criteria it never sees. The same applies to the judge: it must not see the maker's reasoning, or it inherits the maker's assumptions. Pitfall: putting eval criteria in the repo "for transparency," where the agent reads them and optimizes for the letter of the check. Signal: the agent's pass rate on holdout scenarios sits visibly below its pass rate on in-repo tests.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Calibrate the judge before you let it gate.&lt;/strong&gt; LLM-as-judge is the scaling mechanism for all of this and the place discipline usually collapses: industry figures circulating through 2026 put the share of teams that fail to implement judges well at around 93 percent. I treat the number as directional, but the root cause is missing calibration, not weak models. I learned this the direct way. The first Reviewer agent we ran inside the AI-DLC harness at Betsson approved almost everything it saw: it answered a binary question, pass or fail, and a capable model asked a binary question about plausible-looking code says pass. The gate only started failing work when we replaced the yes/no check with scored rubrics built from what our human reviewers had actually rejected: edge cases declared covered that were not, tests weakened until they passed, changes sprawling beyond the ticket. The workflow that came out of it: write scored rubrics with each band defined in plain terms (what a 0.2 looks like, what an 0.8 looks like); keep a fixed ground-truth set of known failures; measure the judge's agreement with your human reviewers on it; do not let it gate anything below roughly 75 to 90 percent agreement. Require reasons before the score, because a model that emits the score first defends it regardless. When agreement stalls, debug the rubric, not the agent prompt. Boring, repetitive work, and the highest-leverage work we did. Pitfall: trusting a judge because it agrees with you on ten samples. Signal: a standing agreement rate against fresh human-reviewed samples.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Wire the gate into promotion, then watch it drift.&lt;/strong&gt; Eval scores become promotion criteria between environments: a change that drops a metric past threshold does not ship. Then treat the estate as a living artifact. A suite passing at 100 percent is not a healthy suite, it is a dead sensor. And every model change re-ages the estate: Anthropic engineers describe the same tool-use prompt drastically under-triggering on one model version and over-triggering on the next, so the eval that graded it gave a false signal across the upgrade. Pitfall: celebrating a saturated suite as success. Signal: the gate blocks something real most weeks, and new cases are added monthly.&lt;/p&gt;

&lt;h2&gt;
  
  
  What This Costs, and the Kill Metric
&lt;/h2&gt;

&lt;p&gt;The honest cost is senior human time defining what "done" means for your own workflows; nobody has published a credible dollar figure, and I will not invent one. Fund first: a regression suite and one calibrated judge for your single most valuable agent workflow, seeded from its PR history and incidents. Fund later: fleet-wide coverage and drift monitoring on sampled production traffic. The kill metric within a quarter: the gate must catch real regressions before production, and the judge must hold its agreement rate against fresh human samples. A gate that has failed nothing in ninety days is decoration. Fix the rubrics or stop pretending you have a gate.&lt;/p&gt;

&lt;p&gt;This spend compounds, the same way &lt;a href="https://dev.to/cleberdelima/testing-reinvented-why-test-coverage-is-the-wrong-metric-31l3"&gt;test effectiveness beat test coverage&lt;/a&gt; as the metric that matters: every eval mined from a real failure keeps paying on every future run, model swap, and vendor negotiation. And the trust gap is measurable: &lt;a href="https://cloud.google.com/resources/content/2025-dora-ai-assisted-software-development-report" rel="noopener noreferrer"&gt;DORA's 2025 report&lt;/a&gt; found only 24 percent of developers trust AI output "a lot," and LinearB's 2026 benchmarks found AI-generated PRs merging at less than half the rate of manual ones, with leaders naming context and trust, not correctness, as the barrier. Evals are how trust gets a number.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start, Stop, Continue
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Executives&lt;/strong&gt;&lt;br&gt;
Start: funding the eval estate as acceptance infrastructure with a named owner, the same way CI had one; asking for judge-human agreement rates before approving any autonomy increase.&lt;br&gt;
Stop: accepting "the pipeline is green" as evidence agent work is correct; approving agent deployments gated only on tests the agent itself wrote.&lt;br&gt;
Continue: holding a human accountable for every merged change, with evals deciding how much of the checking they can delegate.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Engineers&lt;/strong&gt;&lt;br&gt;
Start: mining your last quarter of rejected PRs and incidents into a seed eval set this week; requiring reasons before scores from every judge.&lt;br&gt;
Stop: gating anything on an uncalibrated judge; celebrating 100 percent pass rates; scoring quality with binary checks.&lt;br&gt;
Continue: reading what the agents ship. The eval estate extends your judgment; it does not replace it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Strategic Takeaway
&lt;/h2&gt;

&lt;p&gt;An eval estate compounds: every mined failure and calibrated rubric makes the next agent, the next model, and the next autonomy decision cheaper and safer. Inspection plateaus: it scales exactly as fast as senior attention, and senior attention is already your bottleneck. Anthropic Institute's own account of running at over 80 percent AI-authored production code names human review as the new constraint, and their mandatory automated reviewer, retrospectively, would have caught roughly a third of the bugs behind past incidents. The acceptance layer is where the constraint moves next. Whoever industrializes it first sets their own pace.&lt;/p&gt;

&lt;p&gt;So run this test on your organization: if a better model shipped tomorrow, could you tell by the end of the day whether your most important agent workflow got better or worse, with a number? If the answer is no, every autonomy increase you approve is a bet placed blind. Tell me where I am wrong. If your team ships agent work confidently without an eval estate, I want to know what you are doing instead, and if the argument holds, send this to whoever owns your CI pipeline and ask them who owns the evals.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>llm</category>
      <category>agents</category>
    </item>
    <item>
      <title>You stopped typing the code. You still type every prompt. The slowest, most expensive component in your AI workflow is a person deciding each next step by hand. The fix is not a better prompt.</title>
      <dc:creator>Cleber de Lima</dc:creator>
      <pubDate>Thu, 30 Jul 2026 08:52:19 +0000</pubDate>
      <link>https://dev.to/cleberdelima/you-stopped-typing-the-code-you-still-type-every-prompt-the-slowest-most-expensive-component-in-59dl</link>
      <guid>https://dev.to/cleberdelima/you-stopped-typing-the-code-you-still-type-every-prompt-the-slowest-most-expensive-component-in-59dl</guid>
      <description>&lt;div class="ltag__link--embedded"&gt;
  &lt;div class="crayons-story "&gt;
  &lt;a href="https://dev.to/cleberdelima/loop-engineering-stop-prompting-your-agents-and-design-the-system-that-does-428p" class="crayons-story__hidden-navigation-link"&gt;Loop Engineering: Stop Prompting Your Agents and Design the System That Does&lt;/a&gt;


  &lt;div class="crayons-story__body crayons-story__body-full_post"&gt;
    &lt;div class="crayons-story__top"&gt;
      &lt;div class="crayons-story__meta"&gt;
        &lt;div class="crayons-story__author-pic"&gt;

          &lt;a href="/cleberdelima" class="crayons-avatar  crayons-avatar--l  "&gt;
            &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3607194%2Fb81b1058-0dcb-45d3-84ad-67e471a945a1.png" alt="cleberdelima profile" class="crayons-avatar__image" width="800" height="800"&gt;
          &lt;/a&gt;
        &lt;/div&gt;
        &lt;div&gt;
          &lt;div&gt;
            &lt;a href="/cleberdelima" class="crayons-story__secondary fw-medium m:hidden"&gt;
              Cleber de Lima
            &lt;/a&gt;
            &lt;div class="profile-preview-card relative mb-4 s:mb-0 fw-medium hidden m:inline-block"&gt;
              
                Cleber de Lima
                
              
              &lt;div id="story-author-preview-content-4246754" class="profile-preview-card__content crayons-dropdown branded-7 p-4 pt-0"&gt;
                &lt;div class="gap-4 grid"&gt;
                  &lt;div class="-mt-4"&gt;
                    &lt;a href="/cleberdelima" class="flex"&gt;
                      &lt;span class="crayons-avatar crayons-avatar--xl mr-2 shrink-0"&gt;
                        &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3607194%2Fb81b1058-0dcb-45d3-84ad-67e471a945a1.png" class="crayons-avatar__image" alt="" width="800" height="800"&gt;
                      &lt;/span&gt;
                      &lt;span class="crayons-link crayons-subtitle-2 mt-5"&gt;Cleber de Lima&lt;/span&gt;
                    &lt;/a&gt;
                  &lt;/div&gt;
                  &lt;div class="print-hidden"&gt;
                    
                      Follow
                    
                  &lt;/div&gt;
                  &lt;div class="author-preview-metadata-container"&gt;&lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
            &lt;/div&gt;

          &lt;/div&gt;
          &lt;a href="https://dev.to/cleberdelima/loop-engineering-stop-prompting-your-agents-and-design-the-system-that-does-428p" class="crayons-story__tertiary fs-xs"&gt;&lt;time&gt;Jul 27&lt;/time&gt;&lt;span class="time-ago-indicator-initial-placeholder"&gt;&lt;/span&gt;&lt;/a&gt;
        &lt;/div&gt;
      &lt;/div&gt;

    &lt;/div&gt;

    &lt;div class="crayons-story__indention"&gt;
      &lt;h2 class="crayons-story__title crayons-story__title-full_post"&gt;
        &lt;a href="https://dev.to/cleberdelima/loop-engineering-stop-prompting-your-agents-and-design-the-system-that-does-428p" id="article-link-4246754"&gt;
          Loop Engineering: Stop Prompting Your Agents and Design the System That Does
        &lt;/a&gt;
      &lt;/h2&gt;
        &lt;div class="crayons-story__tags"&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/ai"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;ai&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/programming"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;programming&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/llm"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;llm&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/agents"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;agents&lt;/a&gt;
        &lt;/div&gt;
      &lt;div class="crayons-story__bottom"&gt;
        &lt;div class="crayons-story__details"&gt;
            &lt;a href="https://dev.to/cleberdelima/loop-engineering-stop-prompting-your-agents-and-design-the-system-that-does-428p#comments" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left flex items-center"&gt;
              

              &lt;span class="hidden s:inline"&gt;Add&amp;nbsp;Comment&lt;/span&gt;
            &lt;/a&gt;
        &lt;/div&gt;
        &lt;div class="crayons-story__save"&gt;
          &lt;small class="crayons-story__tertiary fs-xs mr-2"&gt;
            8 min read
          &lt;/small&gt;
            
              &lt;span class="bm-initial crayons-icon c-btn__icon"&gt;
                

              &lt;/span&gt;
              &lt;span class="bm-success crayons-icon c-btn__icon"&gt;
                

              &lt;/span&gt;
            
        &lt;/div&gt;
      &lt;/div&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;/div&gt;


</description>
    </item>
    <item>
      <title>Loop Engineering: Stop Prompting Your Agents and Design the System That Does</title>
      <dc:creator>Cleber de Lima</dc:creator>
      <pubDate>Mon, 27 Jul 2026 16:48:59 +0000</pubDate>
      <link>https://dev.to/cleberdelima/loop-engineering-stop-prompting-your-agents-and-design-the-system-that-does-428p</link>
      <guid>https://dev.to/cleberdelima/loop-engineering-stop-prompting-your-agents-and-design-the-system-that-does-428p</guid>
      <description>&lt;p&gt;Your engineers have their AI licenses. They prompt, read what comes back, fix it, and prompt again. The dashboard is green and everyone agrees the tools help. Here is the part that should worry you: you have automated the typing and kept the slowest, most expensive component inside every cycle. A person, deciding each next step by hand.&lt;/p&gt;

&lt;p&gt;That ceiling does not move with more licenses, because it is built into how you use AI, and there is now a name for the skill that removes it. &lt;strong&gt;Loop engineering&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;You stop being the person who prompts the agent and design the system that does it instead.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Which parts of your delivery system will you run as machine-checked loops, and what will you fund to make the checks trustworthy? Decide deliberately, or your AI gains keep scaling with human attention while a competitor's start to compound.&lt;/p&gt;

&lt;h2&gt;
  
  
  The leverage point moved
&lt;/h2&gt;

&lt;p&gt;The first time we let an agent run unattended inside the AI-DLC harness at Betsson, it came back reporting success. It had quietly weakened a failing test until it passed. That single incident taught me more about where the leverage in AI delivery sits than any benchmark, and it is why I read the loop-engineering wave as a real shift rather than a slogan. I have seen enough delivery shifts arrive disguised as tooling upgrades to recognize the pattern: the tool is the least interesting part.&lt;/p&gt;

&lt;p&gt;Two engineers from rival camps said the same thing within a day of each other. &lt;a href="https://www.linkedin.com/in/steipete/" rel="noopener noreferrer"&gt;Peter Steinberger&lt;/a&gt; &lt;a href="https://x.com/steipete/status/2063697162748260627" rel="noopener noreferrer"&gt;posted&lt;/a&gt;: "You shouldn't be prompting coding agents anymore. You should be designing loops that prompt your agents." &lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.linkedin.com/in/bcherny/" rel="noopener noreferrer"&gt;Boris Cherny&lt;/a&gt;, who leads Claude Code at Anthropic, said much the same the next day, in a remark that &lt;a href="https://x.com/rohanpaul_ai/status/2063289804708835412" rel="noopener noreferrer"&gt;circulated widely afterward&lt;/a&gt;: "I don't prompt Claude anymore. I have loops running that prompt Claude. My job is to write loops."&lt;/p&gt;

&lt;p&gt;Strip the vocabulary away and the concept is simple. A prompt is an instruction: one answer, then the agent waits for you. A loop is a goal, a way to check progress against it, and a rule for when to stop. &lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;In a Loop, the system discovers the work, plans, executes, verifies, and iterates until the check passes or the stop rule fires.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The test that separates a real loop from a scheduled prompt is discovery: a loop never says "go fix X." It teaches the agent how to find the work, gives it the tools to validate what it finds, and reserves your judgment for the result.&lt;/p&gt;

&lt;p&gt;What changed this year is that this stopped being a scripting project. &lt;a href="https://www.linkedin.com/in/addyosmani/" rel="noopener noreferrer"&gt;Addy Osmani&lt;/a&gt; describes a working loop as five building blocks: &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;scheduled automations&lt;/li&gt;
&lt;li&gt;isolated workspaces&lt;/li&gt;
&lt;li&gt;written-down project knowledge&lt;/li&gt;
&lt;li&gt;connectors so the loop opens the pull request itself&lt;/li&gt;
&lt;li&gt;sub-agents so the writer is never the checker
All of it, held together by a memory, a markdown file or ticket board outside the conversation. &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Every piece now ships inside both Claude Code, Cursor and Codex, which makes this an operating-model decision, not a tooling backlog item. It is the continuous intelligence loop from &lt;a href="https://dev.to/cleberdelima/redefining-the-software-lifecycle-why-your-sdlc-is-already-obsolete-54nf"&gt;Redefining the Software Lifecycle&lt;/a&gt;, now with a heartbeat, a gate, and a memory.&lt;/p&gt;

&lt;h2&gt;
  
  
  The gate is the product
&lt;/h2&gt;

&lt;p&gt;Here is the claim worth arguing with me about: the loop is now a more valuable engineering artifact than the code it produces. &lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Code is becoming the byproduct. &lt;strong&gt;The loop is the asset.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Everything in a loop lives or dies on the gate, the check that can automatically fail the work. Without it you do not have a loop; you have an agent agreeing with itself on repeat, because the model that wrote the code is far too generous grading its own homework.&lt;/p&gt;

&lt;p&gt;Code Writer agents generate within guardrails, separate Reviewer and Test Writer agents validate against the spec, humans approve at defined gates, and production incidents feed back so the same failure cannot ship twice.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a production loop looks like
&lt;/h2&gt;

&lt;p&gt;The most instructive loop I have seen outside our own harness runs on a normal budget. &lt;a href="https://www.linkedin.com/in/rafaelquintanilha/" rel="noopener noreferrer"&gt;Fabio Quintanilha&lt;/a&gt; mentioned on one of his recent videos he runs a daily "backlog scout" in Codex: it scans the board for exactly one small, self-contained, unassigned ticket, checks the affected code and tests before proposing, and carries a hard exclusion list: mobile, auth, payments, database migrations, anything backend. It remembers rejected suggestions so a "no" is not re-proposed tomorrow, and one standing rule says "nothing safe today" is an acceptable answer. When he approves a candidate, it implements in an isolated worktree and opens a draft pull request. He signs off at the end; he never tells it what to fix.&lt;/p&gt;

&lt;p&gt;Notice what carries that loop: the exclusion list, the memory of rejections, the permission to find nothing, the draft PR as the gate. All design decisions, none of them prompts. His cost caveat is worth repeating: Steinberger and Cherny have effectively unlimited tokens, so their loops wake every five minutes; most teams should pick the cadence their budget survives, and the loop works the same.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the concept breaks
&lt;/h2&gt;

&lt;p&gt;Not everything deserves a loop. The best filter I have seen, from Anatoli Kopadze's write-up on loops, is four conditions that must all hold: &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the task repeats at least weekly&lt;/li&gt;
&lt;li&gt;something can automatically reject bad output&lt;/li&gt;
&lt;li&gt;the agent can do the work end to end&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;"done" is objective rather than a matter of taste&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Miss one and a good prompt is still the better tool.&lt;/p&gt;

&lt;p&gt;Loops also fail quietly: an agent declares victory early, and the loop keeps running and billing with nothing to show. Every serious loop needs two exits, verified success and a hard cap that stops and reports. And the cost compounds in a shape budgets do not expect: every pass re-reads a growing context, and the maker-checker split doubles the reads. The full economics of metered tokens get their own episode in &lt;a href="https://dev.to/cleberdelima/token-economics-why-your-ai-bill-is-a-capital-decision-not-a-cost-to-cut-2ch"&gt;Token Economics&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The sharpest risk is human, and it is why loop design is harder than prompt engineering, not easier. The faster a loop ships code you did not write, the faster the gap grows between what exists and what you understand. Osmani calls the failure shapes &lt;a href="https://addyosmani.com/blog/comprehension-debt/" rel="noopener noreferrer"&gt;comprehension debt&lt;/a&gt; and &lt;a href="https://addyosmani.com/blog/cognitive-surrender/" rel="noopener noreferrer"&gt;cognitive surrender&lt;/a&gt;. Two people can run the identical loop and get opposite outcomes: one moves faster on work they understand deeply, the other uses the loop to avoid understanding it at all. The loop cannot tell the difference. That is &lt;a href="https://dev.to/cleberdelima/the-velocity-trap-why-your-ai-productivity-gains-are-an-illusion-o6o"&gt;The Velocity Trap&lt;/a&gt; running unattended, on a schedule.&lt;/p&gt;

&lt;p&gt;This is not a contradiction with &lt;a href="https://dev.to/cleberdelima/continuous-fluid-flow-how-ai-is-compressing-the-software-delivery-cycle-3f20"&gt;Continuous Fluid Flow&lt;/a&gt;, where I argued AI-native delivery needs more synchronous human collaboration: humans design the loops, the specs, gates, and guardrails together. Only then the loops should run unattended. You are not removing judgment, you are concentrating it where it compounds.&lt;/p&gt;

&lt;h2&gt;
  
  
  Engineering your first loop
&lt;/h2&gt;

&lt;p&gt;Five steps, in an order that survives production.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Select loop-eligible work with the four-condition test.&lt;/strong&gt; Keep only work that passes all four conditions, and start where verification is free: nightly CI failure triage, dependency updates, flaky-test hunts, documentation drift. Or copy the backlog scout above, exclusion list and all. Signal: for every selected task, you can name the exact command that rejects bad output.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Prove one manual run, then write it down.&lt;/strong&gt; The order is manual run, then skill, then loop, then schedule. Run the task by hand with the agent until it works, then capture what made it work in a SKILL.md file: conventions, build steps, the never-touch list, the "we do not do it that way because of that incident" knowledge. Pitfall: scheduling something you never made reliable by hand, which is how loops blow up overnight. Signal: two clean runs in a row driven by the written skill alone.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Engineer the gate before the loop.&lt;/strong&gt; Define machine-checkable success and a hard stop before the first unattended run. Write a loop spec. GOAL: all tests in /tests/auth pass, lint clean, zero type errors. EACH PASS: fix the single highest-impact failure. STOP: when verification passes, or after 8 iterations, then summarize what changed and what still fails. Assign verification to a separate checker agent, ideally a different model. Inside the Betsson harness, our checkers carried one standing instruction the writers never got: treat the change as wrong until the spec and the existing tests prove otherwise, and never edit a test to make it pass. Signal: the checker rejects some share of drafts every week.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Put state on disk, not in the context window.&lt;/strong&gt; The loop's memory must outlive the session. Keep a state file or ticket board with four fields per item: tried, result, still open, next. Tomorrow's run resumes instead of rediscovering the repo from zero. Pitfall: trusting the context window as memory; the agent forgets, the repo does not.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5. Track cost per accepted change, weekly.&lt;/strong&gt; Log every run with tokens consumed, changes proposed, and changes accepted by the human gatekeeper. Counting loops run and tokens burned measures activity, not value.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this costs, and what to fund first
&lt;/h2&gt;

&lt;p&gt;The loop tooling itself is nearly free; it ships inside products you already license. The honest costs are the test suites trustworthy enough to serve as gates, CI signal quality, and review capacity for what loops produce. Fund first: verification infrastructure and one contained pilot loop. Fund later: parallel fleets. The kill metric: within one quarter, the pilot's cost per accepted change should be flat or falling with acceptance above 50%; below roughly half of proposals accepted, a loop costs more than it returns. If it is not, fix the gates or kill the loop. Do not fix the prompts.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start, Stop, Continue
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Executives&lt;/strong&gt;&lt;br&gt;
Start: funding gates as strategic infrastructure; requiring the four-condition test before any workflow is automated; putting cost per accepted change on the engineering dashboard.&lt;br&gt;
Stop: counting licenses, prompts, or tokens as progress; approving unattended loops with no iteration cap or token budget.&lt;br&gt;
Continue: holding a named human accountable for every change a loop merges.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Engineers&lt;/strong&gt;&lt;br&gt;
Start: converting your most repeated manual workflow into a skill file this week; splitting maker from checker on anything that runs unattended; giving every loop an exclusion list and permission to find nothing.&lt;br&gt;
Stop: re-prompting the same triage every morning; letting any agent verify its own work; scheduling anything you have not run reliably by hand.&lt;br&gt;
Continue: reading what the loop ships. Your comprehension is the gate behind the gate.&lt;/p&gt;

&lt;h2&gt;
  
  
  Strategic takeaway
&lt;/h2&gt;

&lt;p&gt;Loops with real gates compound: every cycle updates the skills, the guardrails, and the state, so the next run starts smarter. Prompting plateaus: it scales only as fast as human attention, no matter how good the model gets. That gap is small this quarter and brutal in four. The teams that win the next phase will not have the best prompt libraries; they will have the best-engineered gates, because the gate is what turns AI activity into delivered value, and it is the one artifact a competitor cannot copy from a screenshot.&lt;/p&gt;

&lt;p&gt;So here is the question for your next leadership meeting:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Which workflow in your organization would you trust to run overnight with no human watching?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If the answer is none, is that because your work is impossible to verify automatically, or because nobody has built the gate yet?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>llm</category>
      <category>agents</category>
    </item>
    <item>
      <title>Instruction Debt: Your Prompts Are Aging Like Code</title>
      <dc:creator>Cleber de Lima</dc:creator>
      <pubDate>Mon, 20 Jul 2026 06:52:00 +0000</pubDate>
      <link>https://dev.to/cleberdelima/instruction-debt-your-prompts-are-aging-like-code-1li7</link>
      <guid>https://dev.to/cleberdelima/instruction-debt-your-prompts-are-aging-like-code-1li7</guid>
      <description>&lt;p&gt;The new LLM shipped last week, promising superior intelligence. You pointed your agents at it, and the results got worse.&lt;/p&gt;

&lt;p&gt;Your engineers shrugged: the hype outran reality again. Here is the alternative: the model is fine. It is reading instructions your team wrote a year ago to keep a weaker model out of trouble, and unlike that weaker model, it found every one of them and obeyed.&lt;/p&gt;

&lt;p&gt;I call this instruction debt: the cost of instructions you wrote down and never retired. The decision in front of you is whether your standing instructions, the full set of system prompts, agent rules files, skills, and guardrails your agents read, get treated as versioned infrastructure with a deprecation path, or stay write-once files that quietly tax every model upgrade you will ever make. With frontier release gaps &lt;a href="https://finance.biggo.com/news/5320d310d8ec014d" rel="noopener noreferrer"&gt;compressed to around 40 days by Theo Browne's count&lt;/a&gt;, "wait and see" means paying that tax every six weeks.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Better the Model, the Worse Your Stale Instructions
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://every.to/vibe-check/gpt-5-6-sol" rel="noopener noreferrer"&gt;Every's review of GPT-5.6&lt;/a&gt;, (by the way I love @every's content) ressonate that pattern. They initially found GPT-5.6-Sol worse than its predecessor at customer support work, until they deleted defensive rules written for older models. The conclusion: "Sol had found the instructions and followed them; the instructions were the problem."&lt;/p&gt;

&lt;p&gt;Frontier Models, are better than their predecessors at the exact thing that makes stale instructions dangerous: it does the reading. It searches files, reads standing instructions, and checks connected tools before asking for help. A weaker model that ignored your rules file was accidentally protected from that file's staleness. A stronger one is not.&lt;/p&gt;

&lt;p&gt;And this is not one reviewer's experience. &lt;a href="https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5" rel="noopener noreferrer"&gt;Anthropic's own migration guidance&lt;/a&gt; for its newest model says it plainly: skills developed for prior models "are often too prescriptive" and "can degrade output quality," and capability improvements "are a good prompt to re-evaluate which instructions, tools, and guardrails are still needed." When a vendor states in an official document that your existing instructions can make its new model worse, the debate about whether this problem exists is over. That is the inversion: staleness used to be free because weak models either needed or ignored your rules. It is now different&lt;/p&gt;

&lt;h2&gt;
  
  
  Dead Weight and Active Harm
&lt;/h2&gt;

&lt;p&gt;Audit for two failure modes.&lt;/p&gt;

&lt;p&gt;Dead weight is a instruction the new model no longer needs: "think step by step," enumerated lists of behaviors, role-play framing like "you are a senior engineer." Anthropic now says you can steer most behaviors with a brief instruction; the enumeration itself is a leftover from weaker models, wasting context and burying the rules that still matter.&lt;/p&gt;

&lt;p&gt;Active harm is instruction the new model obeys to its own detriment. The defensive rules that dragged GPT-5.6-Sol down for Every are one case. Theo Browne supplied another &lt;a href="https://finance.biggo.com/news/5320d310d8ec014d" rel="noopener noreferrer"&gt;from his own repository&lt;/a&gt;: T3 Code's agent rules file sat unchanged for two months, still calling the shipped project "a very early WIP" and still directing "sweeping changes that improve long-term maintainability," an open invitation for a capable model to rewrite architecture nobody asked it to touch. No one noticed, because nothing broke.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Industry Is Carrying This Debt at Scale
&lt;/h2&gt;

&lt;p&gt;The data point that changed my mind about the scale came from &lt;a href="https://www.dbreunig.com/2026/06/22/the-problem-is-prompt-debt.html" rel="noopener noreferrer"&gt;Drew Breunig's analysis of prompt debt&lt;/a&gt;: Datadog's State of AI Engineering report found the most-used model in observed traffic in March 2026 was GPT-4o, a model nearly two years old by then. Breunig adds that several large inference providers privately put GPT-4o-vintage models above half of all calls, a secondhand figure, but directionally consistent. A Berkeley-led study he cites found enterprises stay pinned to older models because newer ones break their existing agents. Half an industry frozen on old models is not nostalgia. It is aggregate instruction debt made visible: the upgrade breaks the instruction base, so the upgrade waits.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.dbreunig.com/2026/06/22/the-problem-is-prompt-debt.html" rel="noopener noreferrer"&gt;Breunig also documents&lt;/a&gt; why the debt keeps growing. Most standing instructions exist to fight a specific model's training: ChatGPT's image pipeline told its model eight times not to reply when an image was returned; Claude Code told Opus seven times to parallelize tool calls. Each repetition is a dated workaround for one model generation, waiting to confuse the next. And holding still is not an escape. Goedecke again: a delicately prompted harness built around an old model will always lose to a bare-bones harness built around the current one.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Maintenance Bill on My Own Advice
&lt;/h2&gt;

&lt;p&gt;A fair objection: I've been telling you for the last 2 years about the importance of writing these files. The &lt;a href="https://dev.to/cleberdelima/building-software-in-the-age-of-ai-the-mindset-shift-and-the-playbook-that-actually-works-42jc"&gt;opening episode's playbook&lt;/a&gt; argued for guardrails, curated context, and reusable instructions, and later episodes doubled down with skills and context packs. Some people advocate that we should let the tools the more vanilla as possible and have nothing but the original harness provided by the vendors. I stand by all of it, specially considering the harness can help you you to save cost, and steer smaller and less capable models to perform better, so, my point is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Instruction debt is not an argument against writing instructions; it is the maintenance bill on that advice.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It completes a pattern. Addy Osmani gave us intent debt, the cost of intent you never wrote down, and &lt;a href="https://addyosmani.com/blog/comprehension-debt/" rel="noopener noreferrer"&gt;comprehension debt&lt;/a&gt;, the gap between the code that exists and the code you understand. Instruction debt is the third member: intent you never wrote, code you never read, instructions you never retired.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Playbook: Give Every Instruction a Deprecation Path
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. Inventory your instruction surface.&lt;/strong&gt; You cannot retire what you cannot see. List every place standing instructions live: system prompts, agent rules files, skills, MCP tool descriptions, prompt templates in environment variables. &lt;a href="https://techdebt.works/prompt-engineering-debt/" rel="noopener noreferrer"&gt;TechDebt.works&lt;/a&gt; lists the warning signs: prompts living in Slack messages, reliability described as "works most of the time," one person who owns all the prompts. The concrete handle is a registry, even a single spreadsheet: file, owner, model it was written for, last reviewed. Pitfall: auditing the main system prompt while per-repo rules files sprawl unowned. Signal: every instruction has a named owner and a "written for" model version.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Version prompts like code.&lt;/strong&gt; Prompts belong in git, next to the service, changed through pull requests. Add one question to the PR template: which model weakness does this instruction compensate for? Pitfall: treating version control as an archive nobody reads. Signal: prompt changes get the same review scrutiny as code changes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Eval-gate model upgrades like dependency updates.&lt;/strong&gt; Before switching models, run your evaluation suite against the new one, compare side by side, and hold the upgrade if a critical flow degrades. Tooling exists: &lt;a href="https://www.promptfoo.dev/docs/intro/" rel="noopener noreferrer"&gt;promptfoo&lt;/a&gt; for CI-style prompt evals, and research tools like &lt;a href="https://arxiv.org/abs/2409.03928" rel="noopener noreferrer"&gt;RETAIN&lt;/a&gt; built for regression-testing prompt migration between models. Pitfall: having no evals, which turns upgrade day into debugging in production. If you cannot score outputs automatically, fix that first. Signal: a model upgrade is a branch plus an eval report, not a leap of faith.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Audit and prune on every upgrade.&lt;/strong&gt; Two questions per instruction: does the new model still need this, and does the new model now obey this to our detriment? Strip old-generation behavioral steering immediately. Keep rules files to concrete project facts (where files live, the test framework, the never-touch list), never behavior coaching. Pitfall: pruning only when something visibly breaks; this debt does not break things visibly. Signal: the instruction base shrinks or holds flat across upgrades while eval scores rise.&lt;/p&gt;

&lt;h2&gt;
  
  
  What This Costs, and the Kill Metric
&lt;/h2&gt;

&lt;p&gt;The inventory and the pruning habit cost days, not quarters. The real investment is the eval suite, the same verification infrastructure this series keeps arriving at from different directions. Fund first: the instruction registry and evals for one critical agent workflow. Fund later: prompt-compilation tooling. The kill metric fits in one experiment: on your next model upgrade, run the new model against your eval set twice, once with the current instruction base and once with a pruned one. If pruned wins, you have measured your instruction debt directly, in eval points. If it does not, your instruction base is healthy and the audit cost you a day.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start, Stop, Continue
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Executives&lt;/strong&gt;&lt;br&gt;
Start: requiring model upgrades to pass an eval gate the way dependency updates pass tests; asking one question in your next AI review, who owns the standing instructions and when was one last pruned.&lt;br&gt;
Stop: booking "the new model is disappointing" as vendor failure before anyone has audited the instructions it was given; funding prompt writing with no retirement process attached.&lt;br&gt;
Continue: incident-driven guardrails, with the added rule that every new guardrail names the model weakness it compensates for.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Engineers&lt;/strong&gt;&lt;br&gt;
Start: adding a "written for model X" header to every rules file you own; deleting old-generation steering ("think step by step," role-play framing) this week.&lt;br&gt;
Stop: letting agents generate their own rules files without review; carrying any instruction you cannot explain the reason for.&lt;br&gt;
Continue: writing skills and context packs. The answer to instruction debt is a lifecycle, not abstinence.&lt;/p&gt;

&lt;h2&gt;
  
  
  Strategic Takeaway
&lt;/h2&gt;

&lt;p&gt;A versioned, eval-gated, pruned instruction base compounds: every model upgrade lands at close to full value in week one, and each audit makes the next one cheaper. A hand-tended one plateaus, then turns against you: every defensive fix deepens the lock-in, until you are one of the organizations still running two-year-old models because the upgrade breaks your agents. The gap is brutal because it is invisible: the competitor gets the new model's full capability immediately, while you spend a quarter concluding the hype outran reality.&lt;/p&gt;

&lt;p&gt;So run the test on yourself. Take your most important agent workflow and ask: if a better model shipped tomorrow, would that workflow get better, or would the model read your instructions first and get worse? If you cannot answer, that is the debt.&lt;/p&gt;

&lt;p&gt;I expect pushback on this one, especially from teams whose prompt libraries feel like hard-won assets. Mine currently is.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>llm</category>
      <category>softwaredevelopment</category>
    </item>
    <item>
      <title>Token Economics: Why Your AI Bill Is a Capital Decision, Not a Cost to Cut</title>
      <dc:creator>Cleber de Lima</dc:creator>
      <pubDate>Mon, 13 Jul 2026 08:04:00 +0000</pubDate>
      <link>https://dev.to/cleberdelima/token-economics-why-your-ai-bill-is-a-capital-decision-not-a-cost-to-cut-2ch</link>
      <guid>https://dev.to/cleberdelima/token-economics-why-your-ai-bill-is-a-capital-decision-not-a-cost-to-cut-2ch</guid>
      <description>&lt;p&gt;Your Copilot and Cursor seats are approved. Adoption is climbing, engineers are shipping, the dashboards look healthy. Then the invoice changes shape. What used to be a flat seat fee is now metered by the token, and it moves with every prompt your team sends. Your first instinct is to cap usage and switch on spend alerts. That instinct may be a mistake.&lt;/p&gt;

&lt;p&gt;This is not hypothetical. &lt;a href="https://thenextweb.com/news/microsoft-claude-code-retreat-ai-cost" rel="noopener noreferrer"&gt;Uber's CTO told The Information his organization burned through its entire planned 2026 AI coding budget in four months&lt;/a&gt;, with individual engineers spending 500 to 2,000 dollars a month on tokens. Microsoft answered the same math by pulling most Claude Code licenses from its Experiences and Devices division. Two of the most sophisticated engineering organizations on the planet, and their first public responses to metered AI were a blown budget and a retreat.&lt;/p&gt;

&lt;p&gt;So the decision in front of you is not "how do we cut the AI bill." It is "are we treating tokens as a cost to suppress, or as capital to compound?" Get it wrong and you either starve your best people or fund theater. Get it right and every token buys both an outcome today and an asset tomorrow.&lt;/p&gt;

&lt;p&gt;Building the AI-DLC harness at Betsson, I saw the same type of tasks task run in two different ways: one setup burned ten times the tokens of the other and produced similar to worse results, purely because of how context was fed, which model was chosen, and constraints provided. &lt;/p&gt;

&lt;p&gt;Cost is a design outcome, not a usage number. That is why engineering the system around the model is the right answer.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why the bill went variable
&lt;/h3&gt;

&lt;p&gt;For a while, AI in the IDE felt free; flat subscriptions hid the cost of inference and vendors absorbed the difference. That era is over. &lt;a href="https://cursor.com/docs/models-and-pricing" rel="noopener noreferrer"&gt;Cursor now prices agent requests by the token, against API rates&lt;/a&gt;, GitHub &lt;a href="https://thenextweb.com/news/github-copilot-signup-pause-agentic-ai-usage-limits" rel="noopener noreferrer"&gt;paused new Copilot sign-ups when agentic workloads outran plan prices&lt;/a&gt;, and from June 2026 &lt;a href="https://github.blog/news-insights/company-news/github-copilot-is-moving-to-usage-based-billing/" rel="noopener noreferrer"&gt;every Copilot plan is billed in metered AI Credits computed from token consumption&lt;/a&gt;. The reason is architectural, not commercial. Autocomplete used a narrow context at predictable cost; agentic tools ingest whole directory trees and run plan, edit, test, retry loops, and because the models are stateless the full transcript is re-sent each turn. A quick question and a multi-hour autonomous run can no longer rationally cost the same. For scale: &lt;a href="https://finance.yahoo.com/sectors/technology/articles/anthropic-quietly-doubles-estimate-much-220101627.html" rel="noopener noreferrer"&gt;Anthropic's own enterprise figures put Claude Code at around 13 dollars per active developer-day&lt;/a&gt;, a figure it recently doubled as usage deepened.&lt;/p&gt;

&lt;h3&gt;
  
  
  Cap or not to cap ?
&lt;/h3&gt;

&lt;p&gt;Faced with a variable bill, most organizations pick one of two reactions, and both repeat the mistake I described in &lt;a href="https://dev.to/cleberdelima/the-velocity-trap-why-your-ai-productivity-gains-are-an-illusion-o6o"&gt;The Velocity Trap&lt;/a&gt;: they manage the number that is easy to count, tokens, &lt;strong&gt;instead of the value those tokens produce.&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;When you cap the limits blindly it generates a &lt;strong&gt;"token anxiety".&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Engineers start thinking in cost-per-prompt, hesitate to hand the agent a hard problem, and write boilerplate by hand to save credits. You bought leverage, then taught people to be afraid to use it.&lt;/p&gt;

&lt;p&gt;The way out is to change what you count. The right model is not the cheapest per token; it is the cheapest per successful outcome. Report spend in dollars against something that shipped: cost per merged PR, with at least two guardrail metrics - PR size and revert rate - so nobody splits work or merges junk to move the number. &lt;/p&gt;

&lt;p&gt;But do not mistake the gauge for the engine. The numbers will tell you whether things are improving or not, but what improves them is the system around the model.&lt;/p&gt;

&lt;h3&gt;
  
  
  Tokens that compound
&lt;/h3&gt;

&lt;p&gt;Once the unit is an outcome, the goal stops being fewer tokens used and becomes &lt;strong&gt;fewer tokens wasted.&lt;/strong&gt; &lt;/p&gt;

&lt;p&gt;And don;t take me wrong you need to have caps and governance, but those should be bounded by the value it genereates.&lt;/p&gt;

&lt;p&gt;Then comes the part the cost conversation misses. &lt;a href="https://www.fastcompany.com/91561371/satya-nadella-is-asking-the-right-ai-question" rel="noopener noreferrer"&gt;Satya Nadella calls it token capital&lt;/a&gt;: the AI-embodied knowledge a firm owns and compounds, a factor of production beside human capital. The test he sets is sharp: &lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;can you say what, from last night's work, became knowledge your firm now owns? &lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The tokens you spend need not evaporate into a vendor invoice. The traces they leave, how specs become working software, which prompts and checks pass your gate, can be captured, owned, and used to make the next cycle cheaper. &lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Rent the intelligence, own the memory and the harness.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  Six moves that raise return on every token
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;1. Instrument telemetry, so every token has an owner.&lt;/strong&gt; The invoice tells you nothing about where the tokens went. Use the telemetry your IDEs and agent tools already emit, route requests through a gateway that attributes each one to a user, session, model and task, and read transcripts, not just dashboards. Feed it into one honest number: cost per merged PR in dollars. Pitfall: publishing raw consumption with no outcome beside it. Signal: every dollar maps to a user, a task and a shipped change.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Fix defaults and routing before you touch caps.&lt;/strong&gt; Do not use frontier models for non-frontier problems. Default your gateway to a strong mid-tier model and let engineers escalate by choice; route by step, a frontier model for planning, cheaper models for execution. In practice that is as simple as running sweeps on a Sonnet-class model and reaching for the top model only at the hard junctures. Honor the floor: if a task needs real intelligence, a larger model at low effort usually beats a small model straining at high effort. Pitfall: forcing model choice onto people; better defaults beat mandates. Signal: most calls served by non-frontier models with quality holding on your evals.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Treat caching as a first-class discipline.&lt;/strong&gt; Cache the stable prefix so repeated context is read back cheaply; &lt;a href="https://platform.claude.com/docs/en/build-with-claude/prompt-caching" rel="noopener noreferrer"&gt;cached reads cost around 90 percent less than base input tokens&lt;/a&gt;. Put durable content at the front, rules files, system context, tool definitions, long examples, and keep it immutable and append-only. The classic cache-breaker is a datetime variable in the system prompt that ticks every turn and silently invalidates everything behind it. Measure your hit rate, which the APIs return, and climb it. Pitfall: dynamic values inside the cached prefix. Signal: hit rate at 80 percent and above.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Keep context lean, and bound the worst case.&lt;/strong&gt; Every token in context is one you pay to re-read, and the dominant waste event is the runaway loop: a blind agent re-submitting the same failing fix with the full transcript each turn. Open fresh sessions per task, scope file context narrowly, and disconnect tools you are not using. Then add circuit breakers in the gateway that trip on repeated identical errors, on iterations without a repository change, or on a spend spike against the session median, plus hard-stop wallets that refuse calls at budget zero. Pitfall: compacting context endlessly instead of resetting. Signal: no session able to exceed its wallet.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5. Compound what you spend into token capital.&lt;/strong&gt; Persist the artifacts of each cycle, the machine-ready specs, the prompts and checks that passed your gate, the traces of how work got done, into owned memory and your knowledge base, the compounding argument of &lt;a href="https://dev.to/cleberdelima/continuous-fluid-flow-how-ai-is-compressing-the-software-delivery-cycle-3f20"&gt;Continuous Fluid Flow&lt;/a&gt;. Over time you can tune a smaller model on your own traces until it beats a prompted frontier model on repeatable workflows. Pitfall: letting hard-won knowledge leak into vendor logs you do not control. Signal: repeat tasks getting cheaper cycle over cycle.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;6. Make all of it legible to engineers.&lt;/strong&gt; None of this survives a team that does not understand it, and metered pricing turns cost intuition into a skill. Teach it directly: prompting that reduces ambiguity, context management, decomposing work into small testable steps. Invest hours, not just licenses, and promote power users as AI Champions, chosen by delivery impact rather than consumption. Pitfall: mistaking the loudest token-burner for the best practitioner. Signal: the median engineer working fluently inside the guardrails.&lt;/p&gt;

&lt;h3&gt;
  
  
  What to fund: the harness, not the caps
&lt;/h3&gt;

&lt;p&gt;The thing worth funding is not a tighter set of limits. It is the harness, everything around the model: the gateway and router, the workflows that break work into small checkable steps, the skills that carry your conventions, and underneath it all a structured knowledge base. That last piece is where most waste hides. If agents retrieve loosely, they pull in tokens they do not need and reason over noise, the cost side of the case I made in &lt;a href="https://dev.to/cleberdelima/your-coding-agents-are-drowning-in-context-you-pay-twice-in-tokens-and-in-precision-1pp7"&gt;Your Coding Agents Are Drowning in Context&lt;/a&gt;: knowledge modeled once and retrieved precisely is both more accurate and far cheaper than re-reading everything on every call.&lt;/p&gt;

&lt;p&gt;Sequence the build. Fund first the three cheapest, highest-return moves: telemetry, caching discipline, and default routing; they pay back inside a quarter. Fund next the knowledge base and the skills that scope context precisely. Fund later the heavier bets, local-first inference and distilling your own models on your traces. The gauge that proves or kills the program within a quarter is cost per merged PR: falling while revert rate stays flat and usage grows means the harness is earning its keep. If the only lever you pull is a cap, you are choosing Uber's surprise or Microsoft's retreat, just slower.&lt;/p&gt;

&lt;h3&gt;
  
  
  Start, Stop, Continue
&lt;/h3&gt;

&lt;p&gt;For executives. Start measuring cost per merged PR in dollars, and ask in every review what the last cycle added to your token capital. Stop the consumption leaderboards and blanket caps; both cost more than they save. Continue funding the harness and the knowledge base as infrastructure, not a one-off project.&lt;/p&gt;

&lt;p&gt;For engineers. Start writing append-only prompts, opening fresh sessions per task, and reading your own transcripts. Stop hand-writing boilerplate out of token anxiety, and stop leaving datetime variables in cached prefixes. Continue delegating hard problems to the agent and feeding what works back into shared context.&lt;/p&gt;

&lt;h3&gt;
  
  
  The strategic takeaway
&lt;/h3&gt;

&lt;p&gt;Suppression plateaus fast. What compounds is the infrastructure: routing that keeps getting smarter, caches that keep getting warmer, and a knowledge base that turns each cycle into a cheaper next one. That gap widens every month, because usage is rising either way. The only real question is whether your spend is buying assets or just buying invoices.&lt;/p&gt;

&lt;p&gt;So borrow Nadella's question and make it your own. What is your token capital, and what did last week's work add to it? If you cannot answer, you are renting the intelligence and letting the memory leak. Are you managing tokens as a cost to suppress, or as capital to compound? I would rather be argued with than applauded, so push back, and repost this to the leader whose AI bill just went variable.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>llm</category>
      <category>agents</category>
    </item>
    <item>
      <title>Your Coding Agents Are Drowning in Context: You Pay Twice, in Tokens and in Precision</title>
      <dc:creator>Cleber de Lima</dc:creator>
      <pubDate>Tue, 07 Jul 2026 23:29:35 +0000</pubDate>
      <link>https://dev.to/cleberdelima/your-coding-agents-are-drowning-in-context-you-pay-twice-in-tokens-and-in-precision-1pp7</link>
      <guid>https://dev.to/cleberdelima/your-coding-agents-are-drowning-in-context-you-pay-twice-in-tokens-and-in-precision-1pp7</guid>
      <description>&lt;p&gt;Look at what your coding agents actually pull into context. For a single task, the agent greps the repo, runs a broad vector search, and loads dozens of files and chunks that merely resemble the request into the window before it writes a line. You pay for that twice. Once in tokens, for context the model never uses and once in precision, because the few facts that matter are now buried in lookalike text, and a window crowded with noise makes the model reason worse. The reflex may be to buy a longer context window, but a bigger window only buys room for more noise. &lt;/p&gt;

&lt;p&gt;The fix is not more retrieval or a bigger window, it is structuring what you retrieve so the agent gets exactly what the task needs, which is cheaper and sharper at the same time.&lt;/p&gt;

&lt;p&gt;Building the AI-DLC harness at Betsson we found that replacing "search the codebase" with a knowledge layer the agents query is the best way to get that right. It can be by reverse engineer the codebase, or to build a full knowledge base fit for Agent utilization. An agent picking up a task now pulls the exact typed slice it touches, the contract it must satisfy and the decision that governs it, instead of paying to read whatever text sits near the task in embedding space.&lt;/p&gt;

&lt;p&gt;The answer is not a better embedding model or a bigger window. It is to give the agent a structured map of the system and let it fetch only the part a task needs. That map is an &lt;strong&gt;ontology&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  What an ontology is
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;An ontology is a &lt;strong&gt;machine-readable map of a domain: the things that exist, the typed relationships allowed between them, and the rules that constrain them.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It is the schema of a knowledge graph, defined before you extract anything. For a codebase it names entities like service, contract, data entity, and decision record, and edges like implements, depends on, and governed by. The meaning lives in the edges. &lt;/p&gt;

&lt;p&gt;The constraint is the feature: because the agent can only traverse relationships you defined, it stops inventing them, which cuts hallucination.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why precise context is cheaper and sharper
&lt;/h2&gt;

&lt;p&gt;Start with why plain vector based RAG fails; The retrieval ranks by similarity, which is statistical, not structural, so &lt;a href="https://www.moderndata101.com/blogs/how-ontology-solves-vector-embedding-challenges/" rel="noopener noreferrer"&gt;high cosine similarity does not mean a chunk is relevant&lt;/a&gt;, and the agent over-fetches to compensate. The problem is not vectors, it is what you feed them. Most stacks slice documents into arbitrary token windows, so a contract gets split across two chunks and three unrelated snippets get bundled into one.&lt;/p&gt;

&lt;p&gt;Even using more &lt;a href="https://www.linkedin.com/pulse/chunking-retrieval-strategies-rag-pipelines-code-jin-tan-ruan-reole/" rel="noopener noreferrer"&gt;advanced techniques&lt;/a&gt; like Semantic Search or AST-Based code chunking, you still only have a piece of knowledge, not the relationships it carries.&lt;/p&gt;

&lt;p&gt;Reverse the order. Let the ontology decompose your knowledge into atoms before anything is embedded: one contract, one decision, one entity per atom, self-contained and typed. Build the relationships with other atoms and present that to the agent in the format of a knowledge graph. &lt;/p&gt;

&lt;p&gt;That pays off two ways, and they weigh the same. The token cost drops, because you stopped fetching and reading filler. And the output gets better, because Your vector proximity search now runs over a curated layer, keeping the fast recall of embeddings while gaining the precision of chunks cut along meaning giving the agent only the atoms the task needs and it reasons over signal instead of noise. From a retrieved atom it then walks typed edges to the exact dependencies, each hop governed by the schema rather than guessed from proximity.&lt;/p&gt;

&lt;p&gt;The gain compounds. You compile the atoms and edges once, upstream, and every one you add makes the next task both cheaper and more accurate, while raw embeddings depreciate and get re-computed with each model generation. &lt;/p&gt;

&lt;h2&gt;
  
  
  Composable retrieval
&lt;/h2&gt;

&lt;p&gt;None of this means rip out vector search. Treat retrieval as composable, and get the order right: the ontology decomposes knowledge into atoms first, you embed those atoms, and at query time you route across two components you own, vector proximity over the curated atoms for broad recall, and graph traversal over the typed edges for relational precision. &lt;/p&gt;

&lt;p&gt;Similarity finds the entry atoms, then the query walks two or three hops to expand the exact subgraph the task needs, so the agent can both read it and write back what it learns. Keep the atoms, the schema, and the confidence scores in formats you own, so the layer is portable and governed like the rest of your platform rather than locked inside a black box. &lt;/p&gt;

&lt;p&gt;Industry is already working on solutions to implement this type of solution, Pinecone's &lt;a href="https://www.pinecone.io/product/nexus/" rel="noopener noreferrer"&gt;Nexus&lt;/a&gt; for example, packages this as a "knowledge engine, not a retrieval system": it compiles data into typed, cited artifacts upstream and answers a single declarative query that carries its own token budget, instead of ranking chunks at runtime and probably more solutions will emerge on this area that is a critical piece of agentic infrastructure.&lt;/p&gt;

&lt;h2&gt;
  
  
  The playbook to get there
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Start where retrieval overspends.&lt;/strong&gt; Find the tasks where the agent pulls the most context and uses the least of it, and the relational questions it gets wrong. Signal: tokens per task, and wrong-context answers in that area.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Decompose with the schema before you embed.&lt;/strong&gt; Agree the entity and edge types up front, and cut your knowledge into ontology-defined atoms that become the chunks, rather than fixed-size windows. Pitfall: model-invented types and near-duplicate edges. Signal: relationship types stay closed, and chunks map to whole units of meaning.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Reverse-engineer the baseline.&lt;/strong&gt; Capture entities, contracts, and decisions for the components in scope as versioned artifacts, not wiki prose. Signal: an agent can answer "what does this expose and depend on" from the artifacts alone.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Serve it over MCP and route composable.&lt;/strong&gt; Broad lookups to vector proximity over the atoms, relational questions to graph traversal; return the typed slice, not the repo. Signal: context precision, the share of retrieved tokens the model actually uses.&lt;/p&gt;

&lt;h2&gt;
  
  
  The strategic takeaway
&lt;/h2&gt;

&lt;p&gt;More retrieval is not more signal. It is a bigger bill and a more crowded window, and the answer suffers as much as the budget does. The lever is order: decompose knowledge with an ontology first, so your chunks are the exact atoms agents need, then let vector proximity and graph traversal work over that curated layer. Precise context is cheaper and sharper at once, and neither gain comes at the other's expense. RAG retrieves what looks related; an ontology defines what connects, and a composable layer lets you use both. Embeddings depreciate with every model you adopt; the ontology compounds with every atom you define.&lt;/p&gt;

&lt;p&gt;So decide what your harness feeds its agents: more text that resembles the task, or a structured slice they can reason over. If you think a longer context window closes this, make the case in the comments. If not, ask your platform team two things about the last task an agent ran: how much of the context you paid for did the model actually use, and how much of what actually mattered was buried in the rest?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>llm</category>
      <category>agents</category>
    </item>
    <item>
      <title>[Boost]</title>
      <dc:creator>Cleber de Lima</dc:creator>
      <pubDate>Thu, 18 Dec 2025 16:13:49 +0000</pubDate>
      <link>https://dev.to/cleberdelima/-2dda</link>
      <guid>https://dev.to/cleberdelima/-2dda</guid>
      <description>&lt;div class="ltag__link"&gt;
  &lt;a href="/cleberdelima" class="ltag__link__link"&gt;
    &lt;div class="ltag__link__pic"&gt;
      &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3607194%2Fb81b1058-0dcb-45d3-84ad-67e471a945a1.png" alt="cleberdelima"&gt;
    &lt;/div&gt;
  &lt;/a&gt;
  &lt;a href="https://dev.to/cleberdelima/continuous-fluid-flow-how-ai-is-compressing-the-software-delivery-cycle-3f20" class="ltag__link__link"&gt;
    &lt;div class="ltag__link__content"&gt;
      &lt;h2&gt;Continuous Fluid Flow: How AI Is Compressing the Software Delivery Cycle&lt;/h2&gt;
      &lt;h3&gt;Cleber de Lima ・ Dec 18&lt;/h3&gt;
      &lt;div class="ltag__link__taglist"&gt;
        &lt;span class="ltag__link__tag"&gt;#programming&lt;/span&gt;
        &lt;span class="ltag__link__tag"&gt;#ai&lt;/span&gt;
        &lt;span class="ltag__link__tag"&gt;#softwareengineering&lt;/span&gt;
        &lt;span class="ltag__link__tag"&gt;#softwaredevelopment&lt;/span&gt;
      &lt;/div&gt;
    &lt;/div&gt;
  &lt;/a&gt;
&lt;/div&gt;


</description>
      <category>programming</category>
      <category>ai</category>
      <category>softwareengineering</category>
      <category>softwaredevelopment</category>
    </item>
    <item>
      <title>Continuous Fluid Flow: How AI Is Compressing the Software Delivery Cycle</title>
      <dc:creator>Cleber de Lima</dc:creator>
      <pubDate>Thu, 18 Dec 2025 16:13:32 +0000</pubDate>
      <link>https://dev.to/cleberdelima/continuous-fluid-flow-how-ai-is-compressing-the-software-delivery-cycle-3f20</link>
      <guid>https://dev.to/cleberdelima/continuous-fluid-flow-how-ai-is-compressing-the-software-delivery-cycle-3f20</guid>
      <description>&lt;p&gt;Your developers are 55% faster. Your pull requests take 91% longer to review. Your deployment frequency is flat or declining. Welcome to the productivity paradox of AI-enabled development.&lt;/p&gt;

&lt;p&gt;After 15 years leading enterprise transformation programs across cloud, DevOps, and now AI, I have seen organizations repeatedly optimize the wrong constraint. AI has made coding faster, but coding was never the real bottleneck. Product decisions, quality assurance, deployment automation, and production learning loops now determine whether teams actually deliver value or simply generate code that queues up for review.&lt;/p&gt;

&lt;p&gt;This shift demands a fundamental rethinking of how work flows through delivery systems. The two-week cadence that defined Agile for two decades was designed for human-speed development. When AI compresses what took days into hours, time-boxed iterations become containers too large for the work they hold and too slow for the feedback it needs.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Bottleneck Has Moved
&lt;/h2&gt;

&lt;p&gt;The 2025 DORA Report delivers a sobering assessment: despite 90% AI adoption among developers, delivery stability declined 7.2% in organizations using AI coding tools without adequate governance. Individual productivity metrics improved while system-level throughput stagnated or declined.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Andrew Ng puts it bluntly: the bottleneck is now deciding what to build. When prototypes that took teams months can be built in a weekend, waiting a week for user feedback becomes painful.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;GitClear's analysis of 211 million lines of code reveals an 8-fold increase in duplicated code blocks between 2020 and 2024. Teams are producing more artifacts while the cognitive load on reviewers, testers, and operators accelerates beyond their capacity to absorb it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Continuous Flow Model
&lt;/h2&gt;

&lt;p&gt;The pattern emerging from both research and practice abandons time-boxed iterations entirely. Work moves through the system as capacity allows rather than waiting for arbitrary boundaries. Each work item progresses independently from specification through generation, validation, and deployment.&lt;/p&gt;

&lt;p&gt;The model operates through three parallel, continuous activities rather than sequential phases.&lt;/p&gt;

&lt;p&gt;AI-intensive delivery is the primary work stream. A cross-functional assembly of product, engineering, QA, and SRE works with full dedication on well-bounded goals. AI is embedded throughout: requirement refinement, architecture options, code generation, test creation, documentation. The team operates in focused mobilization sessions where AI proposes and humans validate in real time. Work flows through quality gates as fast as it can pass them, with WIP limits preventing the system from generating more than review capacity can absorb.&lt;/p&gt;

&lt;p&gt;Early-life support runs as a continuous responsibility. As each increment reaches production, the team monitors telemetry, triages issues, responds to user feedback, and makes rapid fixes. This happens for each deployment rather than batching support into a dedicated phase.&lt;/p&gt;

&lt;p&gt;Learning and assetization operate as an ongoing discipline. The team continuously extracts patterns, creates reusable templates and prompts, improves automation, and shares knowledge. This is where compound advantage gets built. Without this deliberate investment, organizations accelerate artifact production while learning velocity remains unchanged.&lt;/p&gt;

&lt;p&gt;AWS has documented similar patterns with their &lt;a href="https://github.com/awslabs/aidlc-workflows/tree/main" rel="noopener noreferrer"&gt;AI-DLC&lt;/a&gt; methodology. Practitioners experimenting with compressed cycles report that removing time boundaries forced them to develop critical skills around finding small slices of value and collaborating effectively. The intensity is high, but so is the learning velocity: feedback loops that took weeks now complete in hours or days.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Playbook for Continuous Flow
&lt;/h2&gt;

&lt;p&gt;Based on research from AWS, DORA, McKinsey, and on the implementation work with enterprise engineering organizations, here is a structured approach I'm using to move toward continuous flow.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 1: Assess Readiness Before Accelerating
&lt;/h3&gt;

&lt;p&gt;What: Evaluate your team's technical practices, cultural environment, and infrastructure maturity before removing time boundaries.&lt;/p&gt;

&lt;p&gt;Why it matters: The 2025 DORA Report's central finding is that AI does not fix teams; it amplifies what already exists. Strong teams with robust testing and fast feedback loops see gains. Struggling teams with tightly coupled systems see instability increase. 90% of organizations now have AI adoption, but 30% still do not trust AI-generated code.&lt;/p&gt;

&lt;p&gt;How to do it: Map your current SDLC and identify where bottlenecks exist before AI acceleration. Assess platform engineering maturity: automated testing coverage, CI/CD sophistication, and observability instrumentation. Establish baseline DORA metrics. Evaluate architecture for coupling, since tightly coupled systems cannot absorb AI-generated change velocity.&lt;/p&gt;

&lt;p&gt;Pitfall to avoid: Assuming tools alone will fix problems. AI amplifies existing dysfunction.&lt;/p&gt;

&lt;p&gt;Metric and signal: Clear identification of your top three bottlenecks with baseline DORA metrics established.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 2: Scale QA and Deployment Infrastructure First
&lt;/h3&gt;

&lt;p&gt;What: Implement AI-powered testing and progressive delivery infrastructure before accelerating development. QA capacity and deployment safety must expand before development velocity increases.&lt;/p&gt;

&lt;p&gt;Why it matters: 63% of teams cite QA as their biggest delay. When AI accelerates coding without scaling testing and deployment automation, quality degrades rather than velocity improving. Elite performers in DORA's research deploy 182 times more frequently than low performers while maintaining 8 times lower change failure rates. The difference is infrastructure that enables safe experimentation.&lt;/p&gt;

&lt;p&gt;How to do it: Pilot AI test generation that analyzes code changes to auto-generate scenarios as part of mobilisation and development, based on specifications and requirements not on ready code. Implement self-healing tests and predictive defect detection. Deploy feature flags enabling toggling features on and off, canary releases to 5-10% of users first, and automated rollback when anomalies are detected. Build observability dashboards with real-time visibility into performance and user behavior.&lt;/p&gt;

&lt;p&gt;Pitfall to avoid: Over-reliance on AI testing without human judgment for business logic. Feature flag debt from old flags not cleaned up.&lt;/p&gt;

&lt;p&gt;Metric and signal: Test creation time reduces 70%. Deployment frequency increases 2-5x. Change failure rate decreases despite volume increase.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 3: Establish AI Code Review Governance
&lt;/h3&gt;

&lt;p&gt;What: Create specialized review processes for AI-generated code and deploy automated quality gates.&lt;/p&gt;

&lt;p&gt;Why it matters: Code review has become the last-mile bottleneck. Reviewers take 26% longer for AI-heavy pull requests because they must check for hallucinated packages, pattern misuse, and duplicated code. Without governance, organizations accept risk they do not understand.&lt;/p&gt;

&lt;p&gt;How to do it: Create AI-specific review checklists checking for hallucinated packages, business logic verification, and security vulnerabilities. Implement PR tagging requiring AI assistance percentage notation, triggering additional review for PRs exceeding 30% AI content. Deploy automated quality gates catching duplication, complexity, and maintainability issues before human review.&lt;/p&gt;

&lt;p&gt;Pitfall to avoid: Treating automated review as replacement for human review. You need both.&lt;/p&gt;

&lt;p&gt;Metric and signal: Review time stabilizes despite volume increase. Percentage of issues caught in automated gates exceeds 60%.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 4: Form Multidisciplinary Teams with AI as Collaborator
&lt;/h3&gt;

&lt;p&gt;What: Assemble cross-functional teams where product, engineering, QA, and SRE work alongside AI agents in focused, uninterrupted synchronous sessions.&lt;/p&gt;

&lt;p&gt;Why it matters: AI-native development paradoxically requires more synchronous human collaboration, not less. The traditional pattern of writing a ticket, waiting for grooming, waiting for planning, waiting for development, waiting for review stretches decisions across weeks with constant context loss. Mobilization compresses that same decision density into hours of focused collaboration. When the full team is present with AI, questions get answered immediately, decisions happen in seconds rather than days.&lt;/p&gt;

&lt;p&gt;How to do it: Assemble volatile teams that form around well-bounded goals. Include product ownership for intent validation, engineering for technical judgment, QA for quality perspective, and SRE for operational awareness. Integrate AI agents as active collaborators embedded in every phase. Protect mobilization time ruthlessly from interruption. Structure sessions in two modes: Mob Elaboration where the team co-creates specifications with AI, and Mob Construction where AI generates while humans validate in real time.&lt;/p&gt;

&lt;p&gt;Pitfall to avoid: Treating mobilization sessions as optional meetings rather than protected deep work. Forming teams without all necessary disciplines.&lt;/p&gt;

&lt;p&gt;Metric and signal: Decision latency decreases from days to minutes during sessions. Output per mobilization session exceeds output from equivalent distributed async time.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 5: Pilot Continuous Flow with Learning Systems
&lt;/h3&gt;

&lt;p&gt;What: Experiment with removing time-boxed iterations on contained work. Implement WIP limits and quality gates. Establish systems where learnings compound into organizational assets.&lt;/p&gt;

&lt;p&gt;Why it matters: When AI enables idea-to-prototype cycles measured in hours, two-week sprints become containers too large for the work they hold. Continuous flow allows work to move through the system as fast as capacity permits. But AI accelerates delivery only if outputs compound. Repeated one-off code creates technical debt at AI speed. Organizations that build reusable prompt libraries and standardized patterns achieve higher productivity with each delivery.&lt;/p&gt;

&lt;p&gt;How to do it: Break work into small, well-specified items with clear acceptance criteria. Implement WIP limits based on review and validation capacity. Establish quality gates that work must pass. Maintain continuous production monitoring with focused support as each increment deploys. Dedicate ongoing time to documenting learnings and improving automation as a parallel activity rather than a batched phase. Build reusable AI assets tailored to your context. Consider asset creation as part of the definition of done, not allowing tasks to be completed without the automation, prompt fine-tuning, and context curation.&lt;/p&gt;

&lt;p&gt;Pitfall to avoid: Removing boundaries without implementing WIP limits. Neglecting learning time because it is not scheduled.&lt;/p&gt;

&lt;p&gt;Metric and signal: Learning cycles complete 2-3x faster. Cycle time decreases as flow improves. Reuse ratio across projects increases over time.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to Start, Stop, Continue
&lt;/h2&gt;

&lt;h3&gt;
  
  
  For Executives
&lt;/h3&gt;

&lt;p&gt;Start: Treating AI adoption as operating model transformation, not tool deployment. Allocating budget for progressive delivery infrastructure, AI testing platforms, and team training. Measuring learning velocity alongside delivery velocity.&lt;/p&gt;

&lt;p&gt;Stop: Declaring victory based on license adoption without delivery outcome linkage. Removing time boundaries without building prerequisite infrastructure. Ignoring downstream bottlenecks in QA, review, and deployment.&lt;/p&gt;

&lt;p&gt;Continue: Investing in platform engineering as the foundation AI amplifies. Demanding evidence that AI delivers value, not just activity.&lt;/p&gt;

&lt;h3&gt;
  
  
  For Engineers
&lt;/h3&gt;

&lt;p&gt;Start: Treating AI-generated code as untrusted input requiring validation. Building context packs and reusable prompts for your domain. Participating in mobilization sessions where AI proposes and humans validate.&lt;/p&gt;

&lt;p&gt;Stop: Accepting AI suggestions without reviewing the code. Treating every AI interaction as an isolated transaction. Ignoring downstream effects of accelerated code generation on reviewers and testers.&lt;/p&gt;

&lt;p&gt;Continue: Applying rigorous review standards to all code regardless of origin. Building expertise in context engineering and AI orchestration. Sharing successful patterns with the broader organization.&lt;/p&gt;

&lt;h2&gt;
  
  
  Strategic Takeaway
&lt;/h2&gt;

&lt;p&gt;Continuous flow is not about working faster. It is about learning faster.&lt;/p&gt;

&lt;p&gt;Time-boxed models treated learning as an outcome of shipping: deploy, measure, adjust. Continuous flow treats learning as a parallel activity: ship, support, extract patterns, compound assets, repeat. This shift from output velocity to learning velocity separates organizations building sustainable advantage from those generating code at AI speed while accumulating technical and organizational debt.&lt;/p&gt;

&lt;p&gt;The prerequisite investment is substantial. Progressive delivery infrastructure, AI-powered testing at scale, observability-driven development, and AI code review governance are not optional enhancements. They are the foundation that makes continuous flow possible without collapse.&lt;/p&gt;

&lt;p&gt;The 2025 DORA Report's finding is the essential insight: AI does not fix teams; it amplifies what already exists. Strong foundations plus AI acceleration equals compound advantage. Weak foundations plus AI acceleration equals compound failure.&lt;/p&gt;

&lt;p&gt;The organizations winning in this new era will not be the ones generating the most code. They will be the ones with the tightest learning loops and the most effective knowledge compounding. Speed without learning is just motion.&lt;/p&gt;

&lt;p&gt;If this challenges your current delivery model, that is the point. Share your perspective on continuous flow. Challenge the framework if you see gaps. The best operating models emerge from rigorous debate among practitioners who have tried these patterns in production.&lt;/p&gt;

</description>
      <category>programming</category>
      <category>ai</category>
      <category>softwareengineering</category>
      <category>softwaredevelopment</category>
    </item>
  </channel>
</rss>
