Your AI rollout looks finished. Licenses are deployed, the adoption charts point up, most merged work touches an agent somewhere, and every team can show you something impressive it built last month. Then someone on the board asks the only question that matters: what did all of this change in the business? The room goes quiet, because the honest answer is smaller than the dashboards suggest.
You are not alone in that room. Deloitte's research on AI adoption finds that 84 percent of organizations have not redesigned jobs or workflows around AI, and that fewer than 60 percent of workers who have access to AI actually use it in their daily work. McKinsey's State of AI survey shows 88 percent of companies using AI in at least one function, while only 39 percent can attribute any bottom-line impact to it, and most of those say the impact is below five percent. Near-universal usage, near-zero transformation. That gap is not a technology gap. It is an operating model gap.
This article is the map for closing it: the full roadmap from adoption, which you can buy, to impact, which you have to build. For the past twelve months I have been writing about the individual aspects of this transformation, one article at a time: machine-ready specs, testing, continuous flow, loops, evals, token economics, the harness. Each of those pieces answered one question and deliberately left the bigger one open: in which order do these capabilities come together, and what does the whole journey look like when one organization runs it end to end? This is that article, which is why it is longer than my usual ones. The decision it asks of you is simple to state. Will you run AI adoption as a staged transformation, where each stage builds an asset the next stage depends on, or will you keep running it as a rolling tool deployment and hope that depth shows up on its own?
My experience is blunt on this one: choose the latter, and depth will not show up by itself.
Two terms carry the whole piece, so let me define them before anything depends on them.
The operating model is the full system that turns intent into running software: who decides what gets built, how work flows from an idea to production, where quality is checked, what gets measured, and who is accountable for what.
AI transformation is not a tool deployment sitting on top of that system. It is a redesign of the system itself.
Maturity is the second term, and it hides two questions that most scorecards collapse into one: are teams using AI well in their everyday work, and is the organization capable of sustaining and scaling what those teams do? You can score high on the first and fail the second, and actually this is where most companies are today.
Accenture and Carnegie Mellon's Software Engineering Institute published a maturity model in June 2026 built from a review of more than 100 existing frameworks, interviews with executives, and a survey of nearly 600 practitioners. The sentence from SEI's Ipek Ozkaya that should hang in your program office: "True AI maturity is not measured by how much AI an organization deploys, but by its ability to build trustworthy and resilient capabilities, rigorous engineering practices, and governance approaches aligned with business outcomes." That is this article compressed into a scorecard.
The moment that taught me the difference arrived on one slide of my transformation program at Betsson. The adoption telemetry was everything a CTO could ask for: nearly every engineer active weekly, agent features in regular use, a large share of merged work carrying AI assistance. Then we scored the same teams against a five-level maturity ladder, level one meaning ad-hoc individual experimentation and level five meaning self-improving systems inside fixed guardrails, and most teams landed in the bottom two levels. Both numbers were true. The telemetry measured opened doors; the ladder measured changed work.
The proof that the gap was real, and closable, came from one particular team that decided to really lean into the transformation. Same people, completely different approach and ways of working: intent turned into machine-readable specs, machine gates did the first pass of verification instead of tired human eyes, humans approved at defined checkpoints, and a small, focused multidisciplinary squad owned the outcome from definition to production.
That team's delivery cycle went from around 80 days to 5. Code generation was a small part of the story. The breakthrough was structural, and I have been telling it on stage ever since, because it is the whole argument of this article.
Rebuilding the processes and the collaboration around the new capabilities AI makes possible is the quantum leap, and it is what brings the biggest payoff.
Sit back and enjoy the reading: what follows is that journey, from breadth without depth to a whole organization running the way that one team did.
The map at a glance
Before going into the details, let's define the shape.
- Stage 0, Prepare: the tools, the models, and the connectors are evaluated and decided once; budgets and limits are set inside the tools; and the governance baseline and first harness standards are written before the crowd arrives. The asset is a decided stack with rules.
- Stage 1, Adopt: the first training waves run with chosen domains and chosen people, and their learnings feed the harness. The asset is proof, plus a trained first cohort.
- Stage 2, Structure: the harness evolves, the company knowledge base starts to compound, and specifications become machine-ready. The asset is an environment where agents work reliably at falling cost.
- Stage 3, Scale: the proven practice rolls out wave after wave to every team, carried by FDEs and harness engineers, and every new team onboards straight into the harness. The asset is the practice at full coverage.
- Stage 4, Verify: evals, systematic tests for AI behavior, and gates take over the first pass of quality. The asset is trust you can measure.
- Stage 5, Automate: loops run work unattended inside the gates, and a managed model portfolio serves them at the right cost. The asset is delivery capacity that no longer scales with attention.
- Stage 6, Reorganize: the ways of working change shape, product and engineering merge around the new speed, and the same playbook extends beyond engineering. The asset is an operating model that compounds.
Read together, the seven stages are the distance between the two numbers in the opening: adoption is what the early stages secure, and measured impact is what the last stage finally delivers.
Stage 0: Prepare. The asset is a decided stack, not a catalog
There is a stage before the rollout most people call the start, and it is the cheapest insurance in this whole journey. Before the first license lands in a builder's hands, four decisions need a first version: which tools the organization runs, what they are allowed to cost, which rules keep you compliant, and which working standards every team will follow.
The first is the stack decision, and it covers three lists, not one.
- Tools: run a short, honest evaluation of the coding assistants, agentic IDEs, and no-code agent builders against your own work, not against vendor demos, with security and cost vetting built into the process.
- Models: standardize which models are approved for which kind of work, one default per job plus vetted alternatives. This is standardization, not routing; the machinery that dispatches each request to the right model belongs to the harness and the portfolio in later stages, but the menu it dispatches from is decided here, because a menu every engineer composes alone is how spend and quality become impossible to compare.
- Connectors: evaluate the external MCP servers and third-party tools your agents will call, and treat them as the supply chain they are, because a connector receives credentials and reaches into your systems, so nothing gets on the approved list without a security review.
The output across all three is one approved list with three statuses, standard, pilot, and retired, a pilot gate for new entries, and a clear answer to the question every engineer will otherwise answer alone: which tool, which model, which connector, for which job. One boundary keeps this stage honest: it governs what you adopt from the market, not what you build. The MCPs and tools you build yourself belong to the harness work of stage two, and the economics of running models on your own infrastructure at scale belong to the portfolio decisions of stage five.
The second is the money, and it must be set up inside the tools on day one, not discovered on the first invoice. The billing model of this stack changed under everyone's feet: coding assistants moved from flat seats to metered, usage-based pricing, which means an agentic workflow can spend in an afternoon what a seat used to cost in a month. Uber's CTO described burning the entire planned annual AI coding budget in four months, and Microsoft answered the same math by pulling agentic tool access from thousands of its own engineers. So allocate budgets per team and per seat before the rollout, and use the admin planes the tools now ship: default models set to the standard tier, model allow-lists matching your approved menu, spend alerts at the levels where you want a conversation to happen, and hard caps only as the backstop against runaways. Two lessons from operating this. Defaults beat caps: most people never come close to their limits, so the real spend lever is which model the tool reaches for first, not how hard you cap the ceiling. And budget by persona, not by average: a small group of power users will legitimately spend many times the median because they are running the most agentic work, and a cap sized for the average punishes exactly the people getting the most from the stack; give them alerts and a review, not a wall. The organization-wide chokepoint, one gateway with wallets and circuit breakers, comes in stage five; stage zero sets the limits where the tools already provide them.
The third is the governance baseline, and if you operate in Europe this is not optional pre-work, it is law with dates attached. The EU AI Act is in force and phasing in: its prohibitions and its AI literacy obligation have applied since February 2025, which means organizations deploying AI must already ensure the people using it are adequately trained, and further obligations keep arriving through 2027. Three pieces need a first version before rollout. An acceptable-use policy that tells every builder what may and may not go into these tools: which code, which data, which customer information, under which vendor terms, with retention and training-on-your-data clauses checked per tool on the approved list. A risk classification of your intended uses against the AI Act's categories, so you know which are minimal risk, which carry transparency duties, and which would cross into high-risk territory and need a different governance lane before they ship. And named accountability: who approves a new AI use case, who owns incidents involving AI output, and where the record of those decisions lives. Thin, again, is fine: a two-page policy every team has actually read beats a governance framework nobody has finished. And notice the convenient overlap: the literacy the law requires is exactly the training stage one delivers, so compliance and capability are the same budget line here, not competing ones.
The fourth is the first harness definition. Stage two will build the harness in depth, but its skeleton must exist before adoption scales: the instruction-file template every repository starts from, the naming and state conventions, the guardrail baseline naming the actions no agent may ever take, and the default autonomy level a team starts at. Thin is fine; written is what matters. A one-page standard adopted by the first ten teams beats a perfect framework announced after fifty teams have invented their own.
The failure mode of skipping this stage is tool sprawl, and unlike the other failure modes it is almost invisible while it happens. Every team picks its own assistant, negotiates its own licenses, writes its own rules or none, and the organization discovers months later that it runs a dozen overlapping tools with no shared standard, ungoverned spend scattered across cost centers, and a security review backlog nobody planned. The money goes into duplicated licenses and, worse, into the consolidation program you will eventually have to run to undo it all, at many times the cost of deciding early.
You are ready to advance the moment the approved list, the budgets, the policy, and the one-page standard exist. This stage is measured in weeks, and it costs decision effort, not budget.
The asset of this stage, made concrete:
- Artifacts: the approved list covering tools, models, and external connectors, each with standard, pilot, and retired status and a pilot gate for new entries; the model standard, one default per kind of work plus vetted alternatives; the security review record per approved connector; the budget allocation per team and per seat, configured in each tool's admin plane with standard-tier defaults, allow-lists, spend alerts, and hard caps as the backstop; the acceptable-use policy with per-tool data boundaries; the AI Act risk classification of intended uses, with named accountability for approving new ones; the instruction-file template; the guardrail baseline naming the actions no agent may take; the default autonomy level per domain; the one-page harness standard every onboarded team starts from.
- Metrics: tools, models, and connectors per job category, which should be one standard plus at most one pilot each; share of teams on the standard stack; spend per seat and per team against budget, visible from the first week; share of spend landing on the standard model tier; time from a new entry appearing to an approve-or-retire decision; AI spend outside the approved list and unvetted connectors in use, both of which should trend to zero.
- Impact: every team that onboards in stage one starts inside the same rules, on the same models, through vetted connectors, with spend visible and bounded from the first week, so the harness work of stage two lands on one foundation instead of fifty improvised ones, finance never meets this program through a surprise invoice, and the consolidation program nobody budgets for never has to exist.
Stage 1: Adopt. The asset is people, not licenses
With the stack decided, everything starts with tools in hands, and this stage is where most organizations both start and quietly stop. The work is real: roll out the stack the previous stage decided, clear the remaining security and legal steps, and put it in front of every builder. But the deliverable of stage one is not deployed software. It is changed behavior, and behavior does not change by email announcement.
Unless your team is quite small, what works is running adoption in waves, with a repeating change cycle inside each wave, not a one-shot rollout. And the waves come in two distinct movements that ask for different disciplines, the first one here in stage one, the second in stage three once the harness is ready:
The first waves exist to prove and to learn, the scale-out exists to industrialize what they proved. Confuse the two and you either scale what was never proven or keep re-proving what already works.
Sequence the waves deliberately: start with the domains that are most ready and where value is densest, take them through the full cycle, and let their results pull the next wave in. Inside each wave the cycle is the same four moves:
- Art of Possible: Show the domain what is possible with demos built on its own work, not generic ones. You will be surprised by the reaction of some people when they see a good result they did not believe was possible.
- Training: Train in role-based cohorts measured in hours of practice, not a lunch-and-learn, so every wave graduates a group that worked with the tools on its real backlog.
- Champions Network: Grow champions inside the teams, embedded users who scale adoption peer to peer, because engineers believe colleagues before they believe a central program.
- Share the Success: Collect the wave's success stories and feed them forward as the marketing for the next one.
Then keep the waves coming, because the tools change monthly and most feature adoption happens long after the first rollout. Run a few waves and one pattern becomes hard to ignore: the same license produces a power user or a dormant seat depending almost entirely on the enablement around it, which is why hours of practice predict usage far better than tool choice does.
Who goes in the first wave matters as much as what the wave teaches, so compose it deliberately.
Do not just open enrolment, and never roll out everything to everyone at the same time. This is change management, not scheduling.
Open the enrolment to a selected audience only, and let them opt in. The idea here is to ensure the right people are on the first waves, but only the ones who really want to be there. Fill most of the seats with the people already leaning in: the ones experimenting on their own, asking for licenses, showing colleagues what they built. Give preference to your respected architects and senior developers who are willing to commit to an enterprise-level approach rather than a personal toolkit, because they are the cornerstones: when the people others already seek for advice adopt a new technology or way of working seriously, the rest of the organization tends to follow without being pushed.
Here is a tricky one: reserve a few seats, a few, not many, for your sharpest skeptics and open detractors. Their concerns are usually rational, the case I made in a previous article, and a skeptic who converts inside the first wave becomes more persuasive than any champion, while one who stays unconvinced hands you the failure catalogue for free, before it costs anything.
What you should never do is mandate the first wave onto the indifferent: forced adopters produce the dormant seats your dashboards will later mistake for a tooling problem.
That is the whole job of the first waves: not coverage, but proof, converted skeptics, visible wins, and the first honest version of what to teach everyone else.
And do not fund the waves alone. The vendors on your approved list want your adoption to succeed at least as much as you do, because their renewal depends on it, so negotiate enablement into the contracts while stage zero is still deciding the stack: training credits, instructor-led workshops, office hours with their engineers, certification seats, early-access programs. Partners can carry the surge too: bringing outside trainers into the first waves is cheaper and faster than building an internal academy before you know what to teach, and by the time the later waves arrive, your own champions can take over the room. The training budget shrinks considerably when the people selling you the tools co-fund the adoption of them.
Scaling to everyone is deliberately not this stage's job. The first waves exist to prove, to learn, and to feed the harness; carrying the practice to the whole organization is stage three, and it only starts once the harness is ready to receive teams.
Expect the dip. Productivity goes down before it goes up, because people are learning a new way to work while still being paid to deliver: METR measured experienced developers 19 percent slower with AI tools while believing they were 24 percent faster. People need time to adjust and adapt, and that shows up as a cost on cycle time that is recovered as the practice matures. In my experience the payback comes in sprints, not in months or quarters, so insist, and keep the focus on improving and compounding.
This stage fails in two directions, and both burn money. Stopping here is the license illusion, the failure from the opening paragraphs. Deloitte's researchers draw the distinction as adoption versus adaptation: adoption tells you someone opened the door, adaptation tells you they changed how they work. An organization that stops at stage one keeps paying a tool bill that grows with headcount while the value stays anecdotal, sits in a J-curve dip it never climbs out of, and reads dashboards that look like progress. Skipping it is worse: harnesses, gates, and agents rolled out to people who never built the judgment to use them become infrastructure for users who do not show up, and you will meet that cost again at every later stage, because a skipped asset does not disappear, it comes back with interest. You are ready to advance when the first waves have delivered their proof, cornerstones committed, skeptics converted or their objections documented, graduates still active a month later, and when your training has moved past prompting tricks into judgment: what to delegate, what to verify, when to stop trusting the output, how to optimize token usage and generate repeatable workflows.
The asset of this stage, made concrete:
- Artifacts: the first-wave plan, rosters composed for proof: volunteers, cornerstone architects and developers, a few skeptics; a role-based training curriculum with a learning path per discipline; vendor and partner enablement commitments written into the contracts; a named champion in every team; a success-story library tagged by domain; an AI literacy module inside week-one onboarding; adoption dashboards that count people changing how they work, not seats assigned.
- Metrics: first waves completed, with their learnings documented and handed to the harness team; weekly active users as a share of all builders, not of licenses; hours of structured practice per person; the share of trained people still active thirty days after training; the number of domains with a live champion.
- Impact: the J-curve dip gets shorter and shallower, usage holds after the novelty fades, and every later stage inherits users instead of seats, and the harness team starts stage two with a real backlog instead of theory. If thirty-day retention of trained users is high, this asset exists; if usage craters after the launch push, it does not, whatever the license count says.
Stage 2: Structure. The harness is the multiplier
Here is the reframe that separates organizations that scale from those that plateau: agent results are mostly not a model property. They are a property of the environment you run the model in. That environment has a name, the harness, and building it is the central engineering investment of the whole journey. I made this case in Loop Engineering and it deserves its full form here: the harness is everything around the model, the instructions, the tools it may call, the environment it runs in, the state it keeps between sessions, and the feedback that tells it whether its work passed. A prompt file is not a harness. A harness is those five subsystems working together, and the feedback subsystem, the explicit commands that verify work, returns more than any other investment.
The evidence for how much this layer matters is now quantified. A position paper on harness-induced versus model-induced variance found in its variance decomposition on SWE-bench tasks that swapping the harness moved success rates several times more than swapping the model, with harness effects around seven to eight times larger than model effects, enough to reverse model rankings in most of the configurations tested. A separate industry study, The Harness Effect, held tasks and models constant and changed only the orchestration layer: cost per task fell 41 percent, latency fell 44 percent, and token use fell 38 percent at the same output quality, though note it is authored by a vendor evaluating its own harness, so read the direction, not the decimals. The practical rule I give every team: when agent output disappoints, do not reach for a bigger model first. Check the harness. One well-written instruction file routinely outperforms a model upgrade, at a tiny fraction of the cost.
Be clear about the starting point and the destination, because they are far apart. The harness you inherit from stage zero is a skeleton: a set of instructions, the approved MCP tools, and the guardrails. Stage two is where that skeleton grows into an enterprise enabler, and it grows in three directions. First, codified workflows: the recurring jobs of each discipline, turning a ticket into a spec, generating and reviewing code, design to code, infrastructure changes, incident analysis, each captured as a written, versioned workflow that encodes how your best people do the work, so every team inherits it instead of reinventing it. Second, a model router with local and smaller models serving the non-frontier work that fills most of a delivery day, and frontier models reserved for the moments where judgment is dense; most tasks are not frontier problems, and paying frontier prices for them is pure waste. Third, telemetry and analytics on every agent run: tokens, cost, latency, acceptance, so cost and accuracy stop being impressions and become curves you can steer. This is the machinery behind the numbers above, and it is why the harness is the rare investment that cuts cost and raises accuracy at the same time.
None of this grows by itself, so staff it. Harness engineering is a central role, not a side duty: a small team of harness engineers, under the named owner, that treats the instruction standards, the hooks, the workflows, the router, and the telemetry as one product with a roadmap, and hardens what the first waves discover into the standard everyone else inherits. Make it a deliberate career path too: harness engineers make natural FDEs when the scale-out comes, and returning FDEs make the best harness engineers, because they have seen where the harness meets real teams and real deadlines.
One component of the harness deserves its own attention because executives rarely hear about it: hooks, small deterministic scripts that fire at fixed points in the agent's lifecycle, before a tool call, after a file write, at session start. Hooks are how you make the probabilistic system behave predictably at the moments that matter. A hook can compress a ten-thousand-line test log into the hundred lines that matter before the model ever reads it, saving tokens and confusion. A hook can block a risky action, a database migration, a production credential, before it executes rather than after. A hook can index every artifact the agent produces so the next session starts informed. Guardrails written in prose are wishes; hooks are enforcement.
One component deserves to be treated as a first-class company asset rather than a feature of the harness: the knowledge base. This is the shared brain every agent and every workflow draws on: standards, architectural decisions, incident history, domain meaning, the reasons behind the rules. Keep it unified but federated by domain, each area feeding and curating its own slice, because it compounds: the more it is fed, the better every workflow that reads it becomes, which makes it the one part of the stack that gains value with age instead of losing it. It must be engineered with precision, not volume, because as I argued in Your Coding Agents Are Drowning in Context, badly curated context makes you pay twice, in tokens and in precision. Give it a named owner and domain stewards, measure it on coverage and retrieval quality, and defend it in the budget like the asset it is, because context is the moat no competitor can copy.
Put those two ownership decisions together and you get the strategic line of this whole stage, the one I drew in Token Economics:
Rent the intelligence, own the memory and the harness.
Models will keep changing names, prices, and owners, and you should let them: intelligence is the part you buy in a falling market. The harness and the knowledge base are the parts you keep, and they are the only parts that compound.
Stage two is also where the specification discipline from From User Stories to Machine-Ready Specs becomes infrastructure rather than advice. User stories worked because human developers filled the gaps with implicit knowledge. Agents have no implicit knowledge; vague input produces wrong output at scale. Specs with explicit inputs, outputs, constraints, and acceptance criteria become versioned artifacts living next to the code, refined through pull requests by product managers and architects together. This is the first visible change in ways of working: product people start writing for two audiences, humans and machines, and the quality of their writing starts to bound the quality of the software.
In practice this stage does not wait for stage one to complete; it starts almost together with it. The harness team should be standing by the time the first wave begins, because those waves are its laboratory: every friction an early team hits, every workaround a cornerstone invents, every objection a skeptic raises is raw material the harness team hardens into the standard. Keep that loop running, wave learnings in, harness versions out, until the harness is reliable enough to receive teams that were never hand-picked. That reliability, not a date on a plan, is what opens stage three. One rule about timing matters more than any other here: structure starts per team, not per company. The moment a team comes out of onboarding it should enter the harness work, while other teams are still in stage one. If structure waits for the whole organization to finish adopting, the vacuum fills itself: every team invents its own instruction files, its own conventions, its own defaults, and what should have been one harness becomes fifty private ones that later have to be reconciled at painful cost.
The frustration arrives on schedule too, because teams working without structure get inconsistent results and rising bills, conclude the technology is overhyped, and disengage exactly when the discipline that would have fixed both was within reach. Stage zero's one-page standard is what makes this parallelism safe: teams can start early because they all start from the same page.
The failure mode of skipping this stage deserves a name: the model tax. It is paying frontier-model prices to compensate for an environment you never built: noisy context inflating every token bill, endless tool churn as teams shop for the model that will finally fix what the environment keeps breaking, and every upgrade producing a new round of surprises because nothing around the model was stable. The money leaks twice, once in inflated inference costs and once in the engineer-weeks spent re-tuning after each release, and neither leak ever carries the harness's name in the budget, which is why this tax goes unnoticed for years. You are ready to advance when the harness has a named owner and a roadmap like any product, when a new team can onboard agents in days because the environment tells them how, and when your specs are versioned and machine-readable in at least the pilot domains.
The asset of this stage, made concrete:
- Artifacts: an instruction file per repository, around a hundred lines that route to deeper docs on demand; hooks at the risky and noisy points of the agent lifecycle; reproducible environments with lockfiles; state files that survive sessions; explicit verification commands per repo; a versioned specs directory living beside the source; the codified workflow library per discipline; the model router configuration with local models serving the non-frontier work; telemetry and analytics on every agent run; the central harness engineering team, under a named owner, running it all as one product with a roadmap; and the knowledge base, unified, federated by domain, with a named owner and domain stewards.
- Metrics: days for a new team to onboard agents; tokens and cost per completed task; first-pass acceptance rate of agent output; clarification round-trips per feature; share of tasks served by local or smaller models without quality loss; knowledge-base coverage and retrieval quality; and an ablation score, the measured drop in success when you remove one harness subsystem at a time.
- Impact: the studies above put the range on the table: double-digit reductions in cost, latency, and tokens at constant quality from orchestration changes alone, and success-rate swings larger than any model swap. This is the stage where the two curves cross: cost per task falling while accuracy rises, with a growing share of work leaving the frontier tier and a knowledge base that makes every quarter's harness better than the last. If every model upgrade still produces a round of surprises, the asset is not there yet.
Stage 3: Scale. The asset is the practice, carried to everyone
With proof in hand from the first waves and a harness hardened from their learnings, scaling stops being a leap of faith and becomes an industrial operation. Still, it is a different discipline, and the moment to start is when proof stops being the question.
Be clear about what does not change: every scale-out wave still runs the same four moves as the first ones, its own demos on the domain's real work, its trained cohorts, its champions, its success stories feeding the next wave. What changes is the layer you add on top of that cycle.
The landing zone is different now: every scale-out wave onboards into the harness from day one, with the workflows, the routing, and the guardrails already in place. That landing is why scale is its own stage, and why it comes after the harness, never before. Scaling on raw tools is one of the most expensive mistakes on this map, because it makes every team pay for adoption twice: they learn a raw-tool way of working now, and have to re-adapt to the harness when it finally arrives.
This is where you borrow the strongest staffing pattern the AI industry itself uses: the forward deployed engineer, FDE for short. Palantir coined the role years ago. Instead of shipping software and hoping customers would figure it out, it embedded engineers with the customer, inside the customer's real problems, and let what they learned flow back into the product. The frontier labs have adopted the same model for exactly the problem you face at this stage: OpenAI deploys its hardest enterprise products through forward deployed engineers rather than self-serve, with the lessons of every deployment feeding back into the product, and Anthropic is building forward-deployed teams to carry its models into enterprises. Read that carefully.
The companies with the best models on earth do not believe the model sells or installs itself. They believe adoption is an embedding problem, and they staff it accordingly.
The internal version draws from two pools. Take the architects and senior developers who came out of the first waves committed, together with the harness engineers who hardened those waves' learnings into the standard, and rotate them into the next domains as embedded FDEs:
Not trainers who deliver a course and leave, but builders who sit inside the receiving team for a wave, take that team's real backlog, and construct the first AI-assisted workflows with the team, on its code, its constraints, its deadlines.
The difference against classroom training (which still happens, as part of the four-move cycle every wave runs) is the difference this whole stage runs on: a course transfers knowledge, an embedded engineer transfers a working practice, and the practice survives because the team watched it being built on their own work, not on a demo repository.
Give the rotation prime status: a named role, a defined tour of one or two waves, and recognition that this is senior, key work, because it is.
The people you want will not volunteer for something that looks like a side duty.
What separates an FDE program from internal consulting is one thing: the return path. Everything an FDE builds in the field, the instruction files, the workflows, the answers to hard objections, flows back into the central harness and knowledge base, so every embedding makes the next one shorter and cheaper. Without the feedback loop and the two-way communication, FDEs are expensive consultants leaving local snowflakes behind; with it, FDEs become how the harness learns. And they are also a seed: an AI-fluent engineer embedded in a domain is exactly the shape stage six will scale beyond engineering, and your first FDEs are where those pods will come from.
The failure mode of stalling before this stage is the pilot island: proof stranded in a handful of teams while the rest of the organization keeps working the old way. It is a quiet failure, because everything inside the island looks like success, the pilots deliver, the demos impress, the case studies write themselves. The money is what never arrives: the program has paid for the platform, the harness, and the first waves, but captures value only where the island's edge happens to fall, and when the board reads the pilot's numbers next to the organization's flat ones, it concludes the technology overpromised when the truth is that the practice was never distributed.
You are ready to advance when every delivery team has been through a wave, when onboarding into the harness is simply how a new team starts, and when the practice holds its shape in teams nobody hand-picked.
The asset of this stage, made concrete:
- Artifacts: the scale-out wave calendar covering every delivery team; the FDE rotation, first-wave cornerstones and harness engineers on named tours with a defined return path for what they build; per-domain onboarding kits generated from the harness; the champions network as a standing structure, not a launch-phase one.
- Metrics: coverage, the share of teams through a wave and working inside the harness; time and cost per wave, falling as embeddings compound; workflows and instruction files returned to the platform per embedding; variance of adoption depth across teams, shrinking as the practice standardizes; thirty-day retention holding even though rosters are no longer hand-picked.
- Impact: the proof of the first waves becomes the norm of the organization: every team lands on the same harness with a working practice from day one, enablement cost per wave falls while the harness gets richer, and the value case stops resting on pilots because the whole delivery organization runs the practice the pilots proved.
Stage 4: Verify. Trust becomes measurable, or autonomy stays a demo
Every stage so far still assumes a human reads everything the machine produces. That assumption is the ceiling on the whole investment, and this stage is where you lift it, carefully. The instrument is the eval: a systematic test that measures how well an AI system performs on your specific tasks, made of inputs and success criteria encoded in grading logic. If your teams can already recite unit tests and CI, the translation is direct: evals are the tests, and the eval gate is the CI, for behavior instead of code. A small, well-designed private eval tells you more about which model and which configuration to use than any public benchmark, because public benchmarks are saturated and contaminated, and your workload is not on them.
Build the estate from what you already have. Real merged pull requests that fixed real issues become tasks. Review comments become learned rules. Incidents become permanent regression scenarios, which is how the operations loop closes: every production failure becomes a check that the same failure cannot ship twice. Grade with three families of graders, deterministic code checks where right and wrong are objective, model-based judges with calibrated rubrics where nuance is needed, and periodic human spot checks to keep the judges honest. Then wire the result where it can act: into the pipeline, so that a score drop blocks a deploy the same way a failing test does. The cautionary tale for why this must be automated comes from a published post-mortem in my research: a team shipped a support agent after three days of careful manual testing, and it quoted a deprecated refund policy in production for five days. No crash, no hallucination drama, just a plausible wrong answer that manual testing missed and an eval suite would have caught.
Evals are also how you survive the upgrade treadmill. Models now ship monthly, and each new one interacts differently with your standing instructions, a decay I described in Instruction Debt. Without an eval gate, every model upgrade is a leap of faith; with one, it is a canary run and a diff. The same applies inside a model family: effort settings and automatic fallbacks mean the configuration serving your request may not be the one you selected, so log what actually served every request and test the configuration, not the brand name.
This stage is where the autonomy ladder starts to climb, one earned rung at a time: suggest-only, then edit-with-review where a human approves diffs, then execute-in-sandbox, then autonomous-within-guardrails, where an agent may open pull requests but never merge to protected branches and never touch production credentials. The rung you are on should be a written policy per domain, not an accident of which team moves fastest. And the split that makes any of it safe: the agent that writes is never the agent that checks. A separate checker, ideally on a different model, with one standing instruction the writer never sees: treat the change as wrong until the spec and the existing tests prove otherwise, and never edit a test to make it pass.
The failure mode of skipping this stage is the review wall, and it shows up in the delivery statistics. The 2025 DORA report, drawing on around five thousand practitioners, found AI adoption now improves throughput but still increases delivery instability: more speed, less stability, exactly what you would predict when generation scales and verification does not. Industry benchmark data from LinearB's 2026 report, which I could not link directly and cite from my research notes, shows AI-authored pull requests merging at less than half the rate of human ones and idling several times longer before a reviewer opens them: generation scaled, trust did not. The money goes into your most expensive people: senior engineers spending their days reading machine output line by line, pull requests waiting in queues while the agents that produced them sit idle, speed purchased at agent prices and delivered at the pace of human eyes. That is The Velocity Trap at organizational scale. You are ready to advance when eval scores gate deploys in the pipeline, when your checker agents reject real work every week and people fix the work rather than the checker, and when a model upgrade is a routine canary run rather than a project.
The asset of this stage, made concrete:
- Artifacts: an eval suite of graded tasks built from your own merged PRs and incidents; written grader rubrics with anchored scoring examples; the CI configuration where a score drop blocks the deploy; a model-upgrade canary pipeline; the written autonomy ladder per domain; the checker agent's standing instruction, versioned like code.
- Metrics: eval pass rate per model and configuration; checker rejection rate, which must stay above zero every week; escaped-defect rate in AI-assisted lanes against your human-reviewed baseline; agreement between judge scores and human spot checks; days to certify a new model.
- Impact: verification stops consuming your senior engineers' reading hours and runs at machine speed instead; the instability the DORA data warns about gets pulled back toward your baseline while throughput keeps its gains; and a model release becomes an afternoon of canary runs instead of a quarter of anxiety. The review queue, the place where the gains of stages two and three go to die, stops being the bottleneck.
Stage 5: Automate. Loops do the work, the portfolio serves it
Only now, with a harness that makes agents reliable and gates that make them trustworthy, does it pay to carefully reduce human intervention. This is loop engineering, the discipline I covered in depth in Loop Engineering: stop being the person who prompts the agent, design the system that does it instead. A loop is a goal, a way to check progress against it, and a rule for when to stop. The build order is strict, and it is the clearest example of why this whole roadmap is sequenced: get one manual run reliable first, capture what made it work as a written skill, wrap the skill in a loop with a gate and a stop condition, and only then put it on a schedule. Scheduling something you never made reliable by hand is how loops blow up overnight. Notice that the order repeats the stages: the skill assumes a harness, the gate assumes evals, the schedule assumes both.
Not everything deserves a loop. The filter is four conditions that must all hold: the task repeats at least weekly, something can automatically reject bad output, the agent can do the work end to end, and done is objective rather than a matter of taste. Miss one, and a good prompt is still the better tool. Where the conditions hold, the compounding is real: at the far end, Stripe's engineers have described an internal pipeline that merges more than a thousand machine-written pull requests a week through deterministic and LLM gates, a reported figure I keep on hand not as a target but as proof that the ceiling is nowhere near where most teams assume. The metric that keeps loops honest is cost per accepted change, logged weekly: below roughly half of proposals accepted, a loop costs more than it returns, and the fix is almost always the gate, not the prompt.
Automation at this scale changes what you buy, which is why the model portfolio belongs in this stage. Loops consume tokens without fatigue, so the bill stops being a per-seat license and becomes metered infrastructure, the capital decision I described in Token Economics. Three moves keep it governed.
First, put an LLM gateway in front of everything before the volume arrives: one chokepoint that meters spend with hard budgets, trips a circuit breaker when a loop iterates without progress, caches repeated queries, and logs which model actually served each request.
Second, scale the routing the harness began: the model router from stage two becomes governed policy at the gateway, frontier models where the task is judgment-heavy, local and cheaper models where it is not, effort settings tuned per task rather than maximum by default. The routing policy is a living document your eval estate validates, because a cheaper model that fails the evals is not cheaper.
Third, hold optionality: keep a self-hosted open-weight tier in the portfolio, but enter it with honest numbers. Independent cost analyses in my research notes put the break-even for self-hosting a large model at around eleven billion tokens a month at sustained utilization, figures you should re-run against your own vendor quotes, and at low utilization self-hosting costs more than premium APIs. For most organizations it is a late-stage move justified by volume, jurisdiction, or control rather than sticker price.
Control deserves one concrete story, because it is the least obvious reason to hold portfolio optionality. When Hugging Face investigated the breach of its own infrastructure, an attack that ran on an autonomous agent executing over 17,000 recorded actions, the hosted frontier models it reached for first refused much of the forensic work: their safety guardrails could not tell an incident responder from an attacker. The forensics ran instead on a self-hosted open-weight model on Hugging Face's own hardware, and their engineers' published advice is now mine too: have a capable model you can run on your own infrastructure vetted and ready before an incident, because the day you need it is the wrong day to start.
The failure mode of this stage run too early is automated chaos: loops amplifying whatever they sit on, unverified output merging at machine speed, and a token bill growing faster than the value it buys. AI multiplies everything, and if the foundations are wrong it multiplies chaos. You will see it first in the one metric this stage lives by: cost per accepted change rising while the acceptance rate sinks below the point where the loop returns less than it burns, which is the machine telling you the gates beneath it were never real. You are ready to advance when at least one loop has run for a quarter with flat or falling cost per accepted change, when the gateway gives finance a number it trusts, and when your routing policy survives contact with your evals.
The asset of this stage, made concrete:
- Artifacts: a loop spec per automated workflow, its goal, its per-pass rule, its stop condition; the skill files that encode how the work is done; a state board that outlives every session; the gateway configuration with budgets, circuit breakers, and caching; a routing policy the evals validate; one vetted self-hostable model with a tested deployment recipe on the shelf.
- Metrics: cost per accepted change, weekly, per loop; acceptance rate, with roughly half as the floor below which a loop loses money; tokens per accepted change; cache hit rate at the gateway; spend against budget per team; the share of traffic served by each tier of the portfolio.
- Impact: delivery capacity stops scaling with human attention. Work is discovered, executed, and verified while nobody watches, at a unit cost finance can see and cap, and the same gateway that meters the spend is the kill switch when a loop misbehaves.
Stage 6: Reorganize. The operating model catches up with the machines
Everything to this point can be built inside the engineering function. The last stage cannot, because the constraint has moved again. When generation is fast and verification is automated, the bottleneck becomes decision latency: how long intent waits between the person who holds it and the system that can execute it. Squeezing that latency means changing the organization, not the tooling, and this is where the series' longest-running arguments converge. The fixed-cadence ceremonies I questioned in The End of Agile and the compressed delivery chain I described in Continuous Fluid Flow stop being provocations and become the operating instructions.
The delivery unit that works at this stage is small and temporary: two or three people spanning product, design, and engineering, plus multiple agents, plus the harness, owning one outcome end to end from definition through production.
What that squad runs is not a sprint. It is the AI-DLC cycle, three phases turning as one continuous loop: Inception, Construction, Operations. The names sound familiar; the content is not.
Inception is where quality is born, and it is a working session, not a document trail. The squad sits together, product, design, engineering, with the AI in the room proposing: it decomposes the intent into units of work, drafts the stories and acceptance criteria, surfaces the edge cases and risks, asks the clarifying questions nobody wrote down. The humans correct, decide, and remove ambiguity, and the guardrails and success metrics are set here, before a line of code exists. The output is a machine-ready specification, versioned like code, plus the checks that will judge the result. A few focused hours of this compress what used to be weeks of sequential refinement, and every ambiguity removed here is cost avoided everywhere downstream, because an agent executes a vague spec at exactly the speed it executes a precise one.
Construction is fast precisely because Inception was thorough. Agents generate code and tests inside the harness, the gates from stage four do the first pass of verification, and the humans validate live: architecture, trade-offs, the things a spec cannot fully carry. The unit of progress shrinks from the two-week sprint to cycles measured in hours or days, and something structural happens to the oldest split in engineering: building and reviewing merge into one act, because the senior engineer is no longer typing the code and then waiting for a colleague to read it, the senior engineer is verifying while the agents produce.
Operations is where the loop closes instead of ending. Deployment is AI-assisted, observability watches at machine granularity, and agents run the first pass of incident analysis. But the point is what flows backward: incidents become guardrails, so the same failure cannot ship twice; learnings become specs, so the next Inception starts smarter; production telemetry and user behavior feed the backlog, so the loop restarts with better information than it began. Close the week the way I argued in Continuous Fluid Flow: a stretch of early-life support where the squad watches what it shipped, then a compounding session that turns the cycle's learnings into reusable assets, new skills, better guardrails, improved specs, instead of into a list of opinions. Paradoxically, all this automation needs more synchronous human collaboration, not less: a few hours of the right people deciding together replaces weeks of asynchronous alignment, because the machine can execute at whatever speed the humans can decide.
Now hold your ceremony calendar against that loop and watch what survives. The daily standup dies, because status lives on the state boards both humans and agents write to, and telemetry reports better than people do; the meetings that remain exist to decide, not to inform. Sprint planning becomes the Inception workshop, scheduled when an intent is ready rather than when the calendar says so. Story-point estimation goes with it, replaced by measured cycle-time distributions, the collapse I argued in The End of Agile. The sprint review becomes outcome review at the gates: shipped increments and eval results, not slideware demos. And the retrospective becomes that compounding session, the one ceremony that gains weight in this world, because it is the only one whose output is an asset rather than a feeling. The test for any surviving meeting is simple: does it make a decision or produce an asset? If it does neither, it is queue time wearing a calendar invite.
Product management changes the most, which is why product belongs inside this transformation from stage two onward, not as a late guest. When code stops being the constraint, knowing what to build becomes the expensive skill. The product role shifts from managing a backlog of stories toward curating intent: writing specifications precise enough to be contracts, deciding what is worth building at all, and validating outcomes rather than implementation.
The strongest external evidence that pairing domain judgment with engineering fluency is the unlock comes from Uber's agentic pods: the company paired its most AI-proficient engineers with domain experts from business functions, gave each pod two weeks, and ran sixteen pods across sixteen functions in two months. Capital allocation analysis went from 15 hours to 30 minutes; financial pacing reports from two days to ten minutes; marketing quality assurance from two weeks to under an hour. The lesson Uber drew is the one I keep repeating: the biggest wins came from redesigning the whole workflow, not from automating a single task inside the old one. And that is also the proof that stage six is not an engineering stage at all: the same playbook, harness, gates, pods, runs in finance, operations, and marketing, which is where the second wave of value lives.
The same blurring happens on the operations side, and it is just as deep. Operations stops being the downstream department that inherits what delivery throws over the wall, because in the AI-DLC cycle it is a phase of the same loop the squad owns. The SRE sensibility sits inside the squad; the toil that made "you build it, you run it" an empty slogan is now carried by agents watching telemetry, triaging incidents, and drafting the first root-cause analysis; and operational signals stop being complaints and become premium product input, feeding the next Inception directly.
Quality moves the same way: the QA role stops being a manual pass at the end and becomes quality engineering, the discipline that designs and supervises the eval estate, the argument I made in Testing Reinvented carried to its organizational conclusion. Define, build, run, and learn stop being four departments' jobs and become four phases of one team's loop.
Put all of it together and you can see the shape of a future product development organization.
Small, end-to-end squads formed around outcomes.
They can be composed of two or three people plus an agent fleet, dissolving and reforming as intents change. A platform group running the shared machinery, the harness, the eval estate, the gateway, the knowledge base, as internal products with roadmaps. Product people embedded in squads as curators of intent and architects of specification. Quality engineers who own evals the way SREs own reliability. And far fewer pure coordination roles, because most of the layers in today's org chart exist to move information between people who do not sit together, and in this model the specs, the state boards, and the telemetry carry that information instead.
The chart flattens, and the people who remain are the ones whose judgment cannot be automated. The early AI-native companies already show this silhouette, tiny teams with revenue per employee at multiples of the traditional software average, and while an enterprise will never copy a fifty-person startup, the direction of the silhouette is the same.
Be honest about the path, though: you do not get there by announcing a reorganization. You get there team by team, the way this whole journey has moved, letting the squads that master the loop become the template the next ones copy.
Stage six is also where you decide, deliberately, how far autonomy goes. There are now organizations running with no human in the code path at all: StrongDM's software factory ships production security software under a charter that code must not be written or reviewed by humans, with three engineers writing specifications and evaluating outcomes against held-out scenarios the coding agents never see. I do not offer that as your target; for regulated domains it may never be permitted, and it is a bet most boards should not take today. I offer it as the far end of a dial you should set on purpose, domain by domain, in writing. A regulated core capped at supervised autonomy is a defensible policy. An unwritten "wherever the tools take us" is not a policy at all.
The failure mode here is the faster typist: keeping the old organization while running new machinery inside it. AI grafted onto handoff-heavy, ceremony-paced delivery produces faster typing inside the same queue, and the queue still sets the speed: intent still waits days for a prioritization meeting, approvals still travel the same chain, and agents that could build in hours wait on decisions that take weeks. The money goes to capacity you bought and cannot use, the most expensive waste in this whole journey, because it looks like success on every dashboard except cycle time. The exit condition of stage six is the impact the whole journey was for: output rising faster than headcount, with people making the decisions and machines carrying the volume, measured honestly rather than assumed.
The asset of this stage, made concrete:
- Artifacts: squad charters for the two-to-three-person pods and their end-to-end ownership; the AI-DLC cycle definition, Inception, Construction, Operations, as the squad's operating manual; the Inception workshop format and its spec templates; the ceremony map, what died, what replaced it, and what each replacement produces; the value dashboard where every initiative carries a baseline, an owner, and a KPI; the two-lens maturity heatmap; the autonomy ceiling register per domain; the kill-review calendar.
- Metrics: cycle time from idea to production; decision latency, how long intent waits for the person who can decide; the share of squads running the full cycle end to end, from Inception through Operations; output to headcount, delivery per engineer or revenue per employee; the share of initiatives with signed-off baselines; measured value against the fully loaded cost of the program.
- Impact: this is where the two numbers from the opening finally meet. Cycle-time compression of the 80-days-to-5 shape stops being one team's story and becomes the operating norm, workflow redesigns of the kind Uber's pods delivered turn days of work into minutes outside engineering too, and the board conversation changes from adoption percentages to a value line finance has signed.
Why the order is the product
Read the stages backward and the dependency chain is visible. Reorganizing around AI speed presumes loops that deliver unattended. Loops presume gates that catch bad work automatically. Gates presume evals worth trusting. Evals presume a harness stable enough that results mean something. The scale-out presumes that harness carried to every team, and the harness presumes people who use the tools well. The people presume a stack somebody decided. Every arrow in that chain is a place I have watched organizations fail by skipping: loops on top of no evals produce confident garbage on a schedule; evals on top of no harness measure noise; harnesses handed to untrained teams become tools nobody uses, with a maintenance bill attached. Read the seven failure modes back in one line, tool sprawl, the license illusion, the model tax, the pilot island, the review wall, automated chaos, the faster typist, and they share one cause: a stage skipped. Boris Cherny, who created Claude Code, mapped the same territory from the practitioner side as five steps from gated access to AI-native operation, with the sharp observation that the step most teams botch is scaling agent count before the verification loop has earned trust. His ladder and mine are the same lesson at different altitudes: the constraint at each step is never tokens. It is bottlenecks and guardrails, found and built in order.
The stages are not a calendar. A strong organization runs them overlapped, piloting stage-five loops in one domain while stage three is still rolling out in another, and it advances team by team rather than waiting for the whole company to clear a stage together: the onboarded teams move to structure while the rest are still adopting. What the order forbids is not overlap but inversion: no team should run a later-stage practice on top of an earlier-stage gap. That single rule is most of the governance you need.
Running it as a strategy, not a project
The staged roadmap is the delivery half. The other half is how you run it from the top, and here is the shape of what this work has taught me, in principles rather than particulars.
Run it as a named program with an executive sponsor, not a collection of initiatives inside the technology function. The numbers this transformation moves, revenue per employee, cost of delivery, customer experience, belong to the business, so accountability for them must sit where those numbers live, with the platform and capability work beneath.
Give the program a guiding principle with teeth: every AI initiative earns its place with measurable business value, and no initiative launches without a baseline and the KPIs it is expected to move.
Set targets the honest way around: measure baselines first, commit to directional ambitions after, and hold one commitment that needs no baseline at all, that the program must at minimum return its own cost in measured value. Then enforce the uncomfortable half of the discipline: when a use case does not move its KPI for a full quarter, it gets a mandatory review with three exits, kill it, pivot it once, or write down exactly why it stays. Zombie initiatives are how AI programs lose the board's trust.
Measure maturity on both axes you met at the start of this article, quarterly for team adoption, semi-annually for organizational capability, and let the scores drive where the enablement effort goes next. Scoring teams on a five-level ladder sounds bureaucratic until you see what it changes: adoption stops being an opinion, lagging areas become visible while they are still cheap to fix, and the board conversation shifts from anecdotes to a trajectory. And write down what you will not do, in the architecture rather than in a policy document: the actions agents may never take autonomously, the domains where a human gate is permanent, the thresholds where automation stops and judgment begins. Those commitments are not brakes on the strategy. They are why you will still have a license to operate when a competitor's ungoverned agent makes the news.
The playbook: seven moves to start on Monday
1. Place yourself on the map with numbers, not opinions. What: score your organization against the seven stages, and your teams against a five-level maturity ladder per work category, documentation, feature development, quality, operations. Why: every failed transformation I have examined misplaced itself at the start, usually one stage too generous. How: run the scoring workshop with three pilot teams first; the SEI and Accenture model gives you an externally benchmarkable rubric to adapt. Pitfall: letting teams self-report; scores inflate a full level. Signal: one honest heatmap the whole leadership team accepts without re-arguing it.
2. Stand up the adoption engine as a permanent cycle. What: domain-tailored demos, role-based training in hours not minutes, embedded champions, success stories feeding the next domain. Why: behavior change is 80 percent of the outcome and it decays without reinforcement. How: pick the two most willing domains, run the full cycle, publish the wins internally, repeat quarterly; make AI literacy part of week-one onboarding. Pitfall: training everyone on prompting and no one on judgment. Signal: weekly active use rising without mandates.
3. Appoint a harness owner and ship version one. What: one accountable owner, product not project, for the instruction files, tools, hooks, state, and verification commands your agents run inside. Why: the harness moves outcomes more than the model does, and unowned harnesses rot. How: start with a hundred-line instruction file that routes to deeper docs, explicit verification commands per repo, and two hooks, one that compresses noisy logs before the model reads them, one that blocks the actions you never want an agent to take. Pitfall: writing a manual when the agent needs a map. Signal: a new team onboards agents in days, and when output disappoints, people check the harness before blaming the model.
4. Build the eval estate and wire it into CI. What: a private suite of graded tasks built from your own merged PRs, review comments, and incidents, gating deploys and model upgrades. Why: it converts trust from a feeling into a number, and it is the asset that makes every later stage safe. How: start with twenty tasks in one domain, grade with deterministic checks plus one calibrated LLM judge, block the pipeline when scores drop, and run every model upgrade as a canary against the suite. Pitfall: building evals once and letting them go stale; they are a living artifact. Signal: a checker rejects real work every week, and people fix the work.
5. Put the gateway in before the loops. What: one chokepoint for all model traffic: budgets, circuit breakers, caching, full logs including which model actually served each request. Why: metered costs arrive with automation, and retrofitting governance onto a running fleet is painful and expensive. How: deploy a gateway with hard per-team budgets and a circuit breaker that halts any loop iterating without repo change; add a routing policy validated by your evals. Pitfall: treating self-hosting as a cost play at low volume; it only pays at sustained scale or for control reasons. Signal: finance trusts the number, and a runaway loop stops itself before anyone notices.
6. Pilot one loop and one pod, in parallel. What: one unattended loop on verification-friendly work, nightly CI triage, dependency updates, documentation drift, and one two-week pod pairing your most AI-fluent engineer with a domain expert outside engineering. Why: the loop proves stage five mechanics; the pod proves the playbook travels beyond engineering, and each generates the next round of internal believers. How: loop follows the strict build order, manual run, skill, gate, schedule, and reports cost per accepted change weekly; pod follows Uber's shape, understand the work first, build alongside the person who does it, validate with others who do the same job. Pitfall: piloting where failure is expensive. Signal: the loop's acceptance rate holds above half; the pod ships something the domain expert actually uses after the pod ends.
7. Make value the gate for everything. What: a value dashboard where every initiative carries a baseline, an owner, and the KPI it claims to move, reviewed on a fixed cadence with the kill rule in force. Why: it is the difference between a portfolio and a pile. How: baseline before launch, validate with holdouts where the money is material, and count only what finance signs off. Pitfall: counting activity, tokens consumed, agents built, PRs assisted, as value. Signal: at least one initiative killed or pivoted per quarter; paradoxically, that is the sign the discipline is real.
The strategic takeaway
What separates adoption from impact is everything the machines cannot bring with them, and that is exactly what compounds in this journey: the harness, the eval estate, the knowledge layer, the written skills, the redesigned ways of working. Every one of those assets makes the next stage cheaper and every model upgrade more valuable, which is why organizations that build them pull ahead quietly for a while, and then the gap is suddenly obvious. What plateaus is everything most programs celebrate: licenses, adoption charts, prompting skill, review capacity. Those scale with headcount and attention, and headcount and attention are exactly what this technology stops rewarding. The cost of the gap is not visible this quarter, because stage-one organizations and stage-five organizations both have impressive demos. It becomes visible when the compounding curves separate, and by then the distance is not a budget line. It is years of accumulated assets the lagging organization has to build from zero, against a competitor whose machines are already feeding their own improvement.
And if you keep only one funding rule from all of this, keep this one: fund people and the harness first, verification second, automation third, and the heavy options last, then hold the whole program to a single quarterly gauge, measured value against fully loaded cost. Everything else in this article is detail on top of that sentence.
So place your organization on the map, honestly. Which stage are you in, which stage does your board believe you are in, and what would it cost to make those two answers the same?
Take this map to your next leadership meeting and ask those questions out loud. The discussion that follows will tell you more about your real stage than any dashboard, and whatever it reveals, the next move is already on the map.
Top comments (0)