Many companies at this point have arrived at the same conclusion: that AI generated code can increase your engineering capabilities by leaps and bounds. How many of your developers are using tools such as Claude, Copilot, and Codex? What was your spend for the past 30 days? What skills did they use last week? Do your developers use the same skills or has each team researched and concocted different ways to do the same thing? These questions have been on my mind for the past few weeks. I think I have an answer: your AI engineering needs an owner.
3 stages of AI adoption
I see the industry at 3 distinct stages of adoption. The first two are about how people use AI, but only the third has an owner.
Curious
Individuals use AI on their own; everyone has their own prompts, skills, and habits. You might have an informal group who're particularly enthusiastic about AI, but by and large nothing is shared. Their stories and tasks are landing faster, so your organisation is improving little by little. Meanwhile your less excited or more sceptical developers are being left behind, and the gap between these two groups is widening.
Adopted
Your teams are now using AI together. They agree on standards and refine the skills in the repo and share habits. This is a great step forward from the last stage; your sceptical developers now have a channel for feedback, and feedback from the sceptic is incredibly valuable as I'll talk about later. As a whole, the team's capabilities improve.
AI native
There is a real person or team who owns the AI engineering capability. Skills/Plugins/MCP servers ship through one channel, rules are enforced and costs are measured. They work with other teams to improve AI for everyone: there's no difference between a core library which sits upstream from everything in your company and your AI capability.
Non-ownership costs you
Without an owner you pay in two ways. The first is the questions I started with: nobody can say what your agents are doing or what they cost. The second is quality. Both apply to Adopted as much as Curious, because a team can agree standards without any kind of measurement. The quality cost is hard to quantify but the signals should not be ignored. Drops in quality don't show straight away.
- DORA, Sep 2025 shows about 90% adoption rates of AI and that throughput is up but stability is down.
- Faros, Apr 2026 (vendor data) reports that bugs per developer increased by 54% as AI adoption grew, up from 9% in its earlier data.
- CMU, Jan 2026 reports that across 806 repositories adopting Cursor, the velocity gains faded after two months. Static analysis warnings increased by 30% and code complexity by 41%.
Why is it costing you
In my view the reason is everyone's AI is someone else's side project. Skills go stale fast (Anthropic, Jun 2026 saw analytical question accuracy decrease by 30% in a single month) and little inconsistencies in the generated code can go ignored.
With AI, small problems snowball an order of magnitude faster. Your newly created balls of mud can take a massive engineering effort to untangle, and so can poor skill performance when you cannot measure what your agents are doing.
The closest study I've found on ownership is DORA's AI Capabilities Model, Sep 2025: AI only improved organisational performance where the internal platform was good, and a clear AI policy made the gains bigger. Someone has to own the platform and the policy, and DORA itself recommends naming long-term owners.
There's one more thing that mostly goes unmentioned; Snyk, Feb 2026 (vendor data) found security flaws in 36.8% of 3,984 public skills from ClawHub and skills.sh.
What owning your AI engineering means
An owner answers both costs.
Observability
Measure what your agents are doing, what it costs, and (most importantly) whether it helped. This isn't only about usage metrics; we need to see what errors agents encountered, check that the correct skills fired, and correlate bugs in production with the agent session that produced them so we can improve things.
Standardisation
Stale skills and ignored inconsistencies are what happens when nobody enforces anything. LLMs are intrinsically non-deterministic so adding more text to a prompt is not good enough. Rules need to be enforced via tests and hooks. Skills need to be evaluated to ensure they produce a good output, and this needs to be determined for every new version of every model you are using. Skills and loops go through the owner's one channel, and nobody installs one straight from an unverified third party.
Scaling
A loop chains skills into a delivery process: spec to tickets, tickets to code, code to review, with checks between each step. Skills make individuals faster; a loop codifies your ways of working, and that is how you scale across teams.
Going from individual skills to organisational loops is a big job. Someone has to decide the order the steps run in, write the scripts that check each step's output, keep it fast, evaluate the whole thing against every new model, and offer options so teams can decide how they steer the AI. A developer can write a skill in an afternoon but few build a loop as a side project.
The pay-off is a loop that is deterministic even though the model inside it is not. The loop needs guardrails to ensure that your specs have been met in full: scripts that check facts. Small problems snowball, so those checks cannot rest on the agent's judgement.
What I have built
For the past few weeks I have been building this for myself. Skillworks is a plugin for coding agents with an agentic delivery loop, plus a dashboard called Studio. It is built by its own loop, so every change to Skillworks is made by the thing it is building. That is the fastest way I know to find out what does not work.
It is one repo on one machine, so treat what follows as what I found when I built it, not results from a team of fifty.
Studio reads the telemetry events coding agents already send. It lists every skill with how often it activated, what it cost, and which models and repositories it ran in. A Sessions view opens a single run and shows what the agent was given and what it did, step by step, including the errors it hit. Every commit the loop makes carries the id of the session that made it, so a bug in production can be traced back to the agent session that wrote it.
Building it taught me things a vendor dashboard will not tell you. Telemetry sends the signed-in user's email address on every event by default (measured Sep 2026), so switching it on is a decision for legal before it is one for engineering. And of 27 companies I checked (Sep 2026), none publishes which skills their developers use. Some say they log it; nobody shows it. The data exists and nobody owns it.
Rules live in the repo as data, and a skill turns them into architecture tests the team owns. Each test is proven red against a deliberate breach before it goes green on real code. I also measured what a rule in the agent's instructions is worth on its own (Sep 2026, 60 runs in four setups of 15): with no rule the agent followed it in 0 of 15 runs, and with the rule in 10 of 15. Better than nothing but nowhere near enforcement.
The loop takes a spec, splits it into tickets, and builds each ticket test-first. Three reviews follow (standards, spec and architecture), each in a fresh session, so no reviewer marks its own work. A script, not the agent, runs the team's checks and reads the results. At the end, every item in the spec gets a verdict and a script counts them. Anything missing is built in one more round, and anything still missing stops the loop and goes to a person.
The team owns the steering: its rules, its words, and what "green" means. The plugin owns the machinery.
I'm still working on it. The aim is for Skillworks to cover observability, standardisation and scaling, so an owner has one place to run them from.
Bringing sceptics along
The last job for an owner is bringing every team along, including the sceptics who were left behind at the Curious stage.
Training and coaching do not reach everyone. You need to show something that works and have the numbers as proof. The loop lets you walk into a team, pick up a ticket and let them be the judge of what is produced; their criticisms will improve your AI for everyone else. Inconsistencies with API design, verbose comments, idempotency concerns, and more are all ways for agents to generate bad, messy code.
This might also differ for every team and project, whether it is personal taste or something mechanical. They will want ways for the team to decide how to steer the agent but it should sit on top of a solid foundation: the loop.
Back to those questions
How many of your developers use AI? What did it cost last month? Which skills ran last week, and did they help? Do your teams share them, or has each built its own?
If nobody in your company can answer those, you are not short of tools. You are short of an owner.
Top comments (0)