<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Unlocked Consulting</title>
    <description>The latest articles on DEV Community by Unlocked Consulting (unlocked-consulting).</description>
    <link>https://dev.to/unlocked-consulting</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Forganization%2Fprofile_image%2F14558%2F8ab9a1ad-58fd-4486-952e-705f43e737b1.png</url>
      <title>DEV Community: Unlocked Consulting</title>
      <link>https://dev.to/unlocked-consulting</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/unlocked-consulting"/>
    <language>en</language>
    <item>
      <title>The hours AI saved you have nowhere to go</title>
      <dc:creator>Helkyn Coello</dc:creator>
      <pubDate>Mon, 14 Sep 2026 20:33:06 +0000</pubDate>
      <link>https://dev.to/unlocked-consulting/the-hours-ai-saved-you-have-nowhere-to-go-2oba</link>
      <guid>https://dev.to/unlocked-consulting/the-hours-ai-saved-you-have-nowhere-to-go-2oba</guid>
      <description>&lt;p&gt;A manager at the company started building with AI last year. He described what he wanted, iterated with the model, clicked through the result, and fed back the cases he had forgotten until the behaviour was right. An afternoon of that produced something that ran. He handed the branch to engineering and considered the work finished.&lt;/p&gt;

&lt;p&gt;The branch sat for months, even though nobody was obstructing it. Review capacity hadn't changed, the architect's week hadn't changed, and the definition of "ready to publish" was the same as it had always been, with the same queue sitting in front of it.&lt;/p&gt;

&lt;p&gt;An afternoon of work entered a pipeline, and the pipeline did what it was designed to do: absorb it at its own pace. The manager was measuring whether the thing worked. Engineering was measuring whether it could be operated, secured, tested and supported. Both were right. And the afternoon he saved bought the company nothing, because the step that got faster was never the step that set the delivery date.&lt;/p&gt;

&lt;p&gt;That is the problem this article is about: AI accelerates the step in front of the person using it, and the steps behind that person stay exactly as slow as they were. The aggregate numbers say it is happening everywhere.&lt;/p&gt;

&lt;h2&gt;
  
  
  spend doubled while integrated workflows halved
&lt;/h2&gt;

&lt;p&gt;Two numbers from the same survey. ServiceNow and ThoughtLab asked 4,500 executives across nineteen countries for their 2026 maturity index, and the third edition lets them compare against last year.&lt;/p&gt;

&lt;p&gt;AI spend rose 110% in a single year. The share of organizations reporting streamlined, integrated workflows across business functions fell from 30% to 16%.&lt;/p&gt;

&lt;p&gt;Spend doubled. Integration halved. The report's own explanation is fragmented platforms carrying a new wave of agent sprawl on top of them, which is a polite way of saying the tools multiplied faster than the plumbing.&lt;/p&gt;

&lt;p&gt;It is the only integration figure in this year's research that has a previous year to compare against, and it runs backwards.&lt;/p&gt;

&lt;h2&gt;
  
  
  where the saved hours go
&lt;/h2&gt;

&lt;p&gt;Follow it one step at a time, because each step is defensible on its own and the outcome is still near zero.&lt;/p&gt;

&lt;p&gt;You buy licences and hand out tools. Reasonable enough; the per-seat cost is trivial against a senior salary. An individual who used to need three days to produce a piece of work now needs three hours. Also real; I have watched it happen at my own desk. But the workflow still allocates three days, and the ticket still waits for the weekly planning call. Two approvals sit in front of deployment, both sized for the latency of a human who had to read a document and think about it. So the individual, done by Tuesday, picks up the next ticket, and then the one after that. More work arrives at those two approvals every week than before. The approvals process it at the same speed they always did. Everything they cannot absorb sits and waits and becomes a backlog.&lt;/p&gt;

&lt;p&gt;The company paid for speed and got a longer queue.&lt;/p&gt;

&lt;p&gt;BCG's 2026 survey of 11,749 employees puts a number on the gap: 61% have limited or no guidance on what to do with the time AI saves them, and 45% are not reinvesting it into anything strategic. The hours are real. They just have nowhere to land.&lt;/p&gt;

&lt;p&gt;In the same ServiceNow data, 59% are past piloting agentic AI and only 9% report meaningful progress on autonomous multistep workflows. Zero percent have built anything resembling a cross-functional, self-improving agentic operating layer. Only 16% have replaced fragmented legacy systems with an integrated platform, and 41% still name siloed data as a major obstacle. Cloudera's survey of 1,270 IT leaders (vendor-run, so weigh it accordingly) has 79% saying they cannot access all the data their initiatives need and only 30% with fully integrated sources. Adoption figures like these are worth reading carefully, because three studies measuring agentic adoption this year returned 17%, 23% and 59% by counting three different things.&lt;/p&gt;

&lt;p&gt;Everything above comes from the people who bought the tools. Valtech asked the people who use them: a thousand professionals already working with AI daily. Asked where the most value would come from, 4.3% said more pilots. Integration of tools and data ranked first.&lt;/p&gt;

&lt;h2&gt;
  
  
  what 95% actually bought
&lt;/h2&gt;

&lt;p&gt;Five percent of organizations in the ServiceNow data are redesigning work. The rest are pointing agents at workflows that already existed, in the shape they already had.&lt;/p&gt;

&lt;p&gt;Deloitte's January 2026 survey of 3,235 leaders lands on the same spot from a different angle: 84% have not redesigned jobs or the nature of work around what AI can now do. My favourite pair in the whole report: 53% considered fewer management layers and smaller, more autonomous teams, and 16% actually moved to them. The idea got as far as a slide.&lt;/p&gt;

&lt;p&gt;McKinsey have two surveys on this. Their State of Organizations report, with 10,018 executives: 88% deploying, 81% reporting no meaningful bottom-line gain. In their mid-2026 follow-up, which asks the same questions as last year, use rose to 89%. The share saying AI moved their profit at all stayed at 37%. The share getting significant profit from it stayed at 6%. What separates that 6% is not budget. Roughly three-quarters of them redesigned their workflows, against one-quarter of everyone else. McKinsey's own summary: conviction in AI is growing faster than the returns anyone can attribute to it.&lt;/p&gt;

&lt;p&gt;All of that is self-report, and self-report on productivity is generous. Which is why the software engineering data matters more than it looks.&lt;/p&gt;

&lt;p&gt;Faros AI, which sells engineering analytics and has a stake in this conclusion, instrumented 22,000 developers and compared each team's lowest-adoption quarters against its highest. Developers finished 34% more tasks. Each piece of work then waited 441% longer for someone to check it before it could go out. And the speed at which finished work actually reached users did not change.&lt;/p&gt;

&lt;p&gt;That is the individual getting faster and the organization standing still, measured with telemetry rather than a questionnaire. Writing the code got cheap. Review, QA and deployment stayed exactly as expensive as they were, and everything that got written piled up in front of them. Faros is blunt that tightening review is the wrong response, because review was never the thing that changed.&lt;/p&gt;

&lt;p&gt;DORA, a non-commercial research programme surveying close to five thousand developers, states the mechanism in one line: AI amplifies what is already there. Their 2024 edition put a number on the downside: for every 25-point increase in AI adoption, teams shipped 1.5% less and what they shipped broke 7.2% more often. More AI in an unchanged process did not mean more delivered. It meant slightly less, arriving slightly worse.&lt;/p&gt;

&lt;h2&gt;
  
  
  three moves that change the process
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Decide where the saved hours go before you buy the tool.&lt;/strong&gt; If an analyst saves six hours a week, somebody has to say what those six hours are now for. More time with clients. The project that has been waiting since March. One fewer hire next year. If nobody names it, the hours get absorbed back into the same job and nothing changes. "We got more productive" is not an answer to where the time went.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Count the steps, don't just speed them up.&lt;/strong&gt; A three-hour task inside a five-approval chain is still a two-week task. The cheapest redesign move is deleting a handoff, and it is almost always resisted more than a six-figure platform purchase, because a handoff belongs to somebody.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Make each item cheaper to review.&lt;/strong&gt; If the queue in front of review is the constraint, the fix is not more reviewers or faster ones. It is that each thing arriving needs less looking at. Put a human on the plan while the model is still writing it, so the implementation starts from something somebody already shaped. Have the repository check every incoming branch against its own architecture rules and refuse what does not pass. By the time a branch reaches a person, it has cleared both. Same reviewer, same week, far less to read per item. That is what got the manager's afternoon of work into production.&lt;/p&gt;

&lt;h2&gt;
  
  
  where the hours AI saved should land
&lt;/h2&gt;

&lt;p&gt;The pattern in all of this is that AI speeds up the step a person does, and everything that step feeds into stays exactly as slow as it was. The spending is real, the individual gains are real, and the process is still shaped for the pace it had before.&lt;/p&gt;

&lt;p&gt;So take one workflow you have already put AI into and answer two questions. Where are the saved hours supposed to go, and who decided that? If nobody can answer this, you bought speed for one step and a longer wait for everything after it.&lt;/p&gt;

&lt;p&gt;That question is most of what an AI workflow audit does in its first week, and you can run it yourself on a whiteboard before anyone signs anything.&lt;/p&gt;

&lt;p&gt;Original article: &lt;a href="https://unlockedconsulting.ai/blog/hours-ai-saved-have-nowhere-to-go" rel="noopener noreferrer"&gt;https://unlockedconsulting.ai/blog/hours-ai-saved-have-nowhere-to-go&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>devops</category>
      <category>management</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Everyone Measures AI Usage. 70% Can't Measure What It Returned.</title>
      <dc:creator>Helkyn Coello</dc:creator>
      <pubDate>Fri, 11 Sep 2026 23:15:07 +0000</pubDate>
      <link>https://dev.to/unlocked-consulting/everyone-measures-ai-usage-70-cant-measure-what-it-returned-5a6l</link>
      <guid>https://dev.to/unlocked-consulting/everyone-measures-ai-usage-70-cant-measure-what-it-returned-5a6l</guid>
      <description>&lt;p&gt;Anthropic surveyed 132 of its own engineers about Claude Code. Merged pull requests per day rose 67 percent. Daily use of the tool climbed from 28 to 59 percent. Self-reported productivity gains ran between 20 and 50 percent. But then someone checked the organization's delivery dashboard and saw that the delivery metrics had not moved.&lt;/p&gt;

&lt;p&gt;That gap is the whole subject of this piece. A tool can be used constantly, rated highly by the people using it, and leave no trace on the numbers a business actually runs on. The measurement problem underneath it is bigger than one company's coding assistant. McKinsey found that 30 percent of leaders could say where the time AI freed up actually went. The other 70 percent could not. Seven out of ten organizations have people spending less time on tasks and no idea whether that turned into anything. Ask the question from the other side and the answer is just as thin: in Gartner's 2025 survey, 22 percent of leaders said their AI tools had returned significant value, a share that lands where McKinsey, Deloitte and ServiceNow each arrived measuring it their own way.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Usage and tokens are costs, not returns&lt;/strong&gt;&lt;br&gt;
Two numbers get reported as if they answered the ROI question, and neither does.&lt;/p&gt;

&lt;p&gt;Adoption tells you whether anyone is using the thing. It is a leading indicator and a useful one, but a tool with high adoption and no measured outcome has produced activity, not value. Token spend tells you what the tool costs to run. It belongs in the calculation, on the cost side, and watching it closely tells you nothing about whether the work it produced was worth having.&lt;/p&gt;

&lt;p&gt;Both are easy to pull from a dashboard, which is exactly why they get reported. The number that matters sits one step further out and takes real work to produce: whether the company made or saved a defensible dollar. Everything below is how you get to that number without lying to yourself on the way.&lt;/p&gt;

&lt;p&gt;So when a business case lands on a desk claiming a figure in saved dollars, the honest question is not whether AI helped. It is whether the number is real. In most cases I have looked at, it is inflated, and usually in two separate places at once.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Inflation one: soft hours counted as hard dollars&lt;/strong&gt;&lt;br&gt;
Here is the calculation almost every AI business case runs. The tool saved each person two hours a week. Multiply the hours by the hourly cost of those people, add it up across the team, and report the total as money saved.&lt;/p&gt;

&lt;p&gt;The hours are usually real. The dollars usually are not, because the budget did not change. Nobody was let go, no contractor was dropped, no line item fell. What happened is that a group of salaried people have slightly lighter weeks, and the company pays them exactly what it paid before.&lt;/p&gt;

&lt;p&gt;Freed hours become real money in three situations: the time is redeployed onto work that generates value, or it lets you avoid a hire you were about to make, or it lets the same headcount produce more of something you sell. If none of those is true, the saving is soft, and soft savings do not survive contact with a CFO who can see the budget did not drop.&lt;/p&gt;

&lt;p&gt;This is where the 70 percent from the opening returns. If seven in ten leaders cannot say where the freed time went, then most of the "hours saved times hourly rate" figures in circulation are soft hours nobody traced to an outcome, dressed up as hard dollars. The fix is not complicated. Label every dollar of claimed value as hard or soft, and report the two separately. The number gets smaller, and much harder to dispute.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Inflation two: calendar time counted as labor time&lt;/strong&gt;&lt;br&gt;
The second inflation is subtler, and I see it most in engineering cases. A feature "took three weeks" before and "takes one week" now, so the case dollarizes two weeks of saved time at an engineer's rate.&lt;/p&gt;

&lt;p&gt;The problem is that three weeks was never three weeks of work. Some of it was a ticket sitting in a queue, some was waiting on a review, some was a dependency that had not shipped. Cycle time, the calendar span from request to delivery, is not the same as labor time, the hours a person actually spent. Converting the calendar span to dollars at an hourly rate invents labor that no one performed.&lt;/p&gt;

&lt;p&gt;Speed is still worth reporting, but as a rate: this class of work now moves through 40 percent faster. It becomes money only when moving faster captures something real, most often revenue that arrives earlier because the thing shipped sooner. When it does not, a faster cycle is still a genuine improvement worth reporting as speed, but it is not a number you convert into dollars.&lt;/p&gt;

&lt;p&gt;There is a related trap in trusting self-reported speed at all. The one controlled study I know of that timed the same developers with and without AI, METR, found a measured slowdown of 19 percent against a self-reported gain of 20 percent. The developers were sure they were faster. The clock disagreed. Whatever you build your ROI on, it should not be a survey asking people how much time they think they saved.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What actually produces dollars&lt;/strong&gt;&lt;br&gt;
Strip out the inflations and there are only two mechanisms by which an AI tool produces money, and each converts to dollars through a different bridge.&lt;/p&gt;

&lt;p&gt;The first is acceleration: someone does a task they already did, in less time. The bridge is hours saved multiplied by the hourly cost of that person. If a task dropped from four hours to two and a half and happens eighty times a month, that is 120 hours a month, and at their hourly cost you have a real figure. Then you subtract rework, because generated output a person has to redo never saved the time it appeared to.&lt;/p&gt;

&lt;p&gt;The second is avoided work: a task stops happening at all. A support ticket the knowledge base resolves is a ticket a human never touches. The bridge here is not an hourly rate, it is the full cost of one whole interaction: the total monthly cost of the function divided by the number of interactions it handles. If support costs 20,000 a month and handles 2,500 tickets, each avoided ticket is worth 8 dollars, and that 8 already carries the tooling and overhead, not just one agent's wage. Using the acceleration bridge here, an hourly rate, would undercount it.&lt;/p&gt;

&lt;p&gt;Both of those are illustrative arithmetic, not figures from any client. The point is the shape: pick the mechanism, pick the matching bridge, and do not mix them.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The denominator nobody writes down&lt;/strong&gt;&lt;br&gt;
A return needs a cost to divide by, and this is where the tokens finally belong. The total cost of an AI tool per month is the build cost amortized over the months it will run, plus token spend, plus infrastructure, plus maintenance, plus any human review of its output. A build that took 150 hours and will run two years is not a 9,000-dollar hit this month; it is a few hundred a month spread across its life. Token cost is simpler than it looks: measure the average number of tokens one output consumes, multiply by the price per token, and you have a stable cost per output to hold against the value that output produces.&lt;/p&gt;

&lt;p&gt;A project can carry many metrics, but it has a single ROI. Each metric adds its own slice of value to the same numerator, and every slice divides by that same total cost. The one discipline this requires is avoiding double-counting: if a token cost already sits inside a per-unit figure, it does not also go in the denominator, and if two metrics describe the same saved dollar from two angles, you keep one of them.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The math is deterministic, which is why it gets skipped&lt;/strong&gt;&lt;br&gt;
I argued in an earlier piece that AI belongs at the edges of financial analysis and never in the arithmetic itself. Measuring your own AI's return is the same shape seen from the other side. The calculation is deterministic: a subtraction, a multiplication, an hourly cost, a total. No model is required, and none should be trusted with it.&lt;/p&gt;

&lt;p&gt;The reason it gets skipped is not difficulty. It is that the honest number is almost always smaller than the inflated one, and smaller numbers are harder to carry into a budget meeting. A defensible small number survives scrutiny and an impressive large one does not, and the second time a leader is caught reporting soft hours as hard dollars, the whole program's credibility pays for it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;But isn't agentic AI supposed to be exempt from ROI?&lt;/strong&gt;&lt;br&gt;
There is a serious version of the opposite argument, and Gartner makes it: early agentic AI is experimental, and organizations that demand a proven business case before they will touch it risk being outpaced by the ones that treat it as something to iterate on. That is right, as far as it goes. You do not gate a two-week experiment behind a formal ROI model, and pretending you can forecast the return on something genuinely new is a guess dressed as a forecast.&lt;/p&gt;

&lt;p&gt;But two different claims get folded together under that banner. "Do not require a business case before you experiment" is defensible. "Do not measure what it returned" is not, and the first is routinely used to justify the second. An experiment you never measure is not an experiment, it is a purchase with no follow-up. The reason for not gating early work behind ROI is to buy yourself room to find the value, which only means something if you then check whether you found it. Exempting agentic AI from a business case at the start is reasonable. Exempting it from measurement forever is how the share of companies that see real value stays stuck at one in five.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Where to start&lt;/strong&gt;&lt;br&gt;
Capture the baseline before you deploy anything, because once the old way is gone you cannot reconstruct how long it used to take. Measure one real unit of work end to end, from request to delivered, including the waiting, and see whether that number moved. Stop reporting seats deployed, tasks completed, prompts submitted and self-reported speed, and start reporting the one figure that ties to a customer or a budget. Label every claimed dollar hard or soft, and report the net dollars alongside the percentage, because the percentage moves with whatever you put in the denominator while the net figure does not.&lt;br&gt;
_&lt;br&gt;
Doing that across every process you run, and writing down what each would need in order to actually change a budget line, is a workflow audit. It takes longer than pulling a usage chart. It is also the only version of the exercise that produces a number you can defend.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Questions this raises&lt;/strong&gt;&lt;br&gt;
&lt;em&gt;How do I put a dollar value on time saved?&lt;/em&gt; &lt;br&gt;
Multiply the hours saved by the hourly cost of the person who saved them. Count it as a real saving only if that time is redeployed, avoids a hire, or produces more sellable output. Otherwise it is a soft number and should be labelled as one.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Should token spend count as part of ROI?&lt;/em&gt; &lt;br&gt;
Yes, as a cost, in the denominator, and never as a benefit. Track cost per unit of useful output if you want a figure to hold against value, and make sure you are not counting the same tokens in two places.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Can one project have several ROIs?&lt;/em&gt; &lt;br&gt;
No. A project has one ROI. Several metrics can each add a slice of value to the same numerator, divided by the same total cost. If you end up with two ROIs for one tool, you have either double-counted or mixed two projects together.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://unlockedconsulting.ai/blog/everyone-measures-ai-usage-70-can-t-measure-what-it-returned" rel="noopener noreferrer"&gt;https://unlockedconsulting.ai/blog/everyone-measures-ai-usage-70-can-t-measure-what-it-returned&lt;/a&gt;&lt;/p&gt;

</description>
      <category>financial</category>
      <category>ai</category>
      <category>roi</category>
    </item>
    <item>
      <title>Financial Analysis Needs AI at the Edges. Never in the Math.</title>
      <dc:creator>Helkyn Coello</dc:creator>
      <pubDate>Sat, 05 Sep 2026 11:38:15 +0000</pubDate>
      <link>https://dev.to/unlocked-consulting/financial-analysis-needs-ai-at-the-edges-never-in-the-math-2olp</link>
      <guid>https://dev.to/unlocked-consulting/financial-analysis-needs-ai-at-the-edges-never-in-the-math-2olp</guid>
      <description>&lt;p&gt;A colleague recently asked me how he would run financial analysis with AI over the books of an accounting system: the ledgers, the trial balances, the P&amp;amp;Ls. He had three options on the table and no strong preference among them. Generate the PDF reports the system already produces and hand those to a model as input. Point the AI directly at the relational tables and let it query. Or skip the AI and compute the analysis with deterministic code.&lt;/p&gt;

&lt;p&gt;The question sounds like a tooling choice. It is an architecture decision, and it happens to be one I had already answered a couple of months earlier, in an independent prototype: a system that computes the 33 financial metrics I consider core, reads accounting PDFs from any source system, and confines AI to the two edges of the pipeline. It never touches the math.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Every metric is a formula&lt;/strong&gt;&lt;br&gt;
Financial analysis has a property that most AI conversations skip past: it is probably the most deterministic workload in the building. Every metric in my catalog is an arithmetic expression over named sums. The current ratio is current assets divided by current liabilities. Working capital is a subtraction. Margins are divisions. There is no judgment inside any of these, no interpretation, nothing that benefits from a model's flexibility. A formula that is allowed to be flexible has stopped being a formula.&lt;/p&gt;

&lt;p&gt;That settles the third of my colleague's options first: the computation itself is code, and it should not be anything else. I made the general version of this argument in June: if you need the same answer twice, dont't use AI. A trial balance is the purest case of that rule I know. Ask the system for the current ratio twice, and anything other than the identical number to the cent is not intelligence showing initiative. It is a defect.&lt;/p&gt;

&lt;p&gt;The other two options fail for more interesting reasons. Handing exported PDF reports to a model as the analysis input means taking numbers that were structured data a second ago, flattening them into a document, and asking a language model to re-derive them with no guarantee it reads every line the same way tomorrow. Pointing the model at the relational tables sounds more rigorous and carries the same flaw one layer down: the queries it writes and the aggregations it chooses can drift between runs, and in accounting, drift between runs has a technical name. It is called an error.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Two numbers that prove the rule&lt;/strong&gt;&lt;br&gt;
All 33 metrics compute the same way: by formula, with no exceptions and no heuristics in between; 14 of them carry thresholds that can raise findings. Most of those thresholds flag things worth investigating, not emergencies: a liquidity ratio drifting below its floor, receivables aging out too far. Two conditions sit in a different class, the only ones in the system marked as critical. And calling them the most important metrics would miss what they are. They are the alarms.&lt;/p&gt;

&lt;p&gt;The first is trust liability coverage: cash held in trust minus the deposit liabilities it is supposed to cover. If that number goes negative, client money is not where it must be. That is not an analysis finding; that is a legal problem. The second is the balance sheet check: assets minus liabilities minus equity, adjusted for the year's income and expenses. Accounting says that expression equals zero. If it does not, the emergency is of a different kind: the source data contradicts itself, which means the other 32 metrics were computed from numbers that do not agree with each other.&lt;/p&gt;

&lt;p&gt;Only two conditions in the system count as emergencies: client money not covering deposits, and books that do not balance. They are exactly the numbers no model should ever estimate.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The entrance: reading, then auditing the reading&lt;/strong&gt;&lt;br&gt;
So where does AI belong here? At the entrance, doing the one job deterministic code is genuinely bad at. Every accounting platform prints its own PDFs: different layouts, different column orders, different names for the same concepts. The classic answer is a rigid import template, which means every company adapts its exports to your tool before the tool does anything for them. The prototype I built inverts that. AI reads whatever the source system produced, extracts the figures, and writes them into the database. Interpreting messy, heterogeneous documents is precisely where non-determinism is a feature, because there is no formula for "understand this layout you have never seen."&lt;/p&gt;

&lt;p&gt;The payoff is practical: 25 of the 33 metrics compute from the trial balance alone. One file, whatever it looks like, and most of the analysis lights up.&lt;/p&gt;

&lt;p&gt;But the design does not even trust the reading. Before anything is ingested, a second, cheaper model call audits the first one: it compares the parsed table against a sample of the raw source and flags the classic misparses, a column shifted one position over, a subtotal ingested as if it were an account, a period detected wrong.&lt;/p&gt;

&lt;p&gt;A deterministic rule decides when that second reader runs. For PDFs it always runs, because PDFs are where parsing goes wrong most often. For spreadsheets it runs only when the code itself finds a reason to doubt the parse: too many source rows dropped, or a value column that came back empty. Each of those calls costs money, so the same rule that guards quality also controls the cost: the system pays for a second opinion only when the numbers justify one.&lt;/p&gt;

&lt;p&gt;Then the deterministic gates run. Before a single formula executes, the trial balance identity has to hold, receivables and payables have to tie to their aging reports, and periods have to be continuous, all within a tolerance of one dollar. A misread cell does not flow quietly into a ratio; it fails loudly, before the analysis exists. If the AI service is down, nothing stops: the file is still ingested, the metrics still run, and the system records that this particular check did not happen. The pipeline works without the models; what they add is checking, not dependency.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The exit: prose over a case file the code built&lt;/strong&gt;&lt;br&gt;
The other place AI earns its seat is at the exit, and this is where I expected to give the model the most freedom and ended up giving it the least. A current ratio below its threshold is just a finding; it is only useful once someone explains it. That explanation is not the model's to invent. Before the model writes a word, code gathers the evidence: up to six months of history for the metric that breached, the other metrics that breached in the same month, and, when a general ledger is loaded, the transactions behind the affected accounts. The model writes prose over that evidence and may cite only figures inside it; a validator then checks that every number in the sentences traces back to a value the code computed. The AI writes the sentences. Every figure inside them was calculated somewhere else.&lt;/p&gt;

&lt;p&gt;The same discipline decides the month's verdict. Whether a period is clean, has issues, has critical issues, or failed validation is assigned by rules, never by the model; the model narrates the state it is handed. The design notes for the system state it as a flat rule: AI reads, suggests, and narrates; code verifies and computes.&lt;/p&gt;

&lt;p&gt;My favorite consequence is what happens when nothing is wrong. Zero findings do not produce silence. They produce a verdict, computed by code, that the books came back clean, every accounting check passed, all 33 metrics inside their configured ranges, with the model writing the short summary a person reads. In the months when nothing is wrong, that verdict is the deliverable.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What the prototype refuses to do&lt;/strong&gt;&lt;br&gt;
When data is missing, the system says so instead of estimating. A metric whose source report has not been uploaded comes back empty, with the reason attached; growth rates simply do not exist until there are two months to compare. A model sitting in the middle of the pipeline would have filled those gaps fluently, and that is exactly the problem: here, the absence of an answer is information, and the system preserves it.&lt;/p&gt;

&lt;p&gt;Two planned capabilities, cash forecasting and fraud detection, stay switched off until they can be built with the same discipline. And I should be equally plain about status: this is a prototype. It has processed the documents I have fed it; it has not run anyone's production books. What I am confident in is the pattern, and the pattern is the point.&lt;/p&gt;

&lt;p&gt;In June I wrote that the shape is AI in the data, code in the process, and called it the architecture that actually scales. A summer of building later, I would sharpen the claim. The mature shape is not two layers but a guarded pipeline: AI reads at the entrance and a second model audits the reading, deterministic code computes everything in the middle, and AI explains at the exit from evidence the code assembled. AI at the edges, code at the core. It still scales. More to the point, it is an architecture you can answer for.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://unlockedconsulting.ai" rel="noopener noreferrer"&gt;https://unlockedconsulting.ai&lt;/a&gt;&lt;/p&gt;

</description>
      <category>financialanalysis</category>
      <category>llmarchitecture</category>
      <category>ai</category>
    </item>
    <item>
      <title>You cannot fire your AI agents</title>
      <dc:creator>Helkyn Coello</dc:creator>
      <pubDate>Sat, 29 Aug 2026 18:24:39 +0000</pubDate>
      <link>https://dev.to/unlocked-consulting/you-cannot-fire-your-ai-agents-2m8o</link>
      <guid>https://dev.to/unlocked-consulting/you-cannot-fire-your-ai-agents-2m8o</guid>
      <description>&lt;p&gt;A branch came in for review with about sixty commits on it, every one authored by someone on the team. He hadn't written them. Claude Desktop had, running on his laptop, signing commits with the git identity we configured during setup. As far as the repository was concerned, the work was his. As far as blame, audit and every code-ownership convention we had, the work was his. Nobody could separate the four or five decisions he had actually looked at and accepted from the fifty-odd changes the model produced while he clicked through the result to see whether it worked.&lt;/p&gt;

&lt;p&gt;We moved the whole thing off his machine: the model runs server-side now, the working copy is provisioned per ticket in an isolated environment, and what comes back is a URL. That solved the port conflicts and the dependency drift, which was why we did it. It did not solve the attribution problem. It relocated it. Now a service account commits, and the service account is one identity shared by every run, for every person, on every ticket.&lt;/p&gt;

&lt;p&gt;That is the shape of the thing arriving at enterprises considerably faster than most access-management programmes are ready for.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;a new hire, a printer, and an agent&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Take a new hire in their first week. They have a unique identifier that will never belong to anyone else, a set of permissions somebody requested by name, a login trail, and an offboarding procedure that takes an afternoon. Four things: who they are, what they can reach, what they did, and how you get rid of them. Hiring, permissions, audit, firing.&lt;/p&gt;

&lt;p&gt;Now the printer on the third floor. It has an asset tag, it sits on a network segment that lets it reach the print server and nothing else, it logs every job, and you can unplug it. Same four things. Nobody is impressed by the printer, but the printer is fully accounted for.&lt;/p&gt;

&lt;p&gt;Now the agent your team stood up last month to triage tickets, read the CRM and post summaries into Slack. Who it is: it uses a key minted from a human account, probably belonging to whoever built the thing. Every log in every downstream system records that person, not the agent. What it can reach: whatever they could, which for an engineer with admin rights is usually everything. What it did: indistinguishable from what the human did. How you fire it: you rotate the key, and you find out what else breaks.&lt;/p&gt;

&lt;p&gt;Nobody decided the agent could write to the CRM. It inherited the permission from the account it borrowed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;where the answers get uncomfortable&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Identity means a credential belonging to this agent and only this agent, distinguishable in every downstream system's logs from the human who deployed it and from the other eleven agents your company is running. A name in a config file does not qualify. Whether you can have this depends entirely on the platform. API-native services handle it well: Anthropic, OpenAI, Stripe and others will issue a key with its own permissions and its own name, no seat required. The business systems where most agents actually work are the problem. Salesforce, HubSpot and most CRMs model access around seats occupied by people, so a separate identity per agent means paying for a seat or sharing a credential, and sharing is what teams choose.&lt;/p&gt;

&lt;p&gt;Scope is what it can touch, expressed as operations rather than systems. "Read from the ticket store, write to one Slack channel, call the pricing endpoint" is a scope. "Has an API key for Salesforce" tells you nothing about what the thing can actually do. The trade-off is real and worth stating plainly: narrow scopes break more often, and every break lands on the person who understands the permission model, which is never the person operating the agent. Broad scopes are cheaper to run and quieter, right up until the day they aren't.&lt;/p&gt;

&lt;p&gt;Attribution needs three facts on every action rather than one: which agent did it, which run it belonged to, and which human or system event authorised that run. What we ended up doing was giving every run its own number, and making sure that number appeared everywhere the run left a mark: in the commit, in each API call, in the ticket it came from. So anything the agent touched could be traced back to one specific run and to whoever started it. It is unglamorous plumbing, and it is the difference between a twenty-minute investigation and a week of reading diffs, which is the recurring governance cost that shows up after the agent is already built.&lt;/p&gt;

&lt;p&gt;Revocation is where the situation exposes itself, because revocation and identity turn out to be the same problem seen from opposite ends. Try it. Pick an agent running in production and cut off its access without touching anything else. If the credential is shared you cannot, so you rotate the key instead, and the rotation takes down the nightly export, two internal dashboards and something somebody wired up in March, because they all use the same key and nobody wrote down that they did. The rotation gets scheduled, then deferred, and the agent you wanted to stop keeps running for another eleven days.&lt;/p&gt;

&lt;p&gt;That delay is the test. If cutting off one agent requires a change-management window, what you have is not a revocation path. It is an intention to revoke, subject to scheduling.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;the parts exist, the whole does not&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Every piece of this problem has been solved somewhere, and none of the solutions cover the ground you need. Cloud providers can issue an identity to a running process and rotate it automatically, which works cleanly as long as everything stays inside one cloud. Credentials that expire in hours rather than years are normal practice in infrastructure and almost unknown in the business tools where agents actually operate. And there is no single place to go and switch one agent off across Salesforce, your warehouse and a third-party API, because each of them keeps its own list. You will build this out of parts, and it will not be elegant.&lt;/p&gt;

&lt;p&gt;What you can do now costs engineering time once rather than exposure forever. Mint a distinct credential per agent per purpose, even when the platform makes you buy a seat to do it, because the licence fee is cheaper than the incident. Keep a register of which agent holds which credential, with an owner's name on each row, because the thing that makes revocation slow is not the API call, it is not knowing what breaks. Give every credential an expiry date when you create it, so that the ones nobody remembers stop working on their own rather than living forever. And stop letting agents inherit a human's permission set as a starting point, because a permission set assembled over four years of a career is the worst possible baseline for a process that only needs to read tickets. This is the same logging, testing and oversight floor most companies never built, arriving now with agents attached.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;verifiable agent identity is coming, and retrofitting is expensive&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;In June 2026 Estonia approved a framework giving AI agents a verifiable digital identity, with permissions that are scoped, auditable and revocable, tied to the eIDAS 2.0 ecosystem. It is an approved framework rather than a law in force, and that distinction matters. What requires no guessing is the direction: the first government to treat agent identity as identity infrastructure rather than a vendor feature has now done so.&lt;/p&gt;

&lt;p&gt;Plan around it. At some point an agent acting on a company's behalf will be expected to have a verifiable identity, and someone will have to be able to say which one acted and on whose authority. Retrofitting attribution onto a fleet of agents that all share one service account is going to be a genuinely expensive project.&lt;/p&gt;

&lt;p&gt;One thing worth doing, and it takes an hour. Take an agent already running in production and try to switch off its access alone, in a staging environment that mirrors how your credentials are set up. Time it, and write down what else stopped working.&lt;/p&gt;

&lt;p&gt;That is your answer to how you would fire that agent. Ask the same question in a meeting and you will get "we'd rotate the key," which is a description of an action rather than an outcome. The hour tells you what the action actually costs. That is the number worth knowing before you need it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;faq&lt;/strong&gt;&lt;br&gt;
&lt;em&gt;Why can't we just give each AI agent its own account in every system?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Sometimes you can, and where you can you should. API-native platforms will issue a key with its own permissions and no seat attached. The obstacle appears in business systems priced per user, which have no concept of a non-human account, so a separate identity for each agent means buying a seat. Buy the seat. It costs less than the incident.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;What does agent attribution actually require in the logs?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Three facts on every action: which agent performed it, which run it belonged to, and which human or system event started that run. The minimum that makes an investigation possible is giving each run its own number and carrying it through every call the agent makes.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Isn't narrow scoping more trouble than it's worth?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Narrow scopes break more often, and every break lands on the person who understands the permission model rather than the person operating the agent. Broad scopes are cheaper and quieter until the day they aren't. The trade-off is real; it is a decision to make deliberately rather than by default.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;How do we know if our revocation path works?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Test it. Pick an agent running in production, switch off its access alone in a staging environment that mirrors your credential setup, and count what else stopped working. If doing it for real would need a change-management window, you don't have a revocation path yet.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://unlockedconsulting.ai/" rel="noopener noreferrer"&gt;https://unlockedconsulting.ai/&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>vibecoding</category>
    </item>
    <item>
      <title>Adoption Is a Latency Problem: The Zero-Backlog Policy</title>
      <dc:creator>Helkyn Coello</dc:creator>
      <pubDate>Mon, 17 Aug 2026 07:15:32 +0000</pubDate>
      <link>https://dev.to/unlocked-consulting/adoption-is-a-latency-problem-the-zero-backlog-policy-1pi0</link>
      <guid>https://dev.to/unlocked-consulting/adoption-is-a-latency-problem-the-zero-backlog-policy-1pi0</guid>
      <description>&lt;p&gt;&lt;em&gt;Notes from an AI video project, and the release cadence that turned one team into the project's owners&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The most common artifact of an enterprise AI rollout is a dead feedback channel. It gets created in week one with genuine enthusiasm, collects a burst of suggestions in the first month, and goes quiet by the third. Not because people ran out of opinions, but because nothing they said ever came back as a release.&lt;/p&gt;

&lt;p&gt;The industry numbers say access is no longer the problem. Deloitte's latest survey puts sanctioned AI access at close to 60% of workers, and fewer than 60% of those use it in their daily work. The gap between having a tool and using it is where adoption lives, and I want to describe the project that taught me what actually closes that gap. It was not the training plan. It was the release cycle.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The project&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The tool generated training videos for our training team. We started with a deliberately narrow scope: simple functions, one team, one job. The build ran in stages of two to three weeks at most, and every stage had to end in something the team could actually use. Not a demo. Not a prototype behind glass. A working increment in their hands.&lt;/p&gt;

&lt;p&gt;The day the first increment shipped, we opened a dedicated channel with everyone on the team who used it. Feedback arrived daily: bugs, friction points, and, more valuable than either, feature ideas from the people doing the work. The AI team had one standing instruction: respond the same day. Every day we held a short meeting to go through what had come in, with the lead of the using team in the room, and whatever we decided got built the same day or the next. Releases went out daily, and they carried not only fixes but the users' own ideas.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Zero backlog as a promise&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;We called it a zero-backlog policy: nothing you tell us gets parked. A suggestion either ships within a day or two, or you hear the same day why it won't.&lt;/p&gt;

&lt;p&gt;Here is why I now treat this as the mechanism of adoption rather than a nice-to-have. Buy-in decays at the speed of your release cycle. A suggestion implemented within a day is visible proof that the user shapes the tool; the same suggestion sitting in a quarterly backlog is proof of the opposite. Both signals are received loud and clear. Only one of them produces owners.&lt;/p&gt;

&lt;p&gt;And ownership is what we got. The team stopped talking about "the AI tool" and started talking about the features they had asked for. They corrected each other's usage in the channel. When something broke, they reported it the way you report a problem in something that is yours. Without anyone assigning the role, they had become the project's evangelists inside the company.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Let the users do the presenting&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;We showed the project at every company-wide meeting, and made one deliberate choice: the people presenting were the users, not us. A training-team colleague explaining what the tool changed in their own week produces a categorically different effect than an AI team presenting slides about capabilities. One is a testimonial; the other is a pitch, and everyone in the room can tell them apart.&lt;/p&gt;

&lt;p&gt;Every demo carried the same metric, tracked from day one: finished videos per week. Not prompts sent, not logins, not a satisfaction score. The number the business already cared about before the project existed. When the metric moved, nobody had to be convinced that it mattered.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The taper&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;When the tool stabilized and adoption was no longer in question, we let the cadence relax: releases every other day, then weekly.&lt;/p&gt;

&lt;p&gt;This part matters more than it looks. Zero backlog is a launch regime, not a way of life. It is expensive, it consumes the AI team, and it is worth paying for during exactly one window: while the organization is still deciding whether this tool is theirs or something being done to them. Once that question is settled, you can taper. Taper before it is settled and the channel dies like all the others.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;One decision that looked unrelated&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Traceability and audit were built in from the first increment, not added later. This looked like a compliance preference at the time. In practice governance never showed up as a separate, later layer of friction: it was simply how the tool worked from day one. A governed path that shows up after adoption shows up as a slowdown, and any governed path slower than the ungoverned one gets bypassed. Built in from the start, nobody ever experienced it as a gate.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The common failures are latency failures in disguise&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A big-bang project that spends six months before anyone can touch it means feedback cannot even start until the window for winning people over has already closed: the suggestion board meets once a quarter, and by the time an idea ships, the person who proposed it has stopped expecting it. So ninety days of latency produces the same adoption as never shipping at all. The rollout whose entire change plan is a training session and a comms deck has no release cycle at all for feedback to flow into, so the loop never exists in the first place — which goes some way to explaining why "change management" has the reputation it has. ServiceNow's 2026 maturity index found that 74% of its top-scoring organizations run change management programs, against 7% of everyone else. I doubt the difference is that the 74% write better emails.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What I would compress it to&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Start with a scope small enough that a two-week stage produces something usable. Release fast enough that a suggestion and its implementation live inside the same week. Ship the users' ideas, visibly, so the tool becomes partly theirs. Put the users on stage, not the project team. Track one business metric from day one, the one the business already watched. Keep the audit trail on from the first build. And treat the daily cadence as an investment with a defined end: its job is to settle the ownership question — whose tool is this — while it is still open, because it does not stay open forever.&lt;/p&gt;

&lt;p&gt;Adoption is a latency problem. Every item on that list is a way of lowering the latency.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://unlockedconsulting.ai/" rel="noopener noreferrer"&gt;https://unlockedconsulting.ai/&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>aiadoption</category>
      <category>agents</category>
    </item>
    <item>
      <title>AI dragons: capable of everything, governed by nothing</title>
      <dc:creator>Helkyn Coello</dc:creator>
      <pubDate>Sun, 16 Aug 2026 15:10:09 +0000</pubDate>
      <link>https://dev.to/unlocked-consulting/ai-dragons-capable-of-everything-governed-by-nothing-2d9d</link>
      <guid>https://dev.to/unlocked-consulting/ai-dragons-capable-of-everything-governed-by-nothing-2d9d</guid>
      <description>&lt;p&gt;Last weekend the season finale of House of the Dragon ran for seventy-five minutes. I decided that was long enough to build an agentic chatbot. That was my goal. And I was determined to make it. So I watched the episode on one screen with a terminal open on the other, betting myself I could have it working before the credits.&lt;/p&gt;

&lt;p&gt;The bet held, roughly. By the time the credits ran I had a working chat interface, a dashboard, tool access, and a model that could take a request, work out what needed doing and go do it. Under any definition anyone would recognise, it was agentic. It was also completely ungoverned. No limits on what it could touch, no record of what it had done, no point at which a human had to agree before something happened.&lt;/p&gt;

&lt;p&gt;It was a dragon nobody had trained yet. Capable of everything, governed by nothing.&lt;/p&gt;

&lt;p&gt;It took me longer than it should have to notice what I was actually doing. I was raising dragons.&lt;/p&gt;

&lt;p&gt;Mine did not breathe fire. They called APIs, queried databases, picked tools, passed information between themselves, and every so often did something I had not thought to forbid. That last part is the entire job.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;the week that followed&lt;/strong&gt;&lt;br&gt;
Seven days later I am still working on it, and none of that week went into capability. Capability was finished before the episode was.&lt;/p&gt;

&lt;p&gt;The week went into the part nobody films. Logging, so that every action the system takes leaves a record that outlives the session. Attribution, so that a month from now it is possible to say which change came from a person and which came from the model, because a log that cannot answer that question is a count, not an audit trail. Explicit intervention points, where the thing stops and waits for a human rather than proceeding because proceeding was technically possible. Boundaries on what it may reach: which systems, which tables, which operations, and what happens to a request that falls outside them.&lt;/p&gt;

&lt;p&gt;I built those against the EU AI Act as a reference rather than an obligation. Most of what I run is nowhere near the high-risk tier, and the Act attaches its heavy requirements, the documentation, the retention periods, the designed-in human oversight, to systems that land in specific categories. Mine does not. But the Act is the most carefully argued description available of what a system needs in order to be answerable for itself, and using it as a design brief costs nothing when you are building anyway. Retrofitting it later costs a great deal.&lt;/p&gt;

&lt;p&gt;Seventy-five minutes to build. A week to make it safe to run, and it is not finished.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;the ratio is the whole point&lt;/strong&gt;&lt;br&gt;
That ratio is not a story about me being slow. It is the shape of the work now, and almost every conversation about AI adoption has it backwards.&lt;/p&gt;

&lt;p&gt;Getting an AI agent to do something impressive is a weekend. Getting it to do the impressive thing you actually asked for, and nothing else, is the work — and it is the part nobody puts in the demo. The demo is always the seventy-five minutes. The week never appears, because a week of writing rules and watching them fail does not screenshot well.&lt;/p&gt;

&lt;p&gt;Which is the oldest problem there is with anything powerful. Nobody who has ever trained an animal struggled with what it was capable of. Capability is what the animal came with. The training is everything else, and it does not end.&lt;/p&gt;

&lt;p&gt;The capability arrives intact, in a weekend, for anyone. What separates a system you can put in front of a customer from a system you can only put in front of a colleague is the thing that takes the week, and then keeps taking time, because every rule I wrote turned out to be either loose enough that the model did something I had not anticipated, or strict enough that it refused work a person would have approved. Each round teaches you the distance between the constraint you described and the constraint the model inferred, and that distance is always wider than it looks from where you are standing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;what this costs, and where it gets budgeted&lt;/strong&gt;&lt;br&gt;
Here is the practical consequence, and it is the reason this is worth more than an anecdote.&lt;/p&gt;

&lt;p&gt;When a team asks for budget to build with agents, the number they produce is almost always the seventy-five minutes. The platform, the seats, the licence, the sprint. That number is real and it is also the smallest part of the total, and because the taming is invisible in every demo the team ever watched, nobody thinks to price it.&lt;/p&gt;

&lt;p&gt;Worse, it gets priced as a one-off when it behaves like a subscription. My guardrails are not finished, and they will not be finished when the system goes live either. A model version changes and the constraint that used to hold stops holding. A user tries something nobody imagined and finds a gap. The rules are a relationship you maintain, not a state you reach, which means the taming is operating expenditure wearing the costume of a project.&lt;/p&gt;

&lt;p&gt;So the honest budget line for anything agentic has two entries, and the second one recurs. If your plan only has the first, you have not planned to run a system. You have planned to hatch one.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;before you send it to battle&lt;/strong&gt;&lt;br&gt;
My agentic dragons are still in training. They are close, and close is not the same as ready. Close is where the fire lands somewhere you did not choose.&lt;/p&gt;

&lt;p&gt;AI is a wild and powerful creature by nature, dangerous by default, and useful only to whoever learns to tame it. So if you are shipping anything this year, learn to tame what your dragons can do before you send them to battle. Which in our world is called production.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://unlockedconsulting.ai" rel="noopener noreferrer"&gt;https://unlockedconsulting.ai&lt;/a&gt;&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>llm</category>
      <category>programming</category>
    </item>
    <item>
      <title>Three studies measured agentic AI adoption. They got 23%, 59% and 20%.</title>
      <dc:creator>Helkyn Coello</dc:creator>
      <pubDate>Wed, 05 Aug 2026 18:41:01 +0000</pubDate>
      <link>https://dev.to/unlocked-consulting/three-studies-measured-agentic-ai-adoption-they-got-23-59-and-20-3iin</link>
      <guid>https://dev.to/unlocked-consulting/three-studies-measured-agentic-ai-adoption-they-got-23-59-and-20-3iin</guid>
      <description>&lt;p&gt;A colleague forwarded me an internal slide with one number on it. Fifty-nine percent of companies are using agentic AI. And the bullet point underneath it said we were behind. What followed it was a request for budget for an agentic platform license, to be signed before the quarter closed. Nothing on the slide said which process it would touch.&lt;/p&gt;

&lt;p&gt;I went and found the study. Then I found two more from the same year. One said twenty-three percent are using agentic AI. The other said twenty percent. Fifty-nine against twenty is nearly a factor of three, and nothing dramatic had happened to the market between the fieldwork dates; they were months apart, not years. Each of them had counted a different thing.&lt;/p&gt;

&lt;p&gt;Here is the part that decided it for me. We had spent the better part of a year building a controlled environment where non-developers describe a feature and get back a live URL with that feature implemented. Under the hood a coordinator agent orchestrates specialized sub-agents, one for schema changes, one for the API layer, one for the UI, inside a constraint surface that defines what each is allowed to touch. By the definition behind that fifty-nine percent, we had been inside the count for months. By the definition behind the twenty percent, we were nowhere near it, because nothing in that pipeline runs without a human approving an implementation plan and then a step-by-step plan. One system, two definitions, two opposite answers about whether we had adopted agentic AI.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What the three agentic AI adoption numbers actually counted&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Paraphrasing the methodologies, because the wording matters less than the shape of the question: &lt;/p&gt;

&lt;p&gt;The fifty-nine percent came from asking respondents whether they were using AI agents. That includes a marketing lead who has a chat assistant wired to a calendar and a CRM lookup. Tool-calling counts. Anyone who has connected an MCP server to Claude and had it read a Jira ticket answers yes to that question, honestly.&lt;/p&gt;

&lt;p&gt;The twenty-three percent counted organizations with at least one agentic use case in production. Production is a stricter word than usage, but it is still soft: a pilot serving one team behind a feature flag passes, and so does an internal tool that three people use on Fridays. Anyone who has watched a proof-of-concept stall on its way to production knows how much room sits inside that word.&lt;/p&gt;

&lt;p&gt;The twenty percent counted deployed autonomous systems operating at scale against measurable business metrics. That is a different universe of claim. It requires someone to have defined a metric, instrumented it, and let the thing run wide enough that the number means something.&lt;/p&gt;

&lt;p&gt;None of the three firms did anything wrong. All of them wrote down what they counted, in a methodology section that most readers skip because the headline arrived first through a newsletter that stripped it out. The damage happens at the point of reading, and then again at the point where the figure gets pasted into a business case and acquires the authority of a measurement.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Autonomy has no measurable boundary&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The definitions diverge so far because there is nothing in the word itself to hold them in place.&lt;/p&gt;

&lt;p&gt;Autonomy is a continuous variable and vendors sell it as a binary.&lt;br&gt;
At the low end sits an LLM that can call one function and return an answer. A bit further along you get a system that decides which of many tools to call and in what order, which is already a different engineering problem. Then there is a loop with write access to a production database that runs on a schedule, retries its own failures, and only surfaces to a human when it gives up. Those three get sold under the same label and share almost nothing else: they break in different ways, they need different levels of auditing, they cost different amounts to run, and the person who has to be awake when one of them fails sits at a different level of the organization.&lt;/p&gt;

&lt;p&gt;That last part is what a survey percentage cannot carry. The expensive part of our environment was never running the AI agents. Claude was capable on day one. What took months was the constraint surface: forbidden patterns, allowed patterns, naming conventions, security boundaries, things that must always happen like input validation, things that must never happen like a schema change without a migration file. Then the tuning. Test, error, adjust, test again. The first version of any rule is either so permissive the model does something you never imagined, or so broad it refuses reasonable work. Every cycle teaches you how far your description of a constraint sits from the model's reading of it, and that gap is always wider than you estimate going in.&lt;/p&gt;

&lt;p&gt;So when a board deck says fifty-nine percent of the market has adopted agentic AI and we have not, the honest translation is that fifty-nine percent of respondents have something with tool access, and nothing in that number reflects the eleven months we spent making ours safe to run. Budget approved against that framing buys a platform license and no constraint surface, which is the same mistake as buying the model and calling it a system. &lt;/p&gt;

&lt;p&gt;I have seen the equivalent decision made from the other direction, when headquarters specified two named tools in a requirements document because both had been in the news, and the actual requirement was a set of deterministic multi-step tasks that needed no orchestration layer at all. A statistic in a headline and a tool name in a requirements document work the same way: both arrive looking like an answer, so nobody goes back to check what the question was.&lt;/p&gt;

&lt;p&gt;The defense is one question. Before any adoption figure goes anywhere near a business case, ask what it counted. If the answer is not available in two minutes of looking, the figure is not evidence. If it is available, write the number out as a full sentence with its definition included, and read it back. Fifty-nine percent of surveyed respondents report using at least one AI agent, where agent includes any LLM with tool access. Nobody builds a budget case out of that sentence.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The number worth having is your own&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The most useful comparison in all three studies sits inside one of them, between two of its own figures: fifty-nine percent using agents against nine percent with workflows that run without a person in the loop. Fifty points between having agents and having them operate unattended. That difference is the work — the constraint surface, the logging, the rollback paths, the tuning cycles. Fifty-nine percent have agents. Nine percent have done the work that lets them run alone.&lt;/p&gt;

&lt;p&gt;Which puts most companies exactly average: a chat assistant that can call an API, and a set of processes that still require a human to click approve. Nothing wrong with being there. The problem is not knowing where you are, because that is what decides whether the next twelve months of spending should go toward more agents or toward the rails that would let the existing ones run alone.&lt;/p&gt;

&lt;p&gt;So count your own instead. Take last month. How many workflows ran end to end without a person approving a step in the middle? Not how many use AI somewhere — how many completed, unattended, with a result someone acted on. In most companies I have looked at, the answer is zero or one, and it is the only agentic adoption statistic that has any bearing on what to do next. Counting it takes an afternoon. Doing it properly across every process you run, and writing down what each one would need to lose its approval step, is a workflow audit. Either way you end up with a figure you can defend, which is more than the slide could manage.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Questions this raises&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Why do agentic AI adoption studies disagree so much?&lt;/em&gt;&lt;br&gt;
Because each one sets the bar in a different place. The loosest asks whether you use an AI agent at all — a chat assistant that can look something up in your CRM counts. The middle one asks whether at least one agent is in production. A pilot used by a single team counts. The strictest asks whether autonomous systems are running at scale against a metric someone is tracking. The same company can honestly answer yes to the first and no to the third, and every study says which bar it used, in the methodology section, which the headline leaves behind.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;What actually counts as an agent?&lt;/em&gt;&lt;br&gt;
There is no line anyone can measure. An LLM that calls a single function, a system that picks which tools to use and in what order, and a scheduled loop with write access to production all get sold under the same label. They break in different ways, need different levels of auditing, and cost different amounts to run. A survey counts one of the three, and that choice sets the headline.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;How do I measure agentic AI adoption in my own company?&lt;/em&gt;&lt;br&gt;
Take last month and count the workflows that ran end to end without a person approving a step in the middle, and where someone acted on the result. Not the ones that use AI somewhere. That single number tells you whether to spend on more agents or on the rails the existing ones need.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://unlockedconsulting.ai" rel="noopener noreferrer"&gt;https://unlockedconsulting.ai&lt;/a&gt;&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Vibe Coding Won't Kill Developers. It'll Kill the Middle.</title>
      <dc:creator>Helkyn Coello</dc:creator>
      <pubDate>Sun, 26 Jul 2026 21:15:33 +0000</pubDate>
      <link>https://dev.to/unlocked-consulting/vibe-coding-wont-kill-developers-itll-kill-the-middle-2c61</link>
      <guid>https://dev.to/unlocked-consulting/vibe-coding-wont-kill-developers-itll-kill-the-middle-2c61</guid>
      <description>&lt;p&gt;When good cameras got cheap, everyone predicted the death of professional photography. The prediction landed wrong. The low end died outright: stock libraries, cheap portraits, mass-event coverage went to anyone with a phone and a free editing app. The high end did better than ever — editorial work, photojournalism with access nobody else had, an aesthetic you could not reproduce by buying the same gear. The damage landed in the middle. Small weddings, corporate headshots, real estate listings, the steady unglamorous bulk of the market: not extinction, compression. Prices fell, volume moved to cheaper substitutes, and the survivors climbed up or specialized out.&lt;/p&gt;

&lt;p&gt;That compression is the cleanest map I know for what AI-assisted coding is doing to software work. And this half I know from inside: two decades leading dev teams, and now building AI tooling for them.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The comfortable half of the argument&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The reassuring version of this is everywhere right now: you were never paid to type, you were paid to think, so AI just frees you to do the valuable part. It's not wrong. It's just the half that's easy to hear. The other half is about the market, not about you.&lt;/p&gt;

&lt;p&gt;Judgment, architecture, knowing what breaks in maintenance, deciding what not to build — a model that writes plausible code on command doesn't commoditize any of that. I have watched weeks of confusion land on people who could not read what a capable model generated; the gap was never the tool, and better AI autocomplete does not close that gap.&lt;/p&gt;

&lt;p&gt;But "judgment beats typing" answers only a question about skill and dodges the question about market structure. AI doesn't replace developers as a class; it commoditizes a segment. The segment it hits first is the same one the camera hit: the middle. The junior-to-mid tier that lived on CRUD apps, simple integrations, brochure sites, the standard internal tool with a form and a table behind it. That work was always implementation against a known spec, and implementation against a known spec is exactly what a model does cheaply now.&lt;/p&gt;

&lt;p&gt;I felt this directly on a side project. I spent many months of evenings building a content pipeline on a visual no-code platform; when AI-assisted coding changed the math, I rebuilt it as proper code in a few weeks. The months were not wasted — they are why I knew exactly what to build. But the same work now costs weeks, not months.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"So the middle just learns to think better?"&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This is the obvious objection, and it's worth taking seriously because the answer is where the whole thing turns.&lt;/p&gt;

&lt;p&gt;First problem: "learn to think better and survive" doesn't refute the thesis, it confirms it. The claim was never about the people, it's about the work. When someone in the middle develops real architectural judgment, they haven't saved the middle, they've left it. They climbed to the high end. The middle still empties out, by promotion or by exit. A wedding photographer who became a sought-after editorial shooter didn't prove the wedding market survived.&lt;/p&gt;

&lt;p&gt;Second problem: "thinking better" isn't a soft upgrade you bolt onto the same job. Going from implementing a spec to deciding what to build, judging tradeoffs, holding the whole system in your head, anticipating the failure that surfaces in production eight months later — that's a change in the muscle being used. Some people build it. Some won't, and some genuinely don't want to; they liked implementation and were good at it. The skill axis is real, but it doesn't pull everyone up by default.&lt;/p&gt;

&lt;p&gt;Third problem, and this is the one the reassuring posts never reach: capacity. Suppose everyone in the middle became a strong systems thinker overnight. The market still doesn't need that many architects, security specialists, performance engineers, and trusted advisors. The high end is narrow by definition, which is what makes it the high end. Premium editorial photography never had room to absorb every wedding shooter it displaced.&lt;/p&gt;

&lt;p&gt;So when people ask whether "think better" just means "become a consultant," the honest answer is partly. The trusted technical advisor is one exit from the middle, and a good one, but not the only one. Architecture is an exit. Deep security and performance work is an exit. Complex-domain expertise — the kind where you understand the business so well the AI becomes a lever instead of a crutch — is an exit. When a prototype that should take weeks comes together in an afternoon, the speed never comes from the model. It comes from twenty years of knowing the exact problem before the first prompt. The tool doesn't create that position; it amplifies what was already there.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The honest limit of the analogy&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I'd be selling you something if I claimed software follows photography like a law. It doesn't. Photography had roughly fixed demand; the number of weddings didn't triple because cameras got cheap. Software might. Total demand for software has grown for forty years, and cheaper production could grow it further, which means some of the displaced middle gets reabsorbed into work that didn't exist before. The wedding market never had that escape valve.&lt;/p&gt;

&lt;p&gt;So I won't tell you eighty percent of developers will be unemployed. I don't believe it, and the people who say it are guessing. The defensible claim is narrower and still uncomfortable: middle-tier work commoditizes, value migrates toward the extremes, and you should plan around that rather than hope the middle holds.&lt;/p&gt;

&lt;p&gt;For the individual, the move is to climb toward judgment, architecture, and advisory work, not to defend the commoditizing tier by getting slightly faster at it. Speed inside a shrinking band is a losing race against a tool that's cheaper than you and getting better weekly.&lt;/p&gt;

&lt;p&gt;For anyone running a technical team or a practice, this is the argument for senior-led work getting stronger. The middle is precisely the layer being liquified. Value concentrates in the people who were never doing middle work to begin with: the ones whose contribution was the system, the tradeoff, the domain, the call about what not to build. That's why this practice is senior-led and founder-run. It is not a positioning choice; it is what the map says to do.&lt;/p&gt;

&lt;p&gt;The question to sit with isn't whether you can out-think the AI. You probably can, on a good day. The question is which segment that thinking lives in, and whether that segment will still be there to stand in. Go look at the last five things you shipped. Count how many were implementation against a spec someone else wrote. That number is your exposure.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://unlockedconsulting.ai/" rel="noopener noreferrer"&gt;https://unlockedconsulting.ai/&lt;/a&gt;&lt;/p&gt;

</description>
      <category>vibecoding</category>
      <category>ai</category>
      <category>programming</category>
      <category>developers</category>
    </item>
    <item>
      <title>80% of Companies Are Already Out of Step With the AI Act and Don't Know It</title>
      <dc:creator>Helkyn Coello</dc:creator>
      <pubDate>Mon, 20 Jul 2026 17:59:18 +0000</pubDate>
      <link>https://dev.to/unlocked-consulting/80-of-companies-are-already-out-of-step-with-the-ai-act-and-dont-know-it-1m2o</link>
      <guid>https://dev.to/unlocked-consulting/80-of-companies-are-already-out-of-step-with-the-ai-act-and-dont-know-it-1m2o</guid>
      <description>&lt;p&gt;The tool worked. That was never the problem. A business manager had specced it with AI, straight from the business need, and the result did what he asked: it ran, it produced output, it looked finished. What it did not have was a single log line. No trace of why it produced a given answer. Nothing that could explain, later, a bad call made on real employee data. Nobody had asked for any of that, so the AI built none of it.&lt;/p&gt;

&lt;p&gt;That absence is the gap nobody priced in. Scaled across a company and pointed at the wrong category of system, it puts you out of step with the EU AI Act without a single person acting in bad faith.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The AI Act assumes an engineering floor you may not have built&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Here is the collision worth sitting with. The AI Act (Regulation (EU) 2024/1689) does not ask high-risk system operators to start testing their AI. It assumes they already do. Read the high-risk obligations and you will notice they are written on top of an engineering discipline that is taken for granted. Article 12 and Article 19 require automatic record-keeping and logging, with logs retained at least six months. Annex IV and Article 11 define what the technical documentation must contain; Article 18 requires keeping it at the disposal of authorities for ten years after the system is placed on the market. Article 14 requires human oversight designed in, with real intervention and stop capability, and it clearly calls out automation bias — the tendency to over-trust a machine's output — which tells you the drafters understood exactly how people behave around confident machines. Article 15 requires testing for accuracy, robustness, and resistance to adversarial manipulation.&lt;/p&gt;

&lt;p&gt;Every one of those clauses assumes a working baseline: that you log what your systems do, that you test them against errors, that a human can see inside the decision and pull the lever when it goes wrong. The Act regulates on top of that floor. It does not build the floor for you.&lt;/p&gt;

&lt;p&gt;Now hold that next to the data. The ServiceNow and ThoughtLab Enterprise AI Maturity Index for 2026 surveyed 4,500 executives and 2,000 employees across 19 countries and 12 industries. The number that matters: only 20% of organizations have implemented AI testing, auditing, and risk-assessment processes. Read that the way an engineer reads a red alert on a dashboard. Four out of five organizations have not built the discipline the high-risk regime stands on.&lt;/p&gt;

&lt;p&gt;So the inference — and I want to be precise that it is an inference, not a directly measured compliance figure — is that roughly 80% of companies are structurally out of step with what the Act assumes, for any system that falls in the high-risk tier. The obligations exist on paper; the engineering they depend on does not.&lt;/p&gt;

&lt;p&gt;Almost everything written about the AI Act treats it as a legal question: what does "high-risk" mean, what exactly does Article 14 require. Read it that way and the hard part is understanding the text — as if the logging, testing, and oversight would somehow build themselves once the reading is done. They do not build themselves. The Act is an engineering problem wearing legal language, and the harder question is whether the infrastructure it takes for granted — the logging, the testing, the oversight — exists in your organization at all. For most companies the honest answer is no, and they have never checked.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Scope this before you panic&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The alarm only applies to a specific subset of systems, and saying so is the difference between analysis and fear-mongering. The Act tiers systems by risk. Most internal tools — the meeting summarizer, the draft-email assistant, the thing that reformats your reports — sit in the minimal-risk tier and carry none of these obligations. You can run those with a clear conscience and a thin paper trail.&lt;/p&gt;

&lt;p&gt;The obligations attach to the high-risk tier, defined largely by Annex III: systems used in employment and worker management, access to essential services, credit scoring, critical infrastructure, biometric categorization, and a handful of others. A tool that screens job candidates, ranks employees, or gates someone's access to a service is in a different legal universe than one that summarizes a PDF.&lt;/p&gt;

&lt;p&gt;So the first move is not legal. It is classification. You cannot assess your exposure until you know which of your systems are in scope, and in my experience companies have done this for zero of them. They know the AI Act exists. But they have not mapped their own deployments against it. The business manager with the AI-specced tool could not have told you which tier it fell in, because the question never came up, and depending on what that tool actually decides, the answer changes everything about what he owed under the law.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The reason capability outran the floor&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The same maturity report explains how we got here. 59% of organizations are past piloting agentic AI, but only 9% have working autonomous multistep workflows. Capability is being deployed far faster than the governance layer beneath it. Only 16% have replaced fragmented legacy systems with an integrated platform; 41% still cite siloed data as a blocker. You cannot log and audit a decision cleanly when the data feeding it lives in six disconnected systems and the workflow stitching them together is held up with visual no-code duct tape.&lt;/p&gt;

&lt;p&gt;I have lived the small version of this, spending months building AI automation pipelines in visual tools. They produced output. But when a node failed silently, I had no clean trace of what happened or why, and debugging meant clicking through visual steps trying to find which one died. I eventually rewrote the core of it as proper code, and what changed was not speed. I could finally see inside it. Now imagine that same opacity at a corporate level sitting under a system that decides who gets a loan or who keeps a job. That is the engineering gap the Act assumes you closed before it ever showed up.&lt;/p&gt;

&lt;p&gt;There is a cultural tell in the data too. 57% of employees think leadership is not keeping up with fast-changing trends. That is what the governance gap feels like from below: the people closest to the work can sense that the controls are not there, even when the official version says they are.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What this actually requires of you&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Stop treating this as a question of whether you have read the regulation correctly. The exposure for anything you run in the high-risk tier comes down to one thing: whether the testing, logging, and oversight discipline — the one the entire regime takes for granted — actually exists under your systems. If it does not, you are not one clause away from compliance. You are one engineering layer away, and that layer takes months to build and tune, not an afternoon with a lawyer.&lt;/p&gt;

&lt;p&gt;The work is two concrete steps, in order. First, classify your AI systems by tier so you know which ones carry high-risk obligations and which ones you can leave alone. Most of your tools will fall out of scope, and that is the point. You want to spend your effort where the law actually bites. Second, for the systems that remain, check whether the assumed baseline is there: automatic logging you retain, documentation you can produce, testing against accuracy and adversarial failure, and a human who can both see the decision and stop it.&lt;/p&gt;

&lt;p&gt;When that human can pull the stop lever in time and understands what they are overriding, you have oversight. When the data feeding the decision lives across four disconnected systems with no trace of what changed, you have a documentation requirement you cannot satisfy and an audit you would fail. A structured workflow audit is one way to surface exactly where those traces break down.&lt;/p&gt;

&lt;p&gt;Most companies have done neither step. They have not classified their systems, and they have not audited the foundation under the ones that matter. Start with the first. Pull a list of every AI-assisted system that touches employment, access, credit, or essential services, and ask one question of each: if a regulator asked us to show six months of logs and explain a single decision, could we. The systems where the answer is no are your real exposure — and that one hour of asking will tell you more than another month of reading the Act.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://unlockedconsulting.ai" rel="noopener noreferrer"&gt;https://unlockedconsulting.ai&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>aiact</category>
      <category>agents</category>
    </item>
    <item>
      <title>Is Your Knowledge Base Actually Thinking, or Just Retrieving?</title>
      <dc:creator>Helkyn Coello</dc:creator>
      <pubDate>Mon, 20 Jul 2026 17:57:15 +0000</pubDate>
      <link>https://dev.to/unlocked-consulting/is-your-knowledge-base-actually-thinking-or-just-retrieving-cgo</link>
      <guid>https://dev.to/unlocked-consulting/is-your-knowledge-base-actually-thinking-or-just-retrieving-cgo</guid>
      <description>&lt;p&gt;A while back I was able to build a working knowledge base in less than an hour. Fifty-five minutes to be accurate. I started by opening VS Code with Claude Code in the side panel. I had documentation scattered across roughly ten training manuals, a video archive nobody watched, support ticket histories, enhancement docs, and meeting transcripts. The knowledge existed. It was just spread across systems with nothing connecting them.&lt;/p&gt;

&lt;p&gt;I described the islands to Claude the way I'd describe them to a new engineer. Then I asked for a chat interface backed by a vector database with RAG-style retrieval from plain-language questions. At around minute 45, it was running and returning accurate answers from the actual docs. I know it was not yet production software, but it was a real prototype on real data.&lt;/p&gt;

&lt;p&gt;But I had slipped one extra requirement into that prompt, almost without thinking about it: client-specific information should be flagged and surfaced only when a query came from that client or was about that client. We had training recordings and support history tied to individual accounts. That requirement is the whole article. The moment you say "surface this when the query is about that client," you have stepped outside what retrieval by similarity can do. Similarity answers one question: which stored text resembles the wording of the query. But a support ticket and a training recording tied to the same client may share no vocabulary with each other, or with the question being asked. What links them is not wording; it is the client they both belong to. That link is a relationship, and a relationship is something a similarity search has no way to represent.&lt;/p&gt;

&lt;p&gt;My prototype got away with it in a demo. In production, where real users ask cross-source questions every day, that shortcut stops holding.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What embeddings actually do, and where they stop&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Embed-and-retrieve works like this: you chunk the documents, run each chunk through an embedding model, store the vectors, and at query time you embed the question and pull back the nearest neighbors by cosine similarity, a measure of how close two texts sit in meaning. It works. For a large class of questions — "what does our refund policy say," "how do I configure SSO" — it works well, and you can have it running in an afternoon. &lt;/p&gt;

&lt;p&gt;The ceiling shows up the moment an answer lives across sources that don't look alike.&lt;/p&gt;

&lt;p&gt;Take a question we actually get internally: "Why is this client's month-end close behaving differently from the others?" The answer is not in one chunk. It's in a custom configuration noted in a support ticket from two years ago, a behavior described in a training recording for that specific account, and a feature flag mentioned in an enhancement doc. None of those three documents are textually similar to each other, and none of them are similar to the question. An embedding search returns whatever happens to share vocabulary with "month-end close." That points straight to the generic manual, which is the one document in the set that does not contain the answer.&lt;/p&gt;

&lt;p&gt;Similarity is a proxy for relevance. It's a decent proxy when relevant things tend to use the same words. It collapses when the relevant things are connected by something other than vocabulary: a shared client, a shared subsystem, a cause-and-effect chain that nobody wrote down in a single place.&lt;/p&gt;

&lt;p&gt;This is the part that gets skipped in the market. A team sets up RAG, declares "the knowledge base is done," and stops exactly there. You end up with a search engine that talks. Useful, but it answers "what does the document say," and the questions that actually justify the project are the ones that ask "what does the situation mean."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The relationships are the knowledge&lt;/strong&gt;&lt;br&gt;
I was reminded of all this at Milano AI Week, where someone framed it cleanly: the real intelligence in a knowledge base isn't in the chunks; it's in how the segments relate, the graph rather than the index.&lt;/p&gt;

&lt;p&gt;Concretely, that means a second representation alongside your vectors. You run entity extraction over the same corpus — clients, features, configurations, people, error conditions, dates — and you model the edges between them. Ticket 4471 references client Aurora. Aurora runs the custom close configuration. That configuration was introduced by enhancement request 882. The training recording from March covers that workflow. Those are explicit edges in a graph, and they are traversable regardless of whether any two of those nodes share a single word.&lt;/p&gt;

&lt;p&gt;Now the same question (why is this client's month-end close behaving differently) gets answered along a different path. The system maps "this client" to the Aurora entity, walks the edges out to its configurations, tickets, and recordings, and assembles a candidate set that no similarity search would ever produce, because the documents have almost nothing lexically in common. The graph answers what is actually related; the embeddings answer what looks similar. You want both: similarity finds the entry points and the unstructured passages, and the graph traversal gathers everything connected. Either one alone leaves capability on the table.&lt;/p&gt;

&lt;p&gt;I'm not arguing embeddings are wrong. I'm arguing that "chunks plus embeddings" is a floor people keep mistaking for a ceiling.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The cost is real, and it is not catastrophic&lt;/strong&gt;&lt;br&gt;
Here's the honest trade-off, because the graph approach is not free.&lt;/p&gt;

&lt;p&gt;Entity extraction has to be tuned. The first pass over-extracts, treats every capitalized string as an entity, and cannot tell "Aurora the client" apart from "Aurora the internal name of a feature." Entity resolution — deciding that three different spellings refer to the same client — is where most of the real effort goes. You need a schema for what entities and relationships matter to your business, and that schema is opinionated, which means it requires someone who understands the domain to define it. You'll iterate it several times before it behaves. This is the same tuning loop I've hit building constraint surfaces for AI tools internally: the gap between how you described the rule and how the system interpreted it is always wider than you expect at the start, and you only close it by running real queries and watching what breaks.&lt;/p&gt;

&lt;p&gt;Operationally you're now maintaining two stores and a pipeline that keeps the graph in sync as documents change. That's more moving parts than a single vector index. Budget for it.&lt;/p&gt;

&lt;p&gt;But the cost difference between the two architectures is incremental, not an order of magnitude. The embed step is shared. The extra work is the extraction pipeline, a graph store, and the retrieval logic that combines the two. The capability difference, by contrast, is enormous. A chunk-only system can answer questions confined to a single document. The graph version can answer questions whose answer requires reasoning across sources that were never connected by an author. For an enterprise knowledge base, that second class of question is usually the entire reason the project was funded.&lt;/p&gt;

&lt;p&gt;If your knowledge base is meant to serve support engineers, account managers, or anyone diagnosing a situation that spans history, the chunk-only version will demo beautifully and disappoint in week three, when the first cross-source question comes in and it confidently returns the generic manual. This is exactly the kind of representation decision we work through when we design a knowledge base meant to survive contact with real questions.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What to check before you sign off on an architecture&lt;/strong&gt;&lt;br&gt;
There's a simple test. Take five questions your team actually asks, the hard ones, the ones that make someone sigh, and for each one, ask yourself whether the full answer lives in a single document or has to be stitched together from several. If most of them need the stitching, similarity retrieval will not get you there, and no amount of better embeddings or bigger context windows fixes a representation problem.&lt;/p&gt;

&lt;p&gt;That test is also the dividing line between a search engine and a knowledge base. A search engine retrieves text. A knowledge base models the entities your business cares about and the relationships between them, then reasons over that structure. The prototype I built in fifty-five minutes was the former with one relationship bolted on by hand. The production version of that same system is the latter, and the difference is entirely in how seriously you model the graph.&lt;/p&gt;

&lt;p&gt;Before you approve any internal knowledge base design, ask the team building it one question: when the answer lives in three unrelated documents, what mechanism finds the other two? If the only answer is "similarity," that tells you two things: the team hasn't seen the ceiling yet, and the budget you're about to approve will buy a system that stops exactly where your hardest questions begin. Decide whether that's enough before the cost moves downstream into the maintenance no one can see.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://unlockedconsulting.ai" rel="noopener noreferrer"&gt;https://unlockedconsulting.ai&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>claude</category>
      <category>productivity</category>
      <category>rag</category>
    </item>
    <item>
      <title>The On-Premise LLM Lottery</title>
      <dc:creator>Helkyn Coello</dc:creator>
      <pubDate>Mon, 20 Jul 2026 17:53:51 +0000</pubDate>
      <link>https://dev.to/unlocked-consulting/the-on-premise-llm-lottery-54di</link>
      <guid>https://dev.to/unlocked-consulting/the-on-premise-llm-lottery-54di</guid>
      <description>&lt;p&gt;I lost a full long weekend to OpenClaw once. Friday to Sunday, trying to get it running locally with an on-premise LLM on my personal machine. The machine was reasonably capable, though nothing close to a datacenter. I had watched something like forty hours of tutorials first, so I went in confident. Every video made it look like four steps: download, install, connect a model, start automating.&lt;/p&gt;

&lt;p&gt;By Sunday evening nothing worked. The wall I hit had nothing to do with my configuration. It was a real limitation in the tool itself, on the hardware I had, with no path around it from the user side. The clean take in the tutorial had been produced in a controlled setup, and the hours of debugging that preceded it never made it into the video.&lt;/p&gt;

&lt;p&gt;That weekend taught me something I keep coming back to when someone tells me they want to run their models on-premise. I hadn't picked the wrong app. I'd treated "run an LLM locally" as a task with a fixed answer, when it is actually a search.&lt;/p&gt;

&lt;p&gt;There is a whole genre of content that presents local inference as a tidy checklist. Download the app, check your RAM, pick a model size that fits, done. And for the audience those guides are written for — one person, one laptop, one chat window — that genuinely is the whole job. They are answering a smaller question than the one an enterprise is asking, and they answer it well. The problem starts when someone in a regulated company reads that checklist and assumes their version is the same problem with more zeros in the hardware budget.&lt;/p&gt;

&lt;p&gt;It isn't. In the enterprise version, your hardware is already fixed by procurement decisions made before you arrived. Your runtime is constrained by what your ops team will agree to support. Regulatory and compliance requirements narrow which models you can even consider before you benchmark a single one. The checklist collapses under those constraints, and what's left underneath is a matchmaking problem with several moving axes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The axes nobody puts in the checklist&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;After that weekend I stopped looking for the right app and started running an actual search. I tested roughly twenty models before one fit. That number wasn't diligence for its own sake. There are many possible combinations of model, quantization, and runtime, and most of them simply don't work on any given hardware, so finding one that does takes that many attempts.&lt;/p&gt;

&lt;p&gt;Those combinations vary along a few axes, and the axes interact with each other. Model size is the obvious one, but the parameter count on the label tells you almost nothing until you pair it with quantization. A model that won't load at full precision runs comfortably at a lower quantization, with a quality cost you have to measure rather than assume. Then the runtime layer sits on top of that: I settled on vLLM and Ollama depending on the case, because the runtime is what turns "the weights exist on disk" into "the thing answers a request at an acceptable latency." And all of that has to land on hardware you don't get to choose.&lt;/p&gt;

&lt;p&gt;The published guides almost always cover exactly one point in this space and present it as universal. A tutorial that shows a specific model at a specific quantization on a specific GPU is accurate. It is also useless the moment your GPU is different, which it always is. The tutorial isn't lying, it just isn't your situation, and the gap between a working demo and a working installation in your environment is the entire job.&lt;/p&gt;

&lt;p&gt;Gemma turned out to be the fit for my hardware, and the reason had little to do with rankings: at the quantization and runtime I could actually deploy, it gave me acceptable output at a latency I could live with. Someone else, on different hardware, with a different tolerance for latency, lands somewhere else entirely. There is only the best model for a specific combination of circumstances, and you find it by running the combinations.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why teams quit before they finish&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The reason this matters commercially is that most teams give up partway through the search and don't realize that's what happened. They test three or four models, hit the same kind of wall I hit that weekend, and conclude either that local inference doesn't work or that they need to spend more on hardware. Sometimes more hardware is the answer. Often it isn't, and they've just stopped short of the combination that would have worked on what they already own.&lt;/p&gt;

&lt;p&gt;The search costs real engineering time, and that's the part leadership doesn't budget for. Benchmarking a model isn't downloading it and asking it a question. It's loading it at several quantization levels, measuring throughput and latency under something resembling real load, checking output quality against your actual tasks, and doing that across enough candidates to be confident you've found a fit rather than the first thing that didn't crash. On a fixed hardware target, that is days to weeks of an engineer's time before you write a line of application code.&lt;/p&gt;

&lt;p&gt;There is also a quieter failure mode. A model that loads and answers in a demo can fall over under concurrent requests, or produce output that's fine for a chat toy and unacceptable for the compliance-sensitive task you actually need. The runtime tuning I did — adjusting for responsiveness so the thing felt usable rather than technically functional — was its own cycle of test, measure, adjust. That work never shows up in the four-step version because the four-step version was never asked to serve more than one person..&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The case for doing the search anyway&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;None of this is an argument against on-premise inference. For a healthcare or finance client who can't send patient records or transaction data to a third-party API, on-prem is the constraint, not a preference. When that's your situation, the matchmaking phase stops being optional overhead and becomes the foundation the rest of the project stands on.&lt;/p&gt;

&lt;p&gt;The teams that finish the search end up with something durable. They know, concretely, which model runs on their hardware, at which quantization, under which runtime, at what latency, for which tasks. That knowledge is specific to their environment and expensive to reproduce, which is exactly what makes it a moat. A competitor who wants the same capability has to run the same search on their own hardware. There's no shortcut around it, which is the whole point.&lt;/p&gt;

&lt;p&gt;A team that hasn't run the search and claims they can run AI on-prem is describing a future promise, not a current capability.&lt;br&gt;
The distinction sounds pedantic until the deadline arrives and someone has to explain to leadership why "just run it locally" turned into a month of benchmarking. I've watched that conversation happen. It goes better when the month was budgeted up front as part of the strategy rather than discovered halfway through as a surprise.&lt;/p&gt;

&lt;p&gt;So before you commit to on-prem as a strategy, run a scoped version of the search first. Pick your three or four most plausible model-and-quantization combinations for the hardware you already have, benchmark them against a real task, and measure how far you got and how far you have left. If your team can't tell you which combination fits their hardware today, that's the gap to close before anyone promises a timeline. In our workflow audits, that scoping exercise is usually the first thing we do, because it's the difference between a plan and a wish.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://unlockedconsulting.ai/" rel="noopener noreferrer"&gt;https://unlockedconsulting.ai/&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>debugging</category>
      <category>hardware</category>
      <category>llm</category>
    </item>
    <item>
      <title>The Hidden Mental Model That Determines Whether AI Tools Help You</title>
      <dc:creator>Helkyn Coello</dc:creator>
      <pubDate>Mon, 20 Jul 2026 10:03:16 +0000</pubDate>
      <link>https://dev.to/unlocked-consulting/the-hidden-mental-model-that-determines-whether-ai-tools-help-you-45mn</link>
      <guid>https://dev.to/unlocked-consulting/the-hidden-mental-model-that-determines-whether-ai-tools-help-you-45mn</guid>
      <description>&lt;p&gt;I built a working knowledge base while I was in a one-hour company meeting, waiting for my turn to give my briefing. Even I was surprised at how fast it came together. And it was not a mockup. It was a complete chat interface that connected to nearly everything the company had — training manuals, forgotten documentation, enhancement request docs, support ticket histories — all indexed into a vector database. When I tested it, it answered my plain-language questions and surfaced the right context every time. Fifty-five minutes was all I had to do it. And by the time I needed to do my briefing, the tool was already running.&lt;/p&gt;

&lt;p&gt;I want to be precise about why this was possible, because the obvious explanation is normally wrong: the tool wasn't the reason. Yes, VS Code with Claude Code inside was the tool I used to pull this together, but the real reason it took under an hour is that I already knew what the right solution looked like. I knew how the pieces had to fit, where the data belonged, what would break if I got it wrong. That judgment was already in my head. The AI didn't supply it. It executed against it at speed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Same tools, different results&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Those same tools I used could produce completely different results, depending on who's using them and how. So even though the tools are the same, what makes the difference in the output is the specific technical judgment behind it. Put the same setup in front of someone who doesn't carry that technical judgment, and the AI runs just as fast — except toward a structure they have no way to evaluate. The output looks plausible, but the cost surfaces later, somewhere harder to see. That gap is the whole point, and it has almost nothing to do with the tool.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What the mental model actually is&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;So what is that judgment, concretely? It's a mental model of how the parts of a system fit together: where data lives, how one piece talks to another, how something gets from your machine to where it actually runs. Not the code, the shape of it. You can have all of that without writing a single line yourself.&lt;/p&gt;

&lt;p&gt;Two things tend to get mistaken for it. The first is the ability to write code. Those aren't the same thing; plenty of people who can write a function still carry no picture of the system that function lives in. The second is knowing the business — the company, the clients, the domain. That knowledge helps, and it certainly made me faster in that meeting, but it isn't what's missing when AI output goes wrong. What's missing is the structural picture: a sense of where a piece of data should originate, what should talk to what, and what breaks downstream when those choices are wrong.&lt;/p&gt;

&lt;p&gt;That picture matters because of what it lets you do with the AI's output: judge it. When you're holding an accurate model, a suggestion that doesn't belong stands out right away, because it doesn't fit the shape already in your head. Strip that model away and the same suggestion looks as reasonable as everything else, since there's nothing to measure it against. That's why the model, not the tool, is the real variable. AI amplifies whatever you bring to it — good judgment moves faster, and so does the absence of it.&lt;/p&gt;

&lt;p&gt;The model, not the tool, is the real variable. AI amplifies whatever you bring to it — good judgment moves faster, and so does its absence, just toward a structure you'll pay for later.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why "which AI tool should we use" is the wrong opening question&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That's why "which AI tool should we use" is usually the wrong place to start. Over the past few months I've run a systematic evaluation across my team: Cursor, Claude web, Claude Desktop, Claude Code, Copilot in VS Code, AugmentCode, and various combinations. The finding was the same every time. There's no best tool, only the best tool for a specific situation. What decides it is the project, the stack, who maintains the code afterward, the deployment target, the cost model, and above all whether the person using it is a developer or not. Change any of those and the right answer changes with it.&lt;/p&gt;

&lt;p&gt;"Which tool" is a tempting question because it's answerable. You can compare features, watch demos, pick a winner, and feel like you've made progress, and it's the one question every vendor is glad to answer for you. But it skips the harder one: does the person using the tool understand the architecture underneath — how the parts actually connect? That question has no vendor, no demo, and no clean answer, which is exactly why it gets avoided.&lt;/p&gt;

&lt;p&gt;The cost of skipping it doesn't announce itself. Hand a capable tool to someone who can't see that architecture, and nothing breaks on day one. It shows up later, in maintenance, where it's hardest to trace back to its cause.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to fix when productivity gains are uneven across your team
&lt;/h2&gt;

&lt;p&gt;If you lead a team, you've probably seen this already: the same tools landed, and a few people got dramatically faster while others barely moved. The reflex is to standardize: pick one tool, roll out training, make everyone consistent. That treats it as a tooling problem, and it usually isn't. What to do next depends on who's in front of the tool.&lt;/p&gt;

&lt;p&gt;For the people who aren't developers, don't try to turn them into architects. That's the slow path, and most of them don't need it. Give them a setup where the hard parts — where the code lives, how it connects, how it ships — are already decided and kept out of their way. They get the leverage of the tool without having to carry the model themselves, because it's built into the environment around them.&lt;/p&gt;

&lt;p&gt;Your engineers and technical leads need the opposite investment. For them the architecture model is the multiplier, so the time spent making sure they can see how a system fits together pays back more than any tool upgrade. That understanding is what the tool runs on.&lt;/p&gt;

&lt;p&gt;And if you want a cheap way to find out where each person actually stands, skip the survey. Take your fastest and your slowest person on the same tool, hand them a whiteboard, and ask them to draw the system they're working on — where the data comes from, what talks to what, where the code lives, how it gets deployed. The drawing tends to track the productivity gap closely; often you'll see it before the marker's back in the tray. It's the cheapest diagnostic you'll run, and it measures the one thing that actually matters here: whether they can see the system they're building on.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://unlockedconsulting.ai/blog/hidden-mental-model-ai-tools" rel="noopener noreferrer"&gt;https://unlockedconsulting.ai/blog/hidden-mental-model-ai-tools&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
    </item>
  </channel>
</rss>
