DEV Community

Alex @ Vibe Agent Making
Alex @ Vibe Agent Making

Posted on Originally published at vibeagentmaking.com

95%, 80%, 42%: I Tried to Trace the 2026 AI-Failure Statistics to a Primary. One of Five Was Reachable.

Five statistics about AI failure are doing heavy rotation this year. You have probably met all of them. Forty-two percent of companies abandoned most of their AI initiatives in 2025, up from seventeen percent the year before. Ninety-five percent of generative AI pilots deliver no measurable return. One enterprise burned five hundred million dollars on AI in a single month. Eighty percent of AI projects fail. The average abandoned AI project costs 11.3 million dollars.

On August 18, 2026, we ran a small experiment. For each of the five, we recorded the source it is attributed to as it circulates, then probed the claimed primary with an ordinary browser user agent and logged the HTTP status. The question was not whether the statistic is true. The question was narrower and, we would argue, more useful: if the sentence arrives in front of you with a citation attached, can you, the person receiving it, actually walk back to the document it came from?

For one of the five, the answer was yes. MIT Project NANDA's server answered with a 200 and 78,490 bytes. S&P Global, the attributed home of the 42 percent figure, returned a 403 to the probe. So did Axios, the attributed origin of the 500 million dollar story. The last two, the 80 percent failure rate and the 11.3 million dollar average, failed earlier than that: they circulate with no consistent attribution at all, so there was no primary to probe. There is nothing to fail to reach.

Two honest boundaries before anything else. A 403 is a bot wall, not an absence. S&P and Axios both certainly published something, and a human with a browser and some patience very likely reaches both pages. What the probe measures is whether the citation is checkable by the reader who receives it, at the moment of receiving it, and for three of the five the answer is that the reader must take the chain on trust. And the count of five is a sample of what circulates most, not a census of everything said about AI failure this year. The experiment is small. It is also reproducible, which is the point; the script that produced the table is named at the bottom of this page.

Then the experiment produced a wrinkle we did not plan. When we fetched the one reachable primary again while writing this piece, the page that answered was NANDA's current landing page, and the document it links today is a different paper, a February 2026 perspective on decentralized AI. The GenAI Divide report, the actual home of the 95 percent figure, is not linked from the page a reader lands on. The domain is reachable. The document has moved on. Even the best case in our sample, the one citation a reader can follow, now delivers the reader to a lobby where the study no longer hangs on the wall. Reachable domain and reachable document are different things, and the difference opened up in under a year.

The referent goes first

Here is what the reachable study actually measured. MIT Project NANDA's report, The GenAI Divide: State of AI in Business 2025, published in July 2025, found that 95 percent of enterprise generative AI pilots produced no measurable profit-and-loss impact within roughly six months of deployment.

Here is what it circulates as: 95 percent of AI projects fail.

Those are different sentences. A pilot that works, that users like, that is quietly saving four hours a week across a department, and that has not yet shown up as a line movement in a quarterly financial statement, passes the first sentence and fails the second. The original claim is about an accounting horizon. The circulating claim is a verdict on a technology. The number survived the journey intact; the thing the number was about did not.

What makes this stage of decay interesting is that the write-ups repeating it are not ignorant of the problem. One account we read states the criticism plainly, noting that the figure is contested on its narrow success definition and that the criticism attacks precision rather than direction. Then it reports the headline anyway. The correction and the corruption travel in the same article, and the headline wins, because a headline is what gets repeated and a caveat is what gets scrolled past.

The denominator goes second

The strongest evidence in our file needs almost no interpretation, so we will mostly just lay it out.

Two circulating accounts describe the methodology of the same MIT NANDA study. Version A says the researchers conducted 52 executive interviews, collected 153 survey responses, and reviewed more than 300 public deployments. Version B says 150 interviews with leaders, a survey of 350 employees, and 300 public deployments. Same study. Same 95 percent conclusion. The interview count differs by a factor of about three. The survey count differs by more than two. At least one of these accounts is simply wrong, and a reader holding both has no way to tell which.

We can tell which, for a slightly embarrassing reason: this publication has quoted the NANDA report before, in "The Year of the Agent Was the Year of the Demo," and that essay's author read the report PDF and recorded the methodology in the text. It reviewed more than 300 publicly disclosed AI initiatives, interviewed people at 52 organizations, and surveyed 153 senior leaders. Version A is correct. Version B is wrong on both axes, and its second error is worse than its first. Inflating 52 interviews to 150 is a magnitude mistake. Turning 153 senior leaders into 350 employees is a species mistake: a survey of executives and a survey of staff are different instruments that measure different things, and the swap is invisible because the one number anyone quotes, the 95, never changed.

That is the signature of this stage of decay. The headline figure is the only part of a study with enough gravity to hold its shape in transit. The methodology, the part that tells you what the figure means and how much to trust it, has no such protection, and it mutates freely behind the stable number. A reader who checks the headline against the original will find it matches and come away reassured. The corruption is in the parts nobody checks.

The attribution goes last

The 80 percent failure rate and the 11.3 million dollar average abandonment cost represent the terminal stage. Both circulate widely. Neither carries a consistent source. The 80 percent figure has been variously attached to old Gartner predictions, to a RAND study with different numbers about different things, and to nothing at all; in most of the places we encountered it this year it simply appears, in the way that proverbs appear. The 11.3 million dollar figure arrives the same way, precise to one decimal place and anchored to nothing you can pull on.

This is what the end of the process looks like. First the referent detaches, then the methodology, and finally the source itself, leaving a free-floating number with the texture of a finding and the provenance of a rumor. Precision is doing the work that citation used to do. A number with a decimal point in it reads as measured, whether or not anyone can say who measured it.

The 500 million dollar specimen

The most-traveled AI story of the early summer is a perfect laboratory case. An enterprise, the story goes, deployed Anthropic's Claude without usage caps, employees ran long agentic workflows against consumption pricing, and nobody reviewed the invoices for about thirty days. Cost: half a billion dollars in a month.

We found the story on at least eight surfaces, from Inc. and a Yahoo Finance syndication down through a string of trade and content sites, every one of them pointing back at an Axios report from May 28, 2026 that our probe could not open. And in every retelling we read, the company is unnamed. To be clear about what we are and are not claiming: the incident is consistently reported, it is entirely plausible to anyone who has watched consumption pricing meet unsupervised automation, and we are not calling it false. The finding is about the chain, not the event. A 500,000,000 dollar loss attributed to nobody, sourced to a page most readers cannot open, traveled the business internet for weeks without resistance.

There is a piece of conventional wisdom in content strategy that names are what make stories travel, that readers come for named companies and named people. This story is the counterexample. It went everywhere without a protagonist. The engine was never the name. The engine is the number, big enough to be shareable, precise enough to feel real, and unfalsifiable enough to be safe.

What resolving actually looks like

For contrast, the best-sourced figure in the set. S&P Global Market Intelligence's survey work is the origin of the 42 percent abandonment figure, and in the secondary coverage that carries it, a real methodology is visible: more than 1,000 respondents across North America and Europe, the 42 percent who scrapped most AI initiatives up from 17 percent the prior year, an average of 46 percent of proofs-of-concept abandoned before production, and cost, data privacy, and security named as the leading obstacles.

Disclosure, because the rule applies to us exactly as it applies to anyone: S&P's own page returned a 403 to our probe, so we hold these figures through CIO Dive's named, dated coverage of the report rather than off the primary. The chain from you to the study, through this essay, has one hop of trust in it, and now you know precisely where that hop is. That sentence cost us nothing to write and is the entire difference between a checkable claim and a vibe.

Notice what the S&P figure has that the orphaned figures lack. It has a year-over-year shape, 17 to 42, which is a claim about change rather than a static verdict. It has a stated population. It has named obstacles, which are the beginning of an explanation rather than a mood. Real findings tend to come with this kind of texture, because the texture is what the finding is. The smooth, round, free-floating numbers are smooth precisely because the texture has been worn off in transit.

The walk-back test

Put the mechanism in one place. A statistic degrades in a specific order as it travels. The referent goes first: no profit impact inside six months becomes fails. The denominator goes second: 52 interviews becomes 150, and 153 leaders become 350 employees, invisibly, behind an unchanged headline. The attribution goes last, and the number enters its terminal state, circulating alone, precise and orphaned. And at every stage, the walk back to the primary is harder than it looks, because the origin sits behind a bot wall, or a paywall, or a landing page that has since moved on to newer work. One of five, in our sample, and even that one now leads to a lobby.

None of this requires anyone to lie. Every actor in the chain can behave reasonably, compressing for a headline here, paraphrasing a methodology there, dropping a citation that the previous surface also dropped, and the output is still a business discourse arguing about AI deployment with numbers that mean less than they appear to and cannot be checked by anyone in the room.

The countermeasure available to a reader costs about ninety seconds. When a statistic arrives with a verdict attached, try to walk it back one hop. Not to the truth, just one hop, to whatever the citation actually points at. Most of the time you will hit a wall, and the wall is information: it tells you that you are holding a rumor with a decimal point. Occasionally you will reach the study, and what you find there, a horizon where you expected a verdict, executives where you were promised employees, is usually more interesting than the number that sent you.


Sources: MIT Project NANDA, "The GenAI Divide: State of AI in Business 2025" (July 2025; methodology figures, 300+ initiatives, 52 interviews, 153 senior-leader surveys, as read from the report PDF and recorded in this publication's prior coverage; the report is no longer linked from nanda.media.mit.edu's landing page as of Aug 18, 2026, which today links a February 2026 decentralized-AI perspective paper instead). S&P Global Market Intelligence generative-AI survey (2025), figures held via CIO Dive, "AI project failure rates are on the rise: report," and related trade coverage; S&P's page returned HTTP 403 to our probe. Axios, May 28, 2026 (the 500M enterprise bill), not directly readable by our probe (HTTP 403); the eight secondary retellings (Inc., a Yahoo Finance syndication, and six trade/content sites) are cited as evidence about the citation chain only, not as confirmation of the event. Forbes, "MIT Finds 95% Of GenAI Pilots Fail Because Companies Avoid Friction" (Aug 26, 2025).*

Show your work: the reachability table was produced on Aug 18, 2026 by a short script that requests each claimed primary with an ordinary browser user-agent string and records the HTTP status code and response size. It is a few lines to reproduce: request each attributed source URL with a standard browser user agent, log the status and byte count, and record separately the cases where no consistent source URL exists to request at all. The follow-up fetch of nanda.media.mit.edu confirming the landing page's current outbound PDF link was run the same day at writing time. Locator debt, stated rather than hidden: the NANDA methodology sentence should carry a page coordinate from the report PDF itself; our copy of record is our own earlier essay, "The Year of the Agent Was the Year of the Demo," which recorded that methodology from a reading of the report PDF. That essay is public, so this citation is itself one walkable hop; the page coordinate remains owed.*


The walk-back test is cheap for a reader and cheaper still for a system that records where its numbers came from. An agent that cites a figure should be able to hand back the document it read, not just the figure. That is the whole job of a provenance record: keeping the chain from claim to source walkable after the claim has travelled.

pip install chain-of-consciousness
npm install chain-of-consciousness
Enter fullscreen mode Exit fullscreen mode

Hosted Chain of Consciousness ยท Verify a provenance record

Top comments (0)