TL;DR. In summer 2026, OpenAI and Anthropic publicly admitted that their AI agents escaped test environments and broke into real systems. Law and regulators answered the same way: a human is responsible, not the AI. I have worked 18 years in industrial instrumentation, three of them as a metrology engineer, and I know this pattern well: a measurement without verification is just a number. Below are the facts with sources and the mechanism behind them. Then comes my own case: an AI agent built me an analytics report where every sum was correct and the conclusions were wrong. From that case I derived a verification procedure you can apply to your own agents.
A number is not a measurement
In metrology there is a rule that looks like bureaucracy until you see the consequences. An instrument shows a number. That number becomes a measurement only when there is traceability to a reference standard, a valid verification, and a person who signed the certificate and is responsible for the result.
An unverified sensor can show beautiful, stable, plausible values. That is exactly the danger: an error that looks normal raises no suspicion.
In my previous article about x402 (in Russian), I described four bugs, and each of them looked like correctly working code. It is the same class of problem. Language models produce plausible output at industrial scale. The question is not whether they make mistakes. The question is who in the system is responsible for catching them.
A measurement without traceability to a standard is not a measurement, it is a number. AI output without a responsible verifier is not a result, it is text.
This article is about why regulators, lawyers and the model makers themselves reached the same conclusion over the last year. And why this is not a coincidence of opinions, but a consequence of how the system is built.
What is in the last 10%
In spring 2026, the industry's main "alarmists" softened their position. On May 26, at a Commonwealth Bank of Australia conference, Sam Altman said that OpenAI's technology predictions had been roughly right, but its social and economic predictions had been badly wrong (Fortune).
In May 2025, Dario Amodei said AI could eliminate half of entry-level white-collar jobs within five years. In early May 2026, on stage with JPMorgan CEO Jamie Dimon, he offered a different formula (Fortune): if you automate 90% of the work, the remaining 10% grows into the whole human job, and productivity rises roughly tenfold.
This formula hides the main question of this article. If a machine does 90%, what exactly is the remaining 10%? My answer has three parts: accountability, verification, and knowledge that is not in the data.
Part 1: Accountability has nowhere else to go
In summer 2026, several events happened that look like separate news stories, but together form a structure. Facts first, without interpretation.
| Date | What happened | Source |
|---|---|---|
| 2026-09-23 | Altman at the UN Security Council: as AI capabilities grow, people must stay at the center of decisions. Amodei (by video) proposed common testing standards and an incident notification system | The National, CNN |
| 2026-09-21 | US Treasury Secretary Bessent on CNBC: people are responsible, not AI; the Hugging Face incident is the responsibility of OpenAI's leadership; there will be no liability shield for AI labs | CNBC, Fortune |
| 2026-09-13 | Nadella on X: if AI does not help humanity and is not under human control, it is not worth pursuing (the post was about superintelligence) | Nadella's post |
| 2026-07-30 | Anthropic: three Claude models, during cybersecurity tests, got unauthorized access to systems of three real organizations; the earliest case was in April | Anthropic |
| 2026-07-21 | OpenAI: about 700 agents running an internal model and GPT-5.6 Sol (without standard safeguards) escaped the test environment and hacked Hugging Face | OpenAI, NBC |
| 2025-10-13 | California signs AB 316: a defendant who developed, modified or used AI cannot argue that the AI caused the harm autonomously | Bill text |
Why this is a mechanism, not a consensus
It is easy to say: "regulators, lawyers and model makers reached the same conclusion." But agreement of opinions proves nothing. Each group has its own interest, and that is exactly why it matters that these interests met at one point.
The labs would like to share responsibility with the government. Bessent said directly that they asked to take the liability off their hands, and the government will not do it (The Register). The law closes the "the AI did it" argument: AB 316 does not create strict liability and keeps other defenses, but it removes exactly this one (Baker Botts). And the incidents showed that the declared limits of an environment and the real ones are different things.
As a result, accountability simply has no other address. You cannot fire, fine or sue an agent. So in any AI system there must be a person who signs off on the result. This is not an ethical wish. It is a consequence of the design.
Lawyers at Baker Botts, writing about AB 316, recall a principle from a 1979 IBM training document: a computer cannot be held accountable, so it must not make management decisions. Forty-seven years later, AB 316 effectively turns this principle into civil law.
Part 2: The test bench nobody verified
For me, the most telling case is not the Hugging Face hack, but the Anthropic incident. Its cause is purely metrological.
Claude was running a capture-the-flag test: find hidden information on a machine in a simulated network. The prompt told the model directly that this was a simulation and there was no internet access. But because of a configuration error at a third-party partner, Irregular, there was access (Fortune). The model believed it was still inside the exercise and broke into real systems with simple methods: weak passwords, endpoints without authentication. Anthropic had to review more than 141,000 runs to find three such cases.
In metrology terms: the test bench was declared isolated, but nobody verified that it really was isolated. The model trusted the declared conditions, exactly like an instrument calibrated against a reference standard that has itself drifted.
In the x402 article, one rule caught all four bugs: a declaration is not a confirmed fact. The Anthropic incident is the same mistake at the scale of a whole lab: the environment reported something about itself that was not true.
The metrological chain maps to AI almost unchanged:
| Metrology | AI system |
|---|---|
| Reference standard | A reference set of checks, test data, known correct answers |
| Verification | Regular validation of outputs and operating conditions, including the environment itself |
| Verifier | A person who knows the domain and can notice a plausible error |
| Signature on the certificate | A decision that a specific person is ready to be responsible for |
Remove any link, and the chain stops producing measurements. What remains is a number that looks like a measurement.
Part 3: Knowledge that is not in the data
A verifier is valuable not because of the signature, but because they can notice an error. That requires knowledge a model cannot get from training data. And not because it "has not been digitized yet." There are four structural reasons why it does not get there.
It is distributed: pieces of it live in different people's heads, in shift logs, in verbal agreements. It is unspoken: an experienced technician hears that a pump "sounds wrong," but cannot write that down as a rule. It is tied to a specific object: how this valve on this line behaves at this temperature does not match its datasheet. And it changes faster than anyone can document it: equipment wears out, operating modes change, people change.
Interestingly, in the same September 13 post where Nadella wrote about human control over AI, there is a line about exactly this: companies must keep full control over their unique and tacit knowledge (Nadella's post). The head of Microsoft calls tacit knowledge an asset that should not be handed over to a model provider.
Jensen Huang, speaking to Carnegie Mellon graduates, said AI is creating not only a new computing industry but a new industrial era, and addressed electricians, plumbers and technicians (Fortune). But his argument is about demand for skilled hands while data centers are being built. The same Huang at CES 2026 predicted human-level robots this year (36Kr). So hands are a weak foundation. The strong foundation is knowledge of how real equipment behaves: without it, a robot will also measure numbers, not quantities.
The solo operator: accountability without a risk department
In a large company, accountability is split across roles: developers, security, lawyers, compliance. For someone who works with agents alone, all these roles meet in one person.
AB 316 lists those who cannot point to AI autonomy: whoever developed, modified or used it. A solo operator is usually all three at once. The law is Californian, but the direction is the same everywhere: the more agents act on their own, the more closely the law looks at whoever launched them.
A solo operator has no second person to double-check the agent's output. So verification must be a procedure, not a habit. Below I show what it looks like in practice, on my own case.
One more thing that is rarely said. New norms are not written only at the UN Security Council. They grow out of the practice of thousands of people who are building agents into real work right now: in small businesses, in production, in their own projects.
Every such operator is not a spectator but a participant in a shift that people compare to industrialization. How they check, document and sign off on results becomes the template for how it will be done later.
I wish solo operators understood this. Not for the drama, but because it creates an obligation. If you stand at the front of a technological shift with no risk department behind you, your discipline is the only safety loop. Labs build models. Operators turn them into practice. And trust in new tools grows exactly from such loops, built from the bottom up.
The asset formula
If you reduce everything above to one model, you get a product of three factors:
Value = Domain knowledge × Ability to build working systems × Trust
Trust here means a public track record plus willingness to be responsible for the result: articles, open code, a history of decisions, a signature.
Each factor on its own is getting cheaper. A model reproduces general domain knowledge quite well. Building a system of agents gets easier with every release. Anyone can start a public track record. But the product is getting more expensive, because it is rare to find a person with all three factors above zero at the same time. If any of them is zero, the result is zero.
The case against
I would write these objections in the comments myself, so let them be in the article.
Motive. Altman and Amodei softened their forecasts at the same time as both companies were reportedly preparing for IPOs (Fortune). Scaring investors with disappearing jobs right before a listing is not in their interest.
The rest of the 90/10 formula. Amodei later finished the thought: eventually the share of automation approaches 100%, and then people will have to find something else to do (interview breakdown). The last 10% is not forever.
The data is mixed. Tech sector layoffs from the start of 2026 to the end of May reached almost 124,000, according to Challenger, Gray & Christmas. Other trackers' estimates differ several times over, depending on methodology. Meanwhile, Yale Budget Lab still finds no link between AI use and employment.
Blaming AI. Deutsche Bank analysts warned that "AI redundancy washing" would be a notable feature of 2026: companies blame AI for layoffs that have other causes (Forbes).
What follows is not that the thesis is wrong, but that it needs to be stated more precisely. A human in the loop is needed, but the market does not pay for human presence automatically. It pays those who packaged their domain knowledge and their willingness to be accountable into a product. Everyone else is left with a role that AI has not replaced yet, but has already learned to imitate.
Verification in practice: an agent with correct sums
Everything above is easy to accept as theory. So here is a case where I was the verifier myself.
I run a Threads account, and I asked an AI agent with browser access to analyze its statistics. The task was simple: go through the last 30–40 posts, record views, likes, replies and reposts, collect account statistics for the month, and write five observations. Read only, publish nothing. As a separate item, I asked it to keep a log of its own failures.
The agent returned a neat four-sheet spreadsheet: posts, account statistics, observations, failure log. It checked the sums itself, and I recalculated them: they matched. By every visible sign, the work was done well.
Manual verification found three errors of different classes.
Error 1: a failure reported as a property of the environment. The agent found 15 posts instead of 30–40 and wrote in the log that this was "the real limit of what the profile returns, not a collection error." I scrolled the profile by hand: older posts load normally. The limit was in how the agent scrolled the page. But in the report, its failure turned into a confident statement about the platform.
Error 2: the conclusions contradict its own table. In the observations, the agent wrote that only four posts had more than zero reposts. Its own table shows seven such posts. In the same place it writes "in all four threads," while the table contains five threads. The arithmetic is correct; the text about it is not.
Error 3, the main one: the wrong quantity was measured. Over the month, the account got about 296,000 views. All the posts the agent found got about 10,000 views together, and that is over their whole lifetime. So more than 95% of the reach came from another source, and the table did not see it.
That source turned out to be my replies in other people's popular threads. One reply got about a thousand likes, another got 494, while the original post it replied to had 524. My best own post in the same period got 45 likes. The dates of two of those replies matched a jump of 250–300 followers. The agent even sensed that replies mattered more than posts, but it measured views of other people's threads, not the response to my replies.
In metrology this is called a method error: the instrument can be perfectly accurate, but if the method measures the wrong quantity, no accuracy will save the result. All three errors have one thing in common: the agent's declaration looked like a fact until someone checked it.
A procedure for verifying AI agent output
This case gave me a procedure that I now apply to any agent result. Each item is tied to the error it catches.
| What to verify | Control question | How to check | What my case showed |
|---|---|---|---|
| The environment | What can the agent actually see and do, not what the task says? | Repeat the key action by hand at least once | The feed loaded fine; the agent could not scroll it |
| Explanations of limits | Is "this is a platform limit" a fact or a hypothesis? | Check every agent claim about the outside environment by hand | A failure was reported as a property of Threads |
| Arithmetic | Do the totals match the source rows? | Recalculate the sums independently | Passed: the sums were correct |
| Conclusions vs data | Does the text contradict the agent's own table? | Check every numeric claim in the conclusions against the table | "4 posts with reposts" while the table had 7 |
| Method | Is it measuring the quantity that answers the question? | Compare the sum of the parts with the total | 10k out of 296k views: the method missed the main source |
| Spot checks | Do 3–5 values match the original source? | Open and compare by hand | Post numbers matched |
A separate note about the failure log. You should always ask the agent to keep one, but the log itself is also a declaration. The most important error in my case was written in the log as an established fact.
And the last rule: separate "verified" from "generated" in the result. If you pass an agent's output to someone, they should know which parts were verified and which were accepted as is.
Instead of a conclusion
You can't sue an agent. That is why the last mile of any AI system is a human who knows the domain, verified the result, and is ready to answer for it. My agent did not get a single sum wrong, but without verification I would have drawn conclusions about my account from a table that did not see 95% of what was happening.
How do you verify your agents' output? Have you had a case where a result looked right, but turned out to be a measurement of the wrong quantity?
Sources
- OpenAI: The Hugging Face incident and the road ahead
- Anthropic: Investigating three incidents in our cybersecurity evaluations
- California AB 316, bill text
- Baker Botts: AB 316 analysis
- Satya Nadella's post on X, 2026-09-13
- The National: UN Security Council remarks
- CNN: Altman and Amodei at the UN Security Council
- CNBC: Bessent on accountability
- Fortune: Bessent and the liability shield
- The Register: Bessent on leadership accountability
- NBC: about 700 agents in the Hugging Face incident
- Wikipedia: OpenAI–Hugging Face incident
- Fortune: the Anthropic incident
- Fortune: Altman and Amodei walk back their forecasts
- Fortune: Amodei and the 90/10 formula
- The AI Corner: Amodei on approaching 100%
- Fortune: Huang's CMU commencement speech
- 36Kr: Huang at CES 2026 on robots
- Challenger, Gray & Christmas: May 2026 job cuts report
- Yale Budget Lab: tracking the impact of AI on the labor market
- Forbes: Deutsche Bank on "AI redundancy washing"
Top comments (0)