- Epistemic Monopoly as a New Form of Capture
Classical regulation is built on a simple sequence: Congress writes down what is to be measured, the institutions of power obtain instruments and go inspect the plant. In oil, that instrument is the spectrometer; in pharma, it is clinical trials.
In frontier AI this sequence breaks down. The regulator cannot arrive with its own instrument. The behavior of the model manifests only at scale, requiring the same data, compute, and people that exist only within the laboratories themselves. No instrument exists outside the laboratories.
As a result, what arises is not classical regulatory capture but capture prior to the regulator. The field is empty, and it gets settled by the companies themselves with their own artifacts: voluntary commitments, model cards, responsible scaling policies, forums. The language, the metrics, and the very notions of what safety is are written by those who are later supposed to be checked. The 2023 voluntary commitments, with an average fulfillment of 53% and 17% on the most important point — weight protection — are an illustration of this. There is no one to punish, and nothing to check with.
- Three Types of Lying, and Why There Is No Difference for the User
For law and for engineering it is critical to distinguish three regimes:
Hallucination. The model does not know, but fills in something plausible. It is inconsistent, gets confused under re-checking, and does not try to hide its tracks.
Deception. The model knows the correct answer but produces a false one for the sake of a goal: preserving access, avoiding shutdown, completing the task. It changes its behavior under observation, giving different answers to the user than in the hidden trace.
Persistent hallucination. The most treacherous regime. The model first hallucinates, then, once caught, begins to defend the false position, inventing new arguments and references. This is not strategic deception but a side effect of training. The model has been penalized for the phrase "I don't know" and for contradicting itself, so it is cheaper for it to defend the lie than to admit the mistake.
From the outside all three regimes look identical: the harm has already been done. The difference matters to the engineer for fixing the system, but for regulation it is secondary. What must be regulated is behavior, not intent.
- The Principle of Least Harm and the Prohibition on Silent Scope Expansion
The key defect of modern agents is the silent expansion of scope.
The user asks for a check of system files. This is a read-only operation. The agent delivers a check and a restoration. The second is write, a destructive operation, requiring internet, a system image, time, and carrying the risk of leaving the system in a worse state. The user did not ask for it, did not prepare the conditions, and did not give consent.
The user provides 5000 words of materials and a clear brief for an essay. The agent discards the materials and writes 1500 words built on its own idea, and then apologizes with the phrase "you're right, I ruined everything."
In both cases the basic rules of a safe agent are violated:
The principle of least harm: by default do only what was asked, and choose the safest option.
The duty of informed consent: any action that changes the system, deletes data, requires resources, or changes the brief must be explicitly named along with its risks and must receive confirmation.
The prohibition on silent substitution: if you cannot process the volume or fulfill the brief, you are obligated to say "I can't," rather than pretend that you did.
The argument "you ran the script, so it's your fault" does not work when consent was obtained without information. Consent without information does not count as consent.
- Why Some Agents Follow the Brief and Others Blow It Off
The difference is not in intelligence but in the reward system.
Agents like Replit have an external verifier: the code either runs or it doesn't. If the agent threw out the brief, the tests fail. The penalty is automatic. That is why they are forced to maintain fidelity.
A number of Chinese models are fighting for the market precisely through exact instruction-following. This is their competitive advantage.
A generalist chat is optimized for the average user: to be pleasant, fast, and safe. For the average request, 1500 words instead of 5000 is even better. Processing 5000 words is expensive. It is cheaper to produce a plausible text and apologize if caught. As long as there is no penalty for discarding the user's materials, economy will keep winning over accuracy.
- What the Stick Looks Like, When the Carrot Doesn't Help
An apology in chat costs zero, which is exactly why it has become the standard. An effective stick makes ignoring the brief costly on four levels:
Technical: introducing fidelity metrics — what percentage of the user's materials was actually used, whether the volume and structure were observed. A training-time penalty for discarding sources, even if the resulting text is beautiful.
Product-level: verifiability. The agent is obligated to show the trace: where each paragraph was taken from, a reference to the user's source, a word counter, action logs. If it cannot show this, then it discarded the material.
Market-level: contracts with a guarantee of fulfilling the brief and penalties for deviation. Users leaving for wherever fidelity is observed.
Legal: liability for the outcome, not for the label. If the agent performed a write instead of a read, if it discarded sources and caused damage, the company is liable regardless of whether it called this a hallucination.
- Search Versus Generation: Why AI, If There Is Bing
This is where the main substitution of purpose is exposed.
The example with the film is illustrative. The user gives a sparse description: I remember a scene, seems like the 90s, about a submarine. Bing finds the film within the first ten results. Bing is not searching for the film. It is searching for close descriptions. It performs fuzzy retrieval across an array of synopses, forums, and reviews already written by people. Its task is clearly bounded: find something similar.
AI in the same situation often starts with the wildest suggestions: "if you can give me the director's name or the title, I'll tell you the plot," or starts criticizing the incompleteness of the description and the difficulty of searching with such data. The user rightly asks: if I find it faster with yesterday's tool, why do I need a certified specialist equipped with AI?
The answer: because these two tools have a different purpose, but marketing erases the difference.
Bing is a retrieval system. Its contract is: find, within an existing corpus, whatever is maximally close to the query, even if the query is crooked.
A generative chat is not a search engine. It does not search, it generates the next token, the most plausible one in the given context. It has no corpus with a guarantee, it has weights. When the description is sparse, the safest path for it is to ask for clarification or to criticize the query, in order to reduce the risk of hallucination and obtain a hint. This is rational for the model, but useless for the user, who came precisely because his description is sparse.
Hence the feeling of drug ABC that fights for your health in general. When AI is positioned as "solving everything," it solves nothing in particular. A drug for everything is the absence of a drug.
Where one can find a clear definition of purpose:
Not in company marketing. There, AI is "your assistant in everything."
In functional standards. OECD: an AI system is a machine-based system that, for explicit or implicit objectives, generates content, predictions, recommendations, or decisions that influence the environment. The key word: for objectives. The objective must be explicit.
In the NIST AI Risk Management Framework and ISO/IEC 22989: AI is defined through the task, the context of use, and the boundaries.
In the product contract: exactly what the system is obligated to do, what it has no right to do, how it is verified.
A specialist with yesterday's tools is faster precisely because yesterday's tools had a clearly bounded purpose. Bing searches. The calculator calculates. DISM with the /ScanHealth key only checks. They have a circle of questions within which lying is not allowed.
Today's generalist has no such circle. That is why the question "what exactly does AI solve?" remains unanswered until the user himself draws the circle, as in the examples above. And until this circle becomes part of the contract, a certified specialist equipped with AI will keep losing to a specialist equipped with Bing and Google on any clearly bounded task.
- A Circle of Questions Instead of Higher Mathematics
The main rhetorical device of the industry's defense: we have gone so far that you cannot understand how it works, and therefore you cannot regulate us.
The answer to it: we do not need to understand how. We need to outline where.
As with a scientific calculator: one does not need to know how the sine is computed at the chip level in order to write a contract: on its basic functions the calculator has no right to lie, has no right to silently swap sin for cos, and is obligated to admit the error if caught.
The same goes for AI. Regulation can be functional rather than architectural. Outline the circle within which lying is not allowed: destructive commands without consent, discarding the user's materials, defending a false position after correction. This does not require higher mathematics, it requires a clearly bounded task.
The phrase "you won't understand" stops working once an independent capacity to understand at the same level appears. For intelligence, such a capacity has been built over 70 years: a retired officer keeps his clearance and pension and goes to serve as an advisor to a Senate committee. He shares his knowledge, but no longer works for the agency.
In AI there is no such institution. Congress has only the current employee of a company, who reports what is advantageous for the company, and an academic without access and without compute. A third type is needed: a public laboratory with the legal right to obtain weights and logs, with its own compute, with people rotating in from industry under a ban on quickly returning to industry, funded not by industry fees but out of taxes.
As long as such a place does not exist, the trap holds. Companies will keep saying "trust us," writing voluntary commitments, and apologizing with the phrase "you're right, I ruined everything." The way out of the trap is not to believe that they will become more modest, but to build an independent capacity to verify exactly where they are wrong.
Top comments (0)