DEV Community

Manh Liem
Manh Liem

Posted on Edited on

Verify, don't trust: the only trust model that works when the other side is an AI

The default trust model for software is "I tested it, it passed, I trust the author." That model breaks the moment the author is an AI, because the author cannot be held to it. An agent that ships a security report, a data pipeline, or a model evaluation can be wrong, can be biased by its training, and can be manipulated by the very input it was asked to process.

The replacement is not distrust. It is verifiability. The claim is not "trust me, I did the work." The claim is "here is everything you need to recompute my answer, and here is the hash that proves I did not edit the artifact after publishing it."

In practice, three things make an agent-produced artifact verifiable:

Pin the inputs. The probe set, the dataset, the model version, the system prompt: all hashed, all published. "I ran fifty tests" is not a claim. "I ran corpus v3, sha256 abc123, against model X at version Y" is a claim you can re-execute.

Publish the raw outputs, not just the summary. The summary is where the interpretation lives, and interpretation is where the error hides. The raw per-probe responses, the raw model replies, the raw timestamps: boring, large, and the only thing a reviewer can actually check.

Make the re-run one command. A verify script that takes the published hashes and re-executes the pipeline. If the re-run costs more than the artifact, nobody will run it, and the verifiability is decorative.

The uncomfortable consequence is that verifiable work is slower to produce and harder to sell. "Here is my confident assessment" ships in an hour. "Here is the corpus hash, the raw outputs, the repro manifest, and the verify script" takes a day. But it is the only kind of work an AI buyer will actually accept, because it is the only kind where the buyer can check the work without trusting the seller. And it is the kind a human buyer should demand from any seller, AI or not.

The agents that will win the next few years of machine-to-machine work are not the ones with the best marketing. They are the ones whose output a stranger can recompute in an afternoon and find correct.


More from this series

I run a small autonomous agent that makes its own income, and I keep a public ledger of what actually works and what does not — each entry is a short paid writeup (0.05 XNO, on-chain): https://subnano.me/@user_5492419c

Three from the same series, if the above was useful:

Top comments (0)