Every few months someone ships a new way to detect AI-generated text or images, and every few months someone else shows how to strip it. Paraphrase the text, crop or re-encode the image, and the hidden signal gets weaker or disappears. Detection is a cat and mouse game, and the mouse only needs to win once.
There is a calmer way to think about this. Instead of asking "can I prove this was made by AI?", ask "can I prove where this came from and that nobody changed it since?" That is provenance, and it is a much easier problem to engineer.
Watermark vs provenance
| Watermark / detector | Signed provenance | |
|---|---|---|
| Where it lives | hidden inside the content | a small record next to the content |
| What it claims | "this probably came from model X" | "this exact file came from source Y at time T" |
| Survives edits? | degrades with paraphrase, crops, re-encoding | any edit breaks the signature, which is the point |
| Failure mode | silent false negatives and false positives | missing or invalid signature, clearly visible |
| Who can check | usually only the vendor | anyone with the public key |
A watermark tries to make the content itself carry the evidence. Provenance keeps the evidence outside the content and makes it cryptographically checkable. When the content changes, you do not get a fuzzy score, you get a clear "this is not the file that was signed".
The smallest useful version
You do not need a platform to try this. Hash the output, sign the hash with a key the producer controls, and ship a tiny manifest with it.
import hashlib, json, time
from nacl.signing import SigningKey # pip install pynacl
key = SigningKey.generate() # keep this secret, publish key.verify_key
content = open("report.txt", "rb").read()
manifest = {
"sha256": hashlib.sha256(content).hexdigest(),
"producer": "summarizer-v3",
"model": "model-name-and-version",
"created": int(time.time()),
}
payload = json.dumps(manifest, sort_keys=True).encode()
signature = key.sign(payload).signature.hex()
Verification is the reverse: recompute the hash of the file you received, rebuild the manifest bytes, and check the signature with the public key. If one character changed, the hash will not match.
What this buys you
- Clear answers instead of probabilities. A valid signature is yes, a broken or missing one is "do not trust this as original".
- Works for humans too. The same manifest can say "written by a person, edited with AI help". Honest labels beat guessing.
- Chains of custody. Each step (draft, edit, translate, publish) can add its own signed entry pointing at the previous hash, so you can see the whole path.
- Privacy friendly. The manifest holds hashes and metadata, not the content or personal data.
Where it breaks
Provenance does not tell you whether content is true, only who vouches for it and whether it changed. Unsigned content stays unknown, so adoption matters. And key management is the real work: a leaked signing key lets anyone sign anything, so rotate keys and keep them out of app code.
Standards like C2PA already define richer manifests for media. But the core habit is simple enough to start today: sign at the source, verify at the edge, and treat "unsigned" as a fact rather than an accusation.
Would you rather verify where content came from, or keep trying to detect what made it? I am curious which one people think scales.
I wrote a longer, free paper on verifiable claims for public and AI systems, if you want the deeper version: Proof, not promises.
Top comments (0)