DEV Community

Craig Solomon
Craig Solomon

Posted on

What a timestamp proof actually proves, and the bytes you have to keep

Ask this about your own setup: if a timestamp proof verifies clean, what are you entitled to say out loud?

Here's the honest version. This exact digest existed no later than the block it was anchored in.

That's the whole claim. Not that the file is genuine. Not that you wrote it. Not that nobody has touched it since. A hash existed at a time, and nothing past that. Everything useful you can build on a timestamp proof rests on that sentence.

Start with the easy half, because you can run it right now:

pip install verify-proof
verify-proof hash contract.pdf
Enter fullscreen mode Exit fullscreen mode

That gives you a SHA-256 digest of the file's bytes. The verify subcommand takes a proof file and the hash it's supposed to be for, recomputes that hash, and walks the Merkle path to the root itself instead of asking an API whether things look fine. Verification is local, no key and no account. The only subcommand that touches the network is create, which anchors a hash through ProofLedger and sends the hash, never the file.

So far so boring. Here's where it gets interesting.

The claim is attached to a digest, and the digest is attached to exact bytes

Once you accept that the proof is a statement about a digest, the next question is what you have to preserve for that statement to stay worth anything. The obvious answer is the proof file. Archive the proof, keep it somewhere durable, done.

That answer is not enough, and the gap is the point of this post.

A proof file is a pointer to a digest. The digest is a pointer to a specific sequence of bytes. If you still hold the proof but can no longer produce bytes that hash to that digest, you hold a verifiable statement about something you can't show anyone. The math still checks out. It just checks out about an artifact you lost.

And you can lose it without ever losing the file, which is the part worth internalizing. The digest is over bytes, not over content. Not over "the document", not over the text a human reads on screen. Change the container and leave the meaning identical, and the digest is gone.

Things that rewrite bytes while leaving content alone

I'm not going to tell you what's common in other people's pipelines, because I haven't read them. I'll tell you what to go look for in yours. Any step that deserializes and re-serializes is a candidate:

  • A PDF opened and re-saved by a different tool, which may rewrite the object table, the producer string, or embedded dates.
  • JSON parsed and dumped again, with different key order, different whitespace, or different unicode escaping.
  • Line ending conversion on checkout, or an editor normalizing the trailing newline.
  • Images re-encoded on upload, stripped of metadata, or thumbnailed in place.
  • A zip or tarball rebuilt, with different member order, different mtimes, different compression levels.
  • Office formats, which are zips, so everything above applies twice.

None of those are attacks. They're ordinary helpful behavior from ordinary tools. The proof doesn't know the difference between a helpful re-save and a malicious edit, and it shouldn't. That's not weakness, it's the thing you wanted. A digest that forgave cosmetic changes would also forgive the change somebody made on purpose.

Where this bites in practice

The failure doesn't show up when you create the proof. It shows up later, when something matters and you go to produce the artifact. By then the file has moved through a storage bucket, maybe a document viewer, maybe a mail client, maybe an export and a re-import. You still have a file called contract.pdf. It still looks right. It hashes to something else.

So the operational rule falls out of the mechanism rather than out of anybody's policy doc. Anchor the bytes you can commit to storing untouched, and store those bytes, not a convenient copy of them. If your workflow has to transform a file, decide which form is the one of record, hash that one, and let the transformed copies be copies. If you can't keep the raw bytes, anchoring a canonical serialization you can regenerate deterministically is a better bet than anchoring whatever came out of the tool that day.

The same thinking applies to anything you want to anchor from an agent session. There's an optional MCP server, pip install verify-proof[mcp], built on FastMCP, exposing five tools so Claude Desktop or Cursor can verify a proof or create one inside a conversation. Convenient, but it doesn't change the rule. Whatever bytes you handed the hash function are the bytes you're on the hook for.

The test, run it today

Pick a file you have a proof for, or any file you'd want one for. Then:

  1. verify-proof hash it and write the digest down.
  2. Push it through your normal path. Upload it and download it. Commit and check out on another machine. Attach it to a ticket and pull it back. Export and re-import. Whatever your real handling actually is.
  3. Hash the result and compare.

If they match, you know your pipeline is byte-preserving and your proofs will survive it. If they don't, you just found out your archive copy and your anchored digest are about different things, which is much better to learn now than when you need the proof.

Second test, while you're in there. Take a proof you were issued by somebody and see whether you can check it with the issuer's service unreachable. If verifying requires their endpoint to answer, the proof is a claim with a dependency on the claimant staying online and willing. Proofs anchored to Polygon and to Bitcoin can be verified from the proof file and the chain, which includes legacy proofs issued by the retired ProofAnchor service, and that's exactly the case worth checking, because the issuer isn't there to ask anymore.

What I don't have a clean answer for

Formats you can't freeze. A record that lives in a database row, a doc that's edited collaboratively, anything where "the file" is a rendering rather than a stored artifact. You can anchor a canonical export, but then your proof is about your canonicalizer, and the canonicalizer becomes a thing you have to version and preserve alongside the proof. I've picked the raw bytes side of that tradeoff, and the cost is real: it pushes the discipline onto whoever stores the file.

If you've dealt with that for live records, I'd like to hear which way you went and what broke.

The tool is verify-proof, free and MIT licensed, published by Fulcrum Enterprises LLC. I built it, I maintain it, and I run it myself. Source is at https://github.com/Fulcrum-Enterprises/verify-proof and the package is on PyPI.

Top comments (0)