There's one question worth asking about any timestamp proof file you're holding right now: when you verify it, what exactly are you comparing your computed root against, and where did that value come from?
Real question, real answer, and it's coming. But it doesn't mean much until you've folded a Merkle path by hand once, so let's do that part first.
What a proof file is asking you to do
The setup is simple. You want evidence that a file existed by a certain time. Hashing it gives you a digest. Writing that digest to a public chain gives you a timestamp. Writing one transaction per file gets expensive, so anchoring services batch: they take a pile of digests, build a Merkle tree over them, and publish the root.
Your digest is a leaf. The root is what actually got committed. The proof file is the receipt that connects the two, and it usually carries the leaf, the sibling hashes on the way up, the root, and something about where that root landed.
Verifying means doing the walk yourself. Hash your file, fold it upward with the siblings, and see whether you arrive at the root.
The leaf:
import hashlib
def leaf_hash(filename):
with open(filename, "rb") as f:
return hashlib.sha256(f.read()).digest()
The fold:
def fold(leaf, path):
node = leaf
for step in path:
sibling = bytes.fromhex(step["hash"])
if step["side"] == "left":
node = hashlib.sha256(sibling + node).digest()
else:
node = hashlib.sha256(node + sibling).digest()
return node.hex()
That's the mechanism. Your format will name those fields something else, and it may encode the side as an index instead of a word, but the shape is that: a loop, a concatenation, a digest.
The part that bites
Look at the branch in that loop. Order of concatenation is the entire content of the algorithm. sibling + node and node + sibling produce completely unrelated digests, and once you've gone the wrong way at one level, every level above it is garbage. The fold doesn't drift off by a little. It lands somewhere else entirely.
Which means there are a few things to pin down in whatever format you're handed, and they are not documented as loudly as you'd like:
Does the format tell you the side, or imply it? Reading an explicit side field and inferring side from index parity are different algorithms. So is sorting the pair before hashing, which some constructions do specifically to avoid carrying side information at all. Only one of those matches the tree that was actually built. Open your own proof file and find out which one it's describing.
Raw bytes or hex text? A digest has a binary form and a printable form, and hashing the printable form is a different operation from hashing the bytes it prints. If the tree was built by concatenating raw digests and your code concatenates hex strings, nothing will line up. The failure mode is nasty because it reads as "this proof is invalid" rather than "your decoder is wrong."
Once or twice? Some Merkle constructions hash each pair a single time, some apply SHA-256 twice at every internal node. Check which one your format means before you conclude anything about the file.
Every one of those is a place where a correct proof looks broken, or a broken proof looks fine because you and the builder disagree about what the bytes mean. This is why I'd rather have code that does the fold in front of me than a green check mark from somewhere else.
Where the obvious answer runs out
So you write the loop, you fold the path, you land exactly on the root in the file. Feels finished.
It isn't. Look at where each input came from.
The leaf came from your file, and that part is load bearing. It binds the proof to the actual bytes on your disk, so if one byte changes, the leaf changes and the fold stops reaching the root. Good.
But the path and the root both came out of the same document. So the thing you just proved is that the proof file is internally consistent with your file. That's a weaker statement than it feels like, because internal consistency is cheap to manufacture. Given any file at all, you can hash it, build a fresh tree that includes it, and emit a perfectly self-consistent proof today. The fold will check out. It says nothing about time, because nothing in it has been published anywhere.
The root is the only value in that file that was ever committed outside the file. That's the part that carries the timestamp. Everything else is derivation.
So the comparison that matters is not leaf-folded-against-root. It's root-against-chain.
What to look for in your own proof file
Open one. Find the root. Then ask whether the file gives you enough to resolve that root somewhere you don't control: a chain, a transaction identifier, a block. Something you could hand to a block explorer or your own node without asking the issuer's permission.
If the root only ever appears inside the file, and the only way to confirm it was ever anchored is to call the issuer's API and get back a boolean, then the trust boundary hasn't moved. You're back to taking their word for it, with extra hashing in the middle. Correct folding code doesn't save you from that, which is the uncomfortable bit. The crypto can be perfect and the chain of trust can still be a loop.
Where verify-proof sits, and what it won't do
verify-proof is a Python CLI, pip install verify-proof. Three subcommands: hash gives you the SHA-256 of a file, verify checks a proof file against that hash and its Merkle path, and create anchors a hash through ProofLedger, added in 0.3.0. It handles proofs anchored to Polygon and to Bitcoin, including legacy proofs from the retired ProofAnchor service. There's an optional MCP server, pip install verify-proof[mcp], built on FastMCP, exposing five tools, so you can verify or create inside a conversation in Claude Desktop or Cursor.
verify recomputes the hash and walks the path itself rather than asking an API whether the proof is good. Only create touches the network, and it sends the hash, never the file.
Now the limit, because it follows directly from that design decision. Since verification does no network calls, the tool cannot go look at the chain for you. It will do the fold and it will do it locally, on a machine with no route out, with no API key and no account. It will not tell you the root was ever anchored. That step stays yours.
I made that trade on purpose and I think it's the right one, but it is a trade, and I'd rather say so than describe an offline verifier as though it settles the on-chain question too.
The test you can run today
Take a proof file you already have, from whoever issued it.
- Cover up the root. Pretend it isn't there.
- Hash your file, then fold it up the path with the sibling hashes, using the loop above. Adjust for whatever your format says about side, encoding, and single or double hashing.
- Uncover the root and compare.
- Then take that computed root and go looking for it on the chain the file names, in an explorer or a node that is not the issuer's.
Read the result:
- The fold doesn't reach the stated root. Either your reading of the format is off, or the file is. Fix the format questions first.
- The fold reaches the root, but you can't find the root on chain. You have a self-consistent document. You don't have a timestamp.
- The fold reaches the root, and the root is sitting in a block. Now you have something, and be precise about what: that digest existed by that block's time. Not that the file is genuine, not that anyone in particular wrote it. Just that those bytes existed by then.
And if step 4 turns into "call our endpoint and we'll confirm it," you've answered the question at the top of this post. The answer just isn't the one you wanted.
verify-proof is free and MIT licensed, published by Fulcrum Enterprises LLC. I built it and I maintain it. Source is at https://github.com/Fulcrum-Enterprises/verify-proof and the package is at pypi.org/project/verify-proof/, if you want something that does the fold in front of you.
Top comments (0)