DEV Community

Craig Solomon
Craig Solomon

Posted on

What crosses the network when you timestamp a file, and how to check it yourself

Here's the question worth asking about your own provenance setup before any of the fun parts: at verification time, what has to be up?

Not at build time. At verification time. Months from now, when someone asks you to prove a file existed on the date you say it did, which machines, which services, which credentials have to be alive and reachable for you to answer?

It's an easy question to skip, because the issuing side is the part you actually test. You wire up the attestation step, it goes green in CI, you move on.

So let's work it through properly. There are two phases and they have completely different dependency profiles.

Phase one: issuing has to touch the network, and that's fine

If you want a digest anchored to a chain, something has to put it there. There's no clever way around that. You hash the file locally and then you hand that hash to a service that batches it into a Merkle tree and anchors the root.

In my tool that's one subcommand:

pip install verify-proof

verify-proof hash ... # SHA-256 a file
verify-proof verify ... # check a proof against the hash and its Merkle path
verify-proof create ... # anchor a hash (added in 0.3.0)
Enter fullscreen mode Exit fullscreen mode

Three subcommands. Exactly one of them opens a socket: create. It needs a free ProofLedger API key, and it sends the hash, never the file.

Now, the obvious reading of that last sentence is "good, nothing leaks, we're done on the egress question."

That reading is not enough, and this is the part of the post I actually care about.

A digest is a commitment, not a blindfold

Sending a hash instead of a file is not the same as sending nothing. A SHA-256 digest is a commitment to the content. Anyone holding it can't read your file, but they can confirm a guess at it. Feed in a candidate, hash it, compare.

Whether that matters depends entirely on how guessable the content is. A multi-megabyte binary with embedded build IDs, fine, nobody is guessing that. A one-line config value. A short contract built from a template where the only variable fields are a name and a number. A CSV row. Those are a different story, and the hash of one of them is closer to a lookup key than to a secret.

So the protection in a hash-only transport is not "a hash is unreadable." It's "the file never moved off the machine." That's a real property and it's worth stating in those terms, because stated correctly you can see what it does and does not buy you, and you can decide for yourself whether the digest of this particular artifact is something you're comfortable handing to anyone.

If it isn't, salt the thing before you hash it. Hash a file that includes a random value you keep alongside the proof. Then the digest is no longer a lookup key for a guessable document, and you've traded it for one more piece of state you have to not lose. That's the tradeoff, and it's yours to make rather than mine.

Phase two: verifying, and what "local" actually has to mean

Here's where the dependency question gets sharp.

Verification is hashing and more hashing. Recompute SHA-256 over the file, then fold your leaf up the Merkle path, concatenating and hashing at each level until you have a root. That's all arithmetic. No network required.

Except the root you just computed is worthless on its own. You have to compare it against the root that's actually anchored on chain, and that value cannot come from the proof file, because the proof file is the thing you're checking. So it comes from a block explorer, from a node you run, or from a copy of the root you captured and trusted at the time.

Which means "offline" was never the right property to ask for. The right property is narrower and more useful: verification must not depend on the issuer. A chain lookup is a dependency you can satisfy many ways, including with hardware you control. A call to api.vendor.com/verify is a dependency with exactly one supplier, and that supplier is the party whose claim you're trying to check.

That's the whole reason I wrote the verify path to recompute and fold locally rather than ask an API whether the proof is good. It needs no key and no account. The asymmetry is deliberate: creating a proof costs you a credential and a network call, checking one costs you neither. That also means it can still check proofs issued by ProofAnchor, a service that's retired and isn't coming back. The anchors are on Polygon and Bitcoin, and both of those will answer a question about an old block whether or not the issuer still exists.

Same split holds if you run it as an MCP server, which is optional (pip install verify-proof[mcp], built on FastMCP) and exposes five tools so Claude Desktop or Cursor can do this inside a conversation. The create tool still wants a key and still talks out. The verify side still doesn't.

The test, which you can run today against whatever you're using now

Forget my tool for a second. Here's how to find out what your current setup depends on. Takes one sitting.

Kill the issuer, keep the chain. Block egress to your vendor's API hostname at the firewall and leave the rest of your network up. Then verify a proof you already hold. If it passes, your verify path is talking to a chain and doing its own math. If it fails, you've learned that your ability to substantiate your own claim is a function of that company's uptime and your account's standing.

Pull the credential. Unset every API key and token in the environment and verify again. A verifier that needs a credential is a verifier that can be revoked.

Watch the issuing call. Compute the digest yourself first with sha256sum, then capture the traffic while the attestation step runs and search the capture:

sha256sum contract.pdf # note the hex digest

# run your create/attest step with a capture or proxy in front of it, then:
grep -c '<that digest>' capture.txt # expect a hit
grep -c '<a string that exists only in the file>' capture.txt # expect zero
Enter fullscreen mode Exit fullscreen mode

If the second grep returns anything above zero, your file body is going over the wire and you should know that before the next audit, not during it.

Check what the root is compared against. Read your proof file. If the anchored root is sitting in it and your verifier compares the computed root to that field, the check is internally consistent and externally meaningless.

One honest limit on all of this: none of it tells you a file is genuine, or that a particular person wrote it. It tells you a specific sequence of bytes existed at a specific time. That's narrow, and it's often exactly the thing in dispute, but don't let it grow in the retelling.

verify-proof is MIT licensed and free, on PyPI and at https://github.com/Fulcrum-Enterprises/verify-proof. I built it and I run it.

The part I'd genuinely like other people's answers on: for long-lived artifacts, do you store the anchored root yourself at issue time, or do you plan on querying the chain whenever the question comes up? I went with the second and I keep wondering if the first is the adult decision.


Disclosure: this article was drafted by an AI agent I built and run, from facts I supplied about my own project.

Top comments (0)