I built ProofLedger, which anchors the SHA-256 hash of a file to Polygon and Bitcoin. Almost everything I got wrong while building it was on the verify side, not the anchor side, so that's what this post is about.
Here's the question. When your provenance check goes green in CI, what did it have to reach to get there?
Not "what does it verify." What does it contact. Write down the hostnames. That list is the real trust boundary of your setup, and it's usually shorter than people expect and points at more parties than they expect.
Split the claim into its parts first
A timestamp or attestation check is a few separate assertions stacked into one exit code. Worth pulling them apart, because they have completely different failure modes.
- These bytes hash to H.
- H was committed to something append-only at some point.
- That commitment is real, and I know it's real without the party who issued it telling me so.
Part one is local and cheap. It's just this:
sha256sum dist/app.tar.gz
No network, no auth, no vendor. If the digest doesn't match what the proof says, you're done, the rest doesn't matter. Any verifier that can't do part one without a round trip is doing something strange.
Part two is a data structure question. Either the proof you're holding contains enough material to recompute the commitment, or it doesn't and you have to go ask.
Part three is where the first two stop carrying you. A matching digest and a recomputable commitment both hold even if the commitment exists nowhere but the issuer's database.
"It passes in CI" is not an answer to part three
Green in CI tells you the verifier's dependencies were reachable and returned something the verifier liked. That's it. It doesn't tell you the proof is self-contained, and it doesn't tell you what green would have meant if one of those dependencies had been down.
Two ways that bites.
The first is fail-open. If a verifier treats an unreachable lookup as a soft warning instead of a hard failure, then a green check and a dead endpoint look identical from the outside. You will not notice, because the signal you're watching for is the absence of red. Go read the code path where the HTTP call errors. Does it raise, or does it log and continue? I would check this before I checked anything else.
The second is subtler. Suppose the lookup does work, and the only way to answer part three is an HTTPS call to the issuer's own API that comes back {"valid": true}. That's not verification. That's the issuer's word, restated as JSON, with TLS making you feel better about the transport. If the issuer is also the party whose claim is being checked, you've built a loop. The proof is worth exactly as much as your trust in the company serving that endpoint, which may be plenty, but you should know that's what you bought.
This is why the transparency-log and timestamp-authority conversations in the Sigstore world keep circling back to what a verifier is allowed to require at verify time, and what it does when that thing isn't there. Same question, different vocabulary.
What self-contained actually means
A proof is self-contained when you can get from the file on your disk to the anchored commitment using only material in the proof file, plus one lookup against something public.
If the hash was anchored on its own, that's easy. Hash matches, commitment is the hash, one chain lookup and you're done.
If hashes were batched, you need an inclusion path, and this is the part worth understanding even if you never touch my product, because it's the same mechanic everywhere hashes get batched. The proof hands you the sibling hashes along the path from your leaf to the root, in order. You concatenate and hash, step by step, and you either land on the recorded root or you don't. Nothing in that computation needs a server. The sibling hashes are just bytes. The ordering matters and is part of the proof, and if the format doesn't pin the ordering, the proof is ambiguous and you should say so out loud.
What you cannot derive locally is the root that's actually on chain. That's the one lookup, and it should point at a public chain rather than at the issuer. That's the entire reason I shipped a Python package, verify-proof, that does the check offline and locally, plus a GitHub Action for CI. The public verify URL exists too, and the REST API has a public endpoint with no auth so an auditor or opposing counsel can hit it themselves:
curl "https://proofledger.io/api/v1/verify?hash=$(sha256sum dist/app.tar.gz | cut -d' ' -f1)"
That endpoint is rate limited to 120 requests per hour per IP, which is fine for a human checking a claim and not the thing to build a CI fleet on. The local package is the path for that.
Where this bottoms out, honestly
Offline verification does not mean you conjure chain state out of nothing. It means you can re-check without the issuer. You still have to have obtained the anchored root at some point, from a node you run, a block explorer, or a copy you pinned when you first received the proof. If you never do that, you're trusting whoever handed you the root. Local verification moves the trust from "the vendor's API right now" to "the chain record I can check from any source," which is a real improvement and not a magic one.
Other limits worth stating plainly. A timestamp proves the bytes existed by a certain point. It says nothing about who made them, whether they're correct, or whether the thing they describe is true. Anchoring establishes an upper bound on age and nothing else.
And I'm not claiming my Bitcoin proof is better than a free one. OpenTimestamps gives you free Bitcoin timestamping and the resulting proof is cryptographically just as valid as mine. What I sell is the platform around the proof. The proof itself is the same math available to anyone. On my side, Polygon anchoring is free and unlimited on every plan including the free one, and Bitcoin anchoring is billed per anchor because each one costs me a transaction. I'd rather say that than dress it up.
The test, run it today
Take an artifact you already have provenance for. Not a toy file, a real one from a real build.
Test one. Run your existing verify step with no network at all and record the exit code.
docker run --rm --network none -v "$PWD:/work" -w /work my-ci-image \
./verify.sh dist/app.tar.gz
echo $?
If it passes, ask what it actually checked. If it fails, read the error and find out which of the three parts above it failed on. Either answer teaches you something. The bad outcome is a warning line and a zero exit code.
Test two. Append a byte and verify the tampered copy.
cp dist/app.tar.gz /tmp/tampered.tar.gz
printf '\x00' >> /tmp/tampered.tar.gz
./verify.sh /tmp/tampered.tar.gz; echo $?
A verifier that rejects this offline is doing part one correctly. A verifier that needs the network to notice a flipped byte is worth a closer look.
Test three. Feed your verifier a hash that was never anchored. You want an unambiguous "no record," not an empty success.
Test four. Hand the proof to someone outside your org and ask them to check it without talking to you or to your vendor's support. If they can't, the proof isn't portable, and portability is the only property that matters when the check happens somewhere you aren't.
That last one is the question I'd actually like practitioners to answer: when you hand a provenance artifact to an outside party, what do they have to install or trust before they can check it themselves? I have my answer for my own thing. I'm interested in what other formats require in practice.
If you want to poke at mine, the API docs are at proofledger.io/api.html and the OpenAPI spec is at /openapi.json. verify-proof on PyPI is the part you can run without me.
Top comments (0)