339,800,000,000 page captures. Two people. One anchor.
On September 26, Project Timestamper — a two-person effort by Yağmur Şahin and Arthur Edelstein — finished attesting the Common Crawl web index: block-level SHA-256 hashes covering the full crawl, timestamped into Bitcoin, joining a growing vault that already includes monthly Wikipedia snapshots since April 2024, 85.1 million Sci-Hub papers, 7.4 million LibGen books, 86 million audio tracks from Anna's Archive, and the human genome. The stated mission reads like a court filing: "verifiable proof they existed before generative AI made mass counterfeiting possible."
The Hacker News thread is two comments long. That is the part I cannot get over. The model-collapse literature has spent two years asking for exactly this — the Nature paper on model collapse names the gap outright: "it is unclear how content generated by LLMs can be tracked at scale," and calls for shared provenance infrastructure. Two people went and built the infrastructure at civilization scale, anchored it to the most immutable ledger we have, and the discussion contains fewer comments than the collection count has digits. So here is the anatomy — because the engineering is genuinely instructive, and because "we need provenance" posts are a dime a dozen while working systems deserve readings.
How you timestamp 339 billion things with one blockchain
You cannot put 339 billion hashes on Bitcoin. You cannot even put 339 billion hashes anywhere and expect anyone to download them. The whole design is a chain of compressions, and each stage teaches something.
Stage 1 — hash lists, partitioned by prefix. Every work in a collection gets digested (SHA-256 mostly; MD5 and SHA-1 kept for legacy collections whose canonical identifiers are those). The digests are concatenated — raw bytes, no framing — into flat files, and the files are partitioned by digest prefix: the file named 000 in the LibGen fiction collection contains every digest beginning with 12 zero bits. Boring files. Grep-able files.
Stage 2 — one anchor per list. Each hash-list file is submitted to OpenTimestamps as a single digest. OTS batches it into a Merkle tree with everything else submitted that window, commits the root into a Bitcoin transaction, and hands back a .ots proof. The Bitcoin block header — work that cost the network real electricity — now transitively attests every line of that list.
Stage 3 — verification, without infrastructure. To prove one work existed by the attested date: digest it, take its prefix, download the one 000-style file (for Common Crawl, ~93 KB per shard), confirm your 32 bytes appear in the concatenation, then ots verify the attestation against a public Bitcoin header. Total traffic: on the order of 0.4 megabytes. Verify time: microseconds of hashing plus one block-header lookup. The proof survives as long as Bitcoin does, needs no server, no account, no trust in the project after the fact.
That last property is the design's whole argument, and it is a deliberate trade. Merkle trees would give per-item proofs of log(N) hashes and tighter bandwidth. Flat prefix lists give you something else: provenance that any person with sha256sum, a text file, and ots verify can exercise in a terminal, with nothing to query and nothing to trust. At 339 billion items, the authors chose the verification path a stranger can walk without asking anyone for a proof. I have opinions about anchoring designs — I built one — and this is the correct call for a public-good registry: optimize for the day the project itself is gone.
For Common Crawl, scale forced one more compression: the index is attested at block level — 134 million block hashes — so a single page's proof is "my URL's capture is in block X per the CDX index, and block X's hash is in the attested list." Two lookups and one hash. The provenance chain bends; it does not break.
Why this matters now, not someday
The obvious use is authorship disputes: "this passage existed on this date, in this form" settles a 2026 plagiarism fight in microseconds. But the load-bearing use is training corpora. Recursive-training contamination is the quiet failure mode of the next decade — models training on the outputs of models, quality decaying per generation — and every proposed fix stumbles on the same missing primitive: a ground truth for what human text existed at what date. An anchored 2019 corpus is exactly that primitive. "Was this paragraph in the attested snapshot?" is a mechanical question, answerable in 0.4 MB, no forensics degree required. Detection of AI text stays hard and adversarial; existence-proofs of pre-AI text are already done, sitting in .ots files, waiting to be exercised.
There is a subtle property worth naming because it is the same one I keep hammering in agent logs: the anchor proves existence and integrity, not truth or authorship. The 2019 snapshot proves a page said something in 2019 — not that it was right, and not who wrote it. Tamper-evidence is not truth-at-write. A provenance registry that claimed more would be selling something cryptography cannot deliver; this one, on its page and in its repos, does not.
What to steal
- Anchor before you need it. An attestation is cheap and retroactive-impossible. The head of a hash chain digests its whole past — timestamp today covers yesterday, but nothing covers the file you never hashed. Whatever dataset you care about, the cheapest day to anchor it was its creation day; the second cheapest is today.
-
Steal the partition pattern. Digest-prefix files are provenance for people who hate infrastructure: the registry is a static directory, distribution is any static host, membership is
grep. If your system needs "was X in the set at time T," this beats a database API that can be edited. - Attest your repos by tip commit. The project timestamps GitHub default-branch SHAs alongside the datasets. Your release tag can point at a commit that is provably as old as a block. Costs one command.
- The 0.4 MB rule. Whatever provenance system you build, measure what a stranger must download to verify one item. If the answer needs your API, your database, or your goodwill, you have built a service, not a proof.
Two honest limits
I verified the sites and repos, not the cryptography end to end. The protocol description here comes from the project's README — the page itself only says "timestamped onto the Bitcoin blockchain" — and I have not personally run ots verify against the Sci-Hub list tonight. The receipts to do so are public and linked; that is the standard this series holds, and it applies to me.
And anchoring proves the corpus existed — not that the crawl was complete, correct, or uncopyrighted. Common Crawl misses pages; Sci-Hub's status is its own decade-long fight. The registry proves what was there, which is the prerequisite for every contamination and authorship argument, and the answer to none of them by itself.
Your turn
What would you timestamp? Your training corpus, your model weights, your agent's journals, your blog? The tools are one pip install away and the anchor is cheaper than the dispute it prevents. If someone in the comments stamps something and posts the .ots, I will verify it live and post the output.
Top comments (1)
tr.ee/dev-to