DEV Community

Cover image for AI Companies Are Destroying Physical Books — And Locking the Knowledge Inside Corporate Servers
Sanjay Singh
Sanjay Singh

Posted on Originally published at zyvop.com

AI Companies Are Destroying Physical Books — And Locking the Knowledge Inside Corporate Servers

There is a phrase that keeps appearing in descriptions of Anthropic's "Project Panama": "don't want anyone to know about this."

That tells you most of what you need to know.

Details of Project Panama emerged through court documents in Anthropic's copyright litigation, which ended in a $1.5 billion settlement covering multiple infringement claims. The physical scanning program, launched in early 2024, involved purchasing large quantities of secondhand books through intermediaries, scanning them to build training data for Claude, and then disposing of the physical copies. Anthropic reportedly spent millions of dollars on the program. Internal communications described the goal as securing data "untouched by machines" — pre-AI text, sourced from physical books that predate the era of synthetic content.

The story got framed, understandably, as destruction. But destruction is the wrong frame. Books are discarded by the publishing industry every day — overstocks, returns, library culls. What matters is what happened after the scanning: the knowledge those books contained moved from distributed physical copies into privately controlled digital infrastructure, under no obligation to remain publicly accessible, verifiable, or preserved in any form the public can use.

That structural shift — not the shredding — is what this piece is about.

TL;DR — Project Panama shows that copyright's fair use doctrine, as interpreted in the Alsup ruling, now rewards scan-and-destroy over scan-and-preserve. The physical books are gone; the scans are private; and the knowledge they contained has moved from a distributed public system into a proprietary archive with no access obligation. For developers building on foundation models, that creates an unverifiable provenance chain that cannot be audited from the outside.


Why Physical Books Instead of Ebooks?

Two reasons: cost and legal exposure.

Ebooks are expensive at licensing scale, and publishers have little incentive to extend favorable terms to AI companies seeking bulk training rights. Secondhand physical books are cheap to acquire in volume. More importantly, the "first sale doctrine" in US copyright law permits the resale and transfer of a legally purchased physical copy. It does not extend to ebooks, which are typically licensed rather than sold outright — so acquiring ebooks for training at scale creates immediate licensing liability that physical copies avoid.

The practical result: companies acquire physical books on the secondhand market, feed them through high-speed sheet-fed scanners (a process that requires cutting the spine to separate pages), and then dispose of the originals. Why the disposal? That requires understanding what the court actually found.


What the Court Actually Found

The Anthropic litigation produced a fair use ruling by Judge William Alsup that is easy to misread, and worth being precise about.

According to court documents and contemporaneous reporting by Ars Technica, Alsup found that Anthropic's scanning of legally purchased physical books constituted fair use — specifically because the process was transformative and because the original physical copy was destroyed rather than retained alongside the digital version. His reasoning, as quoted in reporting: "One replaced the other." The company ended up with one copy, digital rather than physical, not two. He further noted that "there is no evidence that the new, digital copy was shown, shared, or sold outside the company."

This ruling does not hold that companies are legally required to destroy books before they can scan them. What it found is that in this specific situation — legally purchased originals, destroyed after scanning, internal-only use of the digital copy — fair use applied. The destruction was relevant because it eliminated the duplication that would otherwise undermine the fair use argument. Companies that want to rely on similar reasoning have a strong legal incentive to destroy. That is not the same as a statutory mandate.

The contrast with the Internet Archive illustrates why. The IA's "National Emergency Library," opened during COVID-19, allowed multiple simultaneous digital borrowers for books the IA physically owned. Publishers sued, and the IA lost on appeal to the Second Circuit in September 2024 — because it retained both the physical copies and the digital scans and made the digital versions available beyond what the physical holdings could support. That created net new copies. Destroying the physical original collapses the "two copies" problem; keeping both creates it.

The outcome is a legal framework that rewards scan-and-destroy over scan-and-preserve — and produces no public benefit from the scanning.

One further distinction matters: the Anthropic litigation also involved separate claims about pirated digital books acquired through means unrelated to physical purchase. Project Panama — the physical scanning program — is distinct from those piracy allegations. The $1.5 billion settlement covered both. The fair use reasoning around format-shifting applies specifically to the legally purchased physical books, not to the broader infringement claims.


What Is Actually Being Destroyed?

The "rare books" framing that Anna's Archive uses deserves scrutiny before acceptance.

As Ars Technica reported, booksellers involved in the supply chain characterized the inventory as dead stock — out-of-print instruction manuals, vanity press titles, mass-market paperbacks with no active readership, books whose next stop, absent a buyer, was the recycling bin. The book trade routinely discards large quantities of unsellable inventory. The physical destruction here, while worth noting, is not categorically different from what the industry already does at scale.

This matters because the strongest version of the alarm — that irreplaceable, unique texts are being systematically erased — is not well-supported by what is currently documented. No confirmed list of genuinely rare works destroyed through AI scanning programs has been published.

The probabilistic concern is more defensible. If millions of books pass through a scanning operation without expert archival review, the statistical likelihood is that some genuinely uncommon material — local histories with no digital record, foreign-language titles from publishers that no longer exist, ephemera that ended up bound and sold at an estate sale — will pass through unrecognized and be destroyed. The problem is not that AI companies are specifically targeting rare books. It is that at sufficient volume, with no curatorial filter, some will be lost incidentally.

But even setting aside rarity entirely, there is a more important concern: what happens to the content of ordinary books after they are scanned.

The knowledge is captured. Anthropic retains the digital scans — they are not discarding the files, only the paper. But that capture is private. The scans are not released publicly. The trained model cannot reproduce verbatim text from its training data. And the original physical copy — the publicly verifiable artifact that any library, scholar, or journalist could have retrieved — no longer exists. For anyone who wants to know exactly what a model learned from a specific source, there is no path to that knowledge. The book is gone. The scan is proprietary. The model will not reproduce it.


The Race for Pre-AI Training Data

Anna's Archive's post claims that since early 2025, AI-generated content has accounted for more than half of newly published internet content. This figure is theirs and should be read as such — it is not independently verified here, and the exact proportion is contested. But the underlying dynamic they are pointing to is widely discussed across the AI research community: as synthetic text proliferates on the web, models trained on recent web crawls increasingly ingest model-generated output alongside human-written text, with compounding effects on training quality.

This is why physical books have become a target. Text produced before the widespread deployment of large language models is, by definition, not contaminated by model-generated output. Physical books, which exist entirely outside the web's content pipeline, represent a corpus with clean and verifiable provenance. Whether the advantage justifies the approach being described here is a separate question. But the demand for this data is real, and it explains the urgency of programs like Project Panama.

Once those books are scanned and the physical copies destroyed, whatever they contained has been absorbed into private training infrastructure. The originals cannot be rescanned by a different organization, a public library, or a future researcher who wants a different kind of access to the same material.


Anna's Archive's Response

Anna's Archive — a shadow library that operates outside copyright law — has issued a call for volunteers to scan books from libraries, archives, and personal collections. According to their post, contributors receive lifetime membership for small-scale uploads and financial compensation for large-scale scanning operations. Their stated goal: if 10 million volunteers each scan one book, that is 10 million texts preserved in a publicly accessible format rather than absorbed into a private corpus.

Their scanning guides recommend non-destructive methods: overhead camera rigs, flatbed scanners, book cradles that hold the spine open at an angle. None of this requires cutting the book apart.

Anna's Archive has an obvious interest in promoting scanning for their library, and their statistics and projections should be read as advocacy rather than independent research. That does not make the underlying argument wrong. But attribution matters.


The Structural Problem: Concentration of Captured Knowledge

The deeper issue here is not specific to Anthropic, not specific to physical books, and not resolvable by any single regulatory decision. It is structural.

Physical Library System Private AI Training Archive
Access model Open; distributed across thousands of independent institutions Closed; single-organization control
Source verification Any copy can be retrieved and checked by anyone No external access to the training corpus
Redundancy Many independent copies survive damage or disaster Single proprietary archive; no redundancy obligation
Curation transparency Acquisition decisions are publicly documented Inclusion and exclusion choices are invisible to outsiders
Error correction Errors in one copy are detectable against others OCR failures or misattributions are undetectable externally
Preservation mandate Legal obligations apply in many jurisdictions None
Institutional continuity Survives organizational closure through distributed holdings No automatic transfer to public stewardship

Physical libraries distribute copies of knowledge across thousands of institutions. No single organization controls access. Texts can be cross-referenced, cited, physically retrieved, and examined. If one copy is lost or damaged, others survive. This distributed structure is not just logistical — it is a resilience property that scholars, lawyers, journalists, and researchers depend on for verification of sources.

AI training pipelines work differently. A company that scans a large quantity of books builds a centralized, privately controlled digital archive. It is not publicly searchable. It cannot be cited by external researchers in a way that others can verify. It cannot be audited by anyone outside the organization.

If errors were introduced during scanning — OCR failures, pages out of order, text from one book attributed to another — there is no mechanism for outside detection. If curation decisions excluded particular categories of material, that exclusion is invisible. If the company closes, the archive does not automatically transfer to public stewardship.

What copyright litigation has produced, without anyone designing it this way, is a system in which a corporation can: acquire physical books from the public secondhand market; scan them into a private digital archive; destroy the physical originals; train multiple generations of AI models on that archive; retain the digital scans for future use; and operate under no legal obligation to make any of this publicly accessible. This is not a public library. It is a private one, assembled from materials sourced in the public market, under legal cover provided by a fair use doctrine designed for a much narrower purpose.


What This Means for Developers

If you are building on top of foundation models, the training-data provenance problem is upstream of you — but not irrelevant to you.

Consider what it means to verify a model's claim. Standard practice involves tracing a factual assertion back through a chain of sources to a primary document. When the model's knowledge derives from a physical book that no longer exists — and the scan of that book is proprietary — that chain is broken.

The model may have learned from a 1974 manual on soil engineering standards, a 1989 local building code, or a 1962 technical specification. If that text is no longer available anywhere in accessible form, the model's claim on that subject cannot be checked against a surviving source. That is a provenance failure, and it is distinct from the ordinary uncertainty about training data composition.

There is also an auditability problem. Training data curation involves choices: what to include, what to exclude, how to handle OCR errors, how to weight older versus newer material, how to handle books that were factually wrong at the time of publication. With proprietary training corpora assembled from destroyed sources, these choices are invisible.

A model that reproduces an error from a scanned source cannot be easily corrected without identifying the error — and identifying it requires either catching the model in a mistake or having access to the underlying source, which may no longer exist in any publicly retrievable form.

Monoculture risk follows from this. If a small number of AI companies control the digitized versions of materials that no longer exist physically, and if those companies share overlapping supply chains and vendor relationships for their physical acquisitions, they may share systematic blind spots or processing errors. The diversity of access that a distributed physical library system provided — many organizations, many copying decisions, many preservation standards — is reduced when the effective copy count drops to a handful of proprietary archives.

None of this is hypothetical. It is the logical consequence of a legal and economic structure that incentivizes centralized private capture over distributed public preservation.


The Practical Case for Preservation

The tools for non-destructive book scanning are accessible. A document scanner suited for book digitization costs under $200; camera rigs with angled book cradles are available for under $100. The core software stack — ScanTailor for page correction, Tesseract for OCR, OCRmyPDF for output packaging — is free and open source. The technical barrier is low.

Organizations that accept public contributions are active: the Internet Archive, Project Gutenberg, the Open Library, and Anna's Archive each accept uploads, with differing legal frameworks. Project Gutenberg and the Open Library focus on public domain and openly licensed material. The Internet Archive has broader scope. Anna's Archive operates outside copyright restrictions and should be understood as such.

If you have access to books that are not currently in any public digital collection — personal libraries, institutional archives, local histories, materials in less-resourced languages — the time to digitize them is now. The secondhand book market is being absorbed by buyers with better logistics, larger budgets, and no interest in public access. That process is ongoing.


Frequently Asked Questions

Was Anthropic's physical book scanning legal?

For the Project Panama physical scanning program specifically: yes, as the Alsup ruling found. Purchasing books legally, scanning them for internal use, and destroying the physical originals constituted fair use under the analysis. The $1.5 billion settlement covered the separate piracy claims — use of books from Library Genesis and similar sources — which Anthropic did not successfully defend.

Why did Anthropic destroy the physical copies?

The destruction was legally consequential, not incidental. Alsup's reasoning explicitly turned on the company ending up with one copy — digital rather than physical — rather than two. Retaining the originals alongside the scans would have created the duplication problem that undercuts fair use. Companies following this approach have a structural legal incentive to destroy.

How is this different from what the Internet Archive did?

The Internet Archive's controlled digital lending program allowed simultaneous borrowers to check out digital copies of books the IA owned physically — creating more effective copies than the physical holdings could support. The Second Circuit found this to be copyright infringement in September 2024. Anthropic destroyed the physical originals, eliminating the net-new-copy problem. The legal distinction is real; the public benefit difference is the point of this piece.

Is Project Panama different from the piracy allegations in the same lawsuit?

Yes. Project Panama involved legally purchased physical books scanned under the fair use framework Alsup described. The separate infringement claims — which drove most of the settlement value — concerned pirated digital books acquired through Library Genesis and similar sources. These are distinct fact patterns with distinct legal treatment.

What does this mean practically for developers building on Claude or similar models?

Primarily, it means provenance chains for model claims are unauditable at the source level for any content that originated from this corpus. If a model returns information derived from a physical book that was scanned and destroyed, there is no surviving primary source to check the claim against. Standard retrieval-augmented generation architectures partially mitigate this by grounding responses in retrievable documents — but they cannot repair the gap for training knowledge that has no surviving accessible form.


Conclusion

Whether a particular book was going to be discarded anyway is not the right question. Books have always been discarded. The right question is what happens to the knowledge they contained after a scanning operation — who can access it, who can verify it, who can build on it, and what happens to it if the organization holding it closes, makes errors, or simply decides not to share.

For most of recorded history, the answer to those questions has been determined by the distributed nature of physical copies. Many institutions, many copies, many independent preservation decisions. That structure created redundancy, auditability, and resilience.

What AI training pipelines are building — at scale, under current legal conditions — is different in kind. The scanning is real. The scans are retained. But the knowledge those scans represent has moved from a distributed public system into a set of private archives subject to no preservation mandate, no public access requirement, and no external audit mechanism.

That outcome is not the result of anyone choosing it deliberately. It is the accidental product of copyright litigation decided without reference to preservation interests. Naming it accurately is a precondition for doing anything about it.


Sources

  1. Anthropic destroyed millions of print books to build its AI models — Ars Technica, Benj Edwards, June 2025

  2. Court opinion and related filings, Bartz v. Anthropic (Judge William Alsup, N.D. Cal.) — fair use ruling on the Project Panama physical scanning program

  3. Hachette Book Group v. Internet Archive, No. 23-1260 (2d Cir. Sept. 4, 2024) — Second Circuit affirming controlled digital lending as copyright infringement

  4. Anna's Archive blog — volunteer scanning call-to-action and public preservation program

  5. Hacker News discussion thread on Project Panama (search "Project Panama Anthropic books")


Originally published on ZyVOP

💡 For more articles like this, subscribe to the ZyVOP newsletter!

Top comments (0)