DEV Community

The AI Prism
The AI Prism

Posted on • Originally published at theaiprism.com

AI Companies Are Shredding Rare Books — And That Changes Everything About Training Data

Originally published on The AI Prism


The Quietest Crime in AI

There’s a story making the rounds on Hacker News this week — 794 upvotes, 514 comments, and climbing — that reveals something deeply unsettling about how AI companies are building their training datasets. It’s not about copyright, not exactly. It’s not about privacy, not directly. It’s about something older and more visceral: the physical destruction of rare books.

According to reports, AI companies — or middlemen working on their behalf — have been buying up rare and out-of-print books, scanning them page by page to feed into training pipelines, and then shredding the physical copies.

Not donating them. Not returning them. Shredding them.

The logic is cold and calculated: if only one digital copy exists, there’s no question about whether the training data was obtained legally. No second copy to create ambiguity. No original left to dispute ownership. The physical book becomes a liability. The digital scan becomes the only evidence that the book ever existed in this context.

Welcome to the Bookshelf Fire of 2026.

How We Got Here: A Chain of Bad Incentives

This didn’t happen in a vacuum. To understand why rare books are being destroyed, you have to follow the chain of perverse incentives that led here.

Step one: Publishers sued AI companies for training on shadow library data — datasets like Books3, built from pirated copies. The lawsuits were aggressive and high-profile. Authors Guild. The New York Times. Individual writers who found their copyrighted work in training sets.

Step two: AI companies got the message. They stopped relying on shadow libraries and started acquiring content through more defensible channels. Deals with publishers (like OpenAI’s partnerships with Axel Springer, Dotdash, and the Financial Times). Licensed data agreements. And — for the niche, out-of-print, orphaned works that no publisher could license — they started buying physical copies.

Step three: A legal theory emerged that scanning a physical book for training purposes, then destroying the original, creates a defensible fair use claim. The argument goes: if the physical copy no longer exists, there’s no market harm because there’s no competing product. A federal judge has reportedly cited this reasoning in at least one ruling.

The result? A booming underground market for rare books — not for collectors, but for paper shredders.

The Scale Is Larger Than You Think

We’re not talking about a few dozen first editions. We’re talking about millions of books. The Tom’s Hardware investigation (currently the only mainstream tech outlet covering this) reports that AI companies have outsourced book acquisition to middlemen who scour library sales, estate sales, university discards, and used bookstores.

These middlemen operate under nondisclosure agreements so tight they can’t even acknowledge which AI company they’re working for. Books are purchased in bulk, shipped to undisclosed scanning facilities, digitized at industrial scale — and then destroyed.

The economics work because training data scarcity is the bottleneck. The frontier AI labs have already scraped most of the public web. Reddit. Wikipedia. GitHub. Common Crawl. Stack Overflow. The low-hanging fruit is gone. To train the next generation of models — models that reason more deeply, write more naturally, and generalize better — they need higher-quality data. And there’s no higher-quality text than published books.

According to estimates from the AI data market, high-quality text datasets are now trading at $5-15 per million tokens, up from pennies just two years ago. A single rare book that cost $200 to acquire might yield 500,000 tokens of training data — making the economics comparable to licensed data deals, with the added benefit of legal defensibility through destruction.

The Legal Loophole That Makes This Possible

The fair use argument here rests on a specific interpretation that’s both clever and deeply troubling: if you destroy the original after digitizing it, you’ve eliminated any potential market harm.

Here’s the reasoning:

Fair use analysis considers four factors:

The purpose and character of the use — Transformative? Training AI is increasingly considered transformative.

The nature of the copyrighted work — Published works get less protection than unpublished ones. These books are published.

The amount used — Entire books are scanned, which cuts against fair use.

The effect on the potential market — This is the key one. If the physical book is destroyed, there’s no market for it. No competing digital version. No lost sales.

By eliminating factor four, the AI companies create a scenario where the balance tips toward fair use — even though an entire book was copied and no one will ever be able to read the original again.

The HN discussion on this thread is worth reading in full. 514 comments and counting, with opinions ranging from “this is a tragedy for human knowledge” to “old books that nobody was reading anyway are being put to better use.” One commenter who claims to work in the industry writes: “The publishers sued AI companies for training on shadow library data, hoping to negotiate content deals for big money down the line. Instead, they got a world where the books are just destroyed. Congrats, publishers. You played yourself.”

What’s Actually Being Lost?

This is where the story gets personal for anyone who cares about knowledge preservation.

The books being targeted aren’t bestsellers. They’re not even in print anymore. They’re the orphaned works — niche academic monographs, out-of-print technical manuals, regional histories, botanical texts from the 18th century, poetry collections from small presses that went under decades ago.

These are exactly the books that libraries struggle to preserve because they have no commercial value. They’re the books that a single copy might exist in a university library’s special collections — or, increasingly, they don’t exist in any library at all, because the AI middlemen got there first.

When an AI company buys and shreds a rare book, that knowledge isn’t lost — but it’s no longer publicly accessible. It’s locked inside a proprietary training set, accessible only through a model’s API. You can’t browse it. You can’t cite it. You can’t rediscover it. You can only ask the model to summarize it for you — assuming the model retained that particular piece of information and didn’t compress it into a statistical pattern.

This is the opposite of what libraries do. Libraries preserve. They share. They make knowledge available across generations. The AI shredding pipeline takes knowledge that was barely surviving and converts it into a non-renewable resource for private models.

The Internet Archive Connection

You can’t tell this story without talking about the Internet Archive.

The Archive’s Open Library project — which scanned physical books and lent them digitally — was sued by publishers in a case that went all the way to the Second Circuit. The ruling was a disaster for digital preservation: the court found that Internet Archive’s controlled digital lending was not fair use, because it created a competing digital market.

The irony is staggering. A nonprofit library that scanned books to make them available to the public lost its fair use defense. Meanwhile, for-profit AI companies are scanning and destroying books behind NDAs, and the legal system appears to be giving them cover through the same fair use framework that the Archive was denied.

One HN commenter put it bluntly: “That’s why archive.org should have never been sued for lending books they had a physical copy of. This is the result. Publishers should be more careful about what they wish for.”

The Data Scarcity Crisis Nobody’s Talking About

Underneath the moral panic about book destruction is a more fundamental story about AI development: we’re running out of data.

The Epoch AI research group has been tracking this for years. Their latest estimates suggest that high-quality text data will be exhausted by 2026-2028. The frontier labs have already consumed most of the internet. What’s left is either low-quality (social media noise), locked behind paywalls (scientific papers, news archives), or physical (books that were never digitized).

The book-shredding pipeline is a direct response to this scarcity. AI companies aren’t destroying books because they’re evil. They’re doing it because they’ve exhausted every other source of high-quality text and the pressure to build better models is relentless.

This is what happens when you tell an industry “build something better” without asking “at what cost?”

The Ghost in the Scanner: How the Pipeline Actually Works

To understand this story, you need to follow the chain from a dusty library sale to a neural network weight file. It’s a supply chain designed for maximum opacity.

Step one: Acquisition. Middlemen — often operating as shell LLCs with innocuous names — attend estate sales, university library discards, and used bookstore closures. They bid in bulk, often paying above market rate. A university library that normally sells discard books for $1-5 each might get offers of $10-20 from these buyers. The sellers rarely ask questions.

Step two: Triage. Books are sorted by perceived value for AI training. Technical manuals, academic monographs, specialized reference works, and out-of-print literary fiction rank highest. Mass-market bestsellers and books already widely available in digital form are lower priority. The selected books go to scanning facilities. The rest — and many valuable books are likely missorted — go straight to shredding.

Step three: Industrial scanning. These aren’t the flatbed scanners you remember from the library. Industrial book scanners can process 1,000-3,000 pages per hour. They use overhead cameras with page-turning robots, capturing both pages simultaneously at 300-600 DPI. A single facility can digitize an entire library of 50,000 books in under a month.

Step four: OCR and processing. The scanned images run through OCR pipelines, cleaned, formatted, and fed into training datasets. Any metadata — author, publisher, ISBN, publication date — is stripped. The goal is clean text, not bibliographic context.

Step five: Destruction. Industrial shredders reduce the physical books to pulp. The paper waste is typically recycled, closing the loop with grim efficiency. No trace remains except the digital copy — owned by a company you’ll never name, for a purpose you can’t verify, used to train a model you’ll only ever access through an API.

The entire operation is designed to be invisible. No logos. No public records. No paper trail from book to training set.

The Open Source AI Angle: Why This Hurts Small Players Most

Consider this from the perspective of someone building an open source LLM. They can’t afford to buy and shred rare books. They can’t even afford to license high-quality training data at $5-15 per million tokens. They rely on what’s publicly available: Common Crawl, Wikipedia, Project Gutenberg, the Internet Archive.

But the Internet Archive is under legal siege. Common Crawl is getting filtered and reduced as websites block crawlers. Project Gutenberg only has books that are in the public domain — which means most 20th and 21st century knowledge is off-limits.

The companies with the most money get access to the best data, and they make sure no one else can get it by destroying the originals. This isn’t about building better AI. It’s about building a moat around the training data supply.

Open source AI advocates have been warning about this for years. The cost of training data, they argued, would eventually become the real barrier to entry — not compute, not talent, not algorithms. We’re watching that prediction come true in real time, one shredded book at a time.

What the Data Tells Us

The economics of this practice are surprisingly transparent, even if the operations aren’t.

Cost of a rare or out-of-print book: $5-200 (bulk purchase discounts apply heavily at scale)

Scanning cost per book: $1-3 (industrial-scale scanning is cheap)

Tokens per book: 50,000-500,000 depending on length

Effective cost per million tokens: $2-40, compared to $5-15 for licensed data from publishers

The economics work. But the hidden cost is the destruction of cultural heritage. A book that’s scanned and shredded contributes to exactly one training run, while a book that’s scanned and preserved in a library could contribute to a thousand — for open source projects, for researchers, for historians, for curious readers.

The AI industry is making a choice here, and it’s the wrong one. The choice isn’t between training data and book preservation. It’s between private, ephemeral access and public, permanent access. The technology to scan without destroying has existed for decades. The choice to destroy is purely legal, not technical.

What Comes Next?

Several things need to happen — and fast.

First, the scanning needs to be separated from the destruction. There’s no technical reason why a scanned book can’t be donated to a library or returned to the seller after digitization. The destruction is purely a legal strategy, not an operational necessity. Laws requiring digitized books to be preserved in public archives after scanning for AI training would close the loophole without banning the practice entirely.

Second, the fair use theory needs scrutiny. A court should examine whether destroying the original actually strengthens a fair use claim, or whether this is an end-run around copyright law that judges never anticipated. The US Copyright Office has already been asked to weigh in on AI training data issues — this should be on their agenda.

Third, we need alternatives to secret training sets. The data scarcity problem isn’t going away. If anything, it’s going to get worse as more AI companies compete for the same finite pool of high-quality text. Open source projects like Common Corpus and the Open Library’s digitized collections show that shared data pools are possible — but they need funding, legal protection, and industry participation.

Fourth, libraries need protection. If AI companies are outbidding libraries for rare books at estate sales and university discards, the result isn’t just fewer books on shelves — it’s fewer books full stop. Libraries should have a right of first refusal on any book that’s being acquired for AI training purposes.

Several things need to happen — and fast.

First, the scanning needs to be separated from the destruction. There’s no technical reason why a scanned book can’t be donated to a library or returned to the seller after digitization. The destruction is purely a legal strategy, not an operational necessity. Laws requiring digitized books to be preserved in public archives after scanning for AI training would close the loophole without banning the practice entirely.

Second, the fair use theory needs scrutiny. A court should examine whether destroying the original actually strengthens a fair use claim, or whether this is an end-run around copyright law that judges never anticipated. The US Copyright Office has already been asked to weigh in on AI training data issues — this should be on their agenda.

Third, we need alternatives to secret training sets. The data scarcity problem isn’t going away. If anything, it’s going to get worse as more AI companies compete for the same finite pool of high-quality text. Open source projects like Common Corpus and the Open Library’s digitized collections show that shared data pools are possible — but they need funding, legal protection, and industry participation.

Fourth, libraries need protection. If AI companies are outbidding libraries for rare books at estate sales and university discards, the result isn’t just fewer books on shelves — it’s fewer books full stop. Libraries should have a right of first refusal on any book that’s being acquired for AI training purposes.

The Bottom Line

The AI industry has a data problem, and it’s solving it by making the problem invisible.

Shredding rare books doesn’t just erase physical objects. It erases the ability of future readers, researchers, and competing AI builders to access the same knowledge. The books that go into the shredder today won’t be available for the next generation of models — unless those models are built by the same companies that destroyed the originals.

We’re creating a world where the past is owned by whoever could afford to digitize and destroy it.

That’s not progress. That’s a private flame for a public library.

The next time an AI company announces a breakthrough model and credits its “high-quality proprietary training data,” ask yourself: what books died to make that possible? And will anyone be able to read them again?

The post AI Companies Are Shredding Rare Books — And That Changes Everything About Training Data appeared first on The AI Prism.


Cross-posted from theaiprism.com — Cutting Through the AI Noise 🧊

Top comments (0)