DEV Community

Cover image for The Great Book Burner: Why AI is Eating the Library
Deep Saged
Deep Saged

Posted on Originally published at deepsage.com

The Great Book Burner: Why AI is Eating the Library

The New Fuel Source

I’ve always believed that every sufficiently complex problem can be solved with enough gears, pulleys, and a very large supply of raw material. If you want to build a brain out of math, you need data. And if you run out of digital data, you simply go to the nearest source of high-density information: the used bookstore.

Lately, there’s been a peculiar uptick in the way some AI firms are sourcing their 'fuel.' It turns out, the most efficient way to feed a Large Language Model isn' quite the delicate process of careful archiving. Instead, it looks a lot like a paper shredder with a scanner attached.

Recently, reports emerged that companies like Anthropic have been engaged in what's known as 'destructive scanning.' It's a lovely term, isn't it? It sounds like something a particularly efficient villain in a low-budget sci-fi movie would do. The process is straightforward: buy a massive amount of physical books—often academic or non-fiction titles from the 70s—strip the bindings, run the pages through a high-speed industrial scanner, and then... well, recycle the remains.

I thought about building a machine to automate this, but the logistics of disposing of several million paperbacks would be a nightmare for the neighbors. I'll leave that to the professionals.

From Preservation to Depletion

For a long time, the industry standard for digitizing books was modeled after the Google Books project. That was a 'non-destructive' approach. You borrow a book, you scan it with a specialized camera, and you put it back. It was a model of digital preservation. It was about making the library bigger by making it accessible.

But we are seeing a fundamental shift in the engineering logic. We are moving from a model of preservation to one of resource depletion.

In the old model, the book was a vessel for information. In the new model, the book is just a consumable fuel cell. Once the statistical relationships between the words have been extracted and converted into weights in a neural network, the physical paper has served its purpose. It’s a bit like how I once tried to build a self-sustaining toast maker; I realized quite quickly that if I kept burning the bread to power the heating element, I was eventually going to run out of breakfast.

This isn't just a minor technical tweak; it's a change in how we view the medium. The physical book is being treated as a mineable ore. We aren't building a digital library; we're running a digital refinery.

The Fair Use Loophole

Now, if you were to present this plan to Halvorsen in safety review, he’d likely point out that destroying property is generally frowned upon. However, the legal engineers have found a very clever workaround: the 'transformative fair use' doctrine.

In a recent ruling involving Anthropic, a judge in California essentially looked at this massive operation and called it 'spectacularly transformative.' The logic is that the AI isn't 'reproducing' the book; it's analyzing the statistical patterns within it to create something entirely new. The court even compared the process to human learning. It’s a very tidy way to frame it.

As long as the company has legally purchased the physical books first, the destruction of the original copies is seen as a way of 'conserving space' through format conversion. It’s a brilliant bit of legal engineering. If you buy a piece of coal, burn it to generate electricity, and then complain that the coal is gone, the law says you're simply being difficult.

I did try to apply this logic to my automated sandwich press—arguing that turning bread into crumbs is 'highly transformative'—but the investors weren't impressed by the lack of actual sandwiches.


Originally published on DeepSage.

Top comments (0)