Judge William Alsup ordered Anthropic to pay a $1.5 billion copyright settlement last year, but his written ruling made a more profound point: training AI on copyrighted books is lawful. This isn't the paradox it seems, according to TechCrunch. It's the clearest signal yet of a massive, unresolved legal transfer of creative capital.
The multi-billion dollar penalty was for piracy. The judge found Anthropic liable for downloading books from "illegal online shadow libraries." The fair use finding was separate, declaring the act of training on those same books, once acquired legally, to be “spectacularly transformative” and permissible. This dual outcome reveals the battlefield: the method of taking the data is punishable, but the use of that data to train is, for now, being validated in court. As attorney Cathy Gellis noted, a $1.5 billion fine is a manageable cost to a company projecting $200 billion in annual revenue by 2028, making the fair use precedent the far more valuable corporate asset.
Is Copyright About Reading or Copying?
The legal fight hinges on this old question applied to a new process. AI companies argue their models are “reading” the world's text to learn patterns, much like a human author studies literature. As Gellis explained, “Copyright law hinges on copying, but it doesn't hinge on using the work or experiencing the work, consuming the work, reading the work.”
Judge Alsup embraced this analogy, writing, “Like any reader aspiring to be a writer, Anthropic’s LLMs trained upon works not to race ahead and replicate or supplant them, but to turn a hard corner and create something different.” This framing treats the AI training process as a form of consumption, not reproduction, which pushes it toward the fair use safe harbor. A similar corporate data-hoarding strategy was on display when Amazon bulldozed rare books for its AI-fueled data war.
What Makes a Use "Transformative" Enough?
Fair use is the critical exception. U.S. copyright law allows for the use of copyrighted materials without permission if the use is "transformative" enough. Judges weigh four factors:
- The purpose and character of the use.
- The nature of the copyrighted work.
- The amount and substantiality used.
- The effect on the market for the original.
In the Anthropic ruling, the judge found the purpose, training a generative AI model, to be “exceedingly,” “quintessentially,” and “spectacularly transformative.” He reasoned that because the models weren't designed to output exact copies and did not displace demand for the original books, the use was fair.
“The copies used to train specific LLMs were justified as a fair use,” Alsup wrote. “Every factor but the nature of the copyrighted work favors this result. The technology at issue was among the most transformative many of us will see in our lifetimes.”
But this logic is not absolute. In a separate case cited in the source, Judge Stephanos Bibas ruled against AI research firm Ross Intelligence. Ross had used Thomson Reuters' legal content to build a competing AI platform. The key difference? The output was a direct market substitute. As attorney Jason Henderson summarized, “What’s tending to win is if what you’re doing is you’re training on somebody’s property because your purpose is to directly compete, then the courts will frown on it… If what you’re doing is not going to compete, then the courts are tending to find ways that it will be okay.”
This creates a precarious distinction for authors. While an AI model trained on all novels may not directly compete with a single author's specific book, its ability to generate summaries, adaptations, or stylistic imitations certainly competes with an author's broader professional capacity.
If It's Fair Use, Why the Billion-Dollar Piracy Fine?
Here, the ruling draws a bright, practical line. Converting lawfully purchased print books to digital format for a private training library was deemed fair use, akin to having a “more favorable” digital bookshelf. Downloading pirated copies, however, was “inherently, irredeemably infringing.”
The court rejected the notion that a transformative end goal excuses illicit means: “Anthropic is wrong to suppose that so long as you create an exciting end product, every ‘back-end step, invisible to the public,’ is excused.”
This bifurcation offers a narrow path for rights holders: suing over the source of the data, not the use of it. It turns litigation into a forensic accounting exercise, tracing whether a specific book in a training set of millions was legally acquired. For AI companies, the lesson is operational: build your “central library of all the books in the world” through purchase and scanning, not torrents.
Can a 1976 Law Define an AI Future?
The core tension is temporal. U.S. copyright law was last updated in 1976. Judges are now interpreting its analog-era concepts of "copying" and "transformative use" for a process where a model ingests a trillion words to generate statistical probabilities.
Gellis points out this forces a re-examination of basics: “If you write your novel in Microsoft Word and run spell check, we kind of feel comfortable with the idea of saying that Word does not own your novel. [AI] is forcing us to look at a whole bunch of decisions that we kind of ignored for a while.”
The legal system is playing catch-up through piecemeal litigation. “Everybody is very worried right now because the law is all over the place, and it’s because of this question,” Henderson said. The rulings so far are “initial opening volleys” that shape behavior while higher courts and possibly Congress work toward a final answer.
XOOMAR Analysis | What to Watch
The Anthropic ruling is a landmark, but not the last word. It establishes a powerful pro-training precedent in the crucial Ninth Circuit, but other cases are moving through other courts with different fact patterns, especially around direct market competition. Watch for two developments:
- The Licensing Pivot: The clearest signal will be if major AI companies, seeking legal certainty and ESG cover, begin proactively striking large-scale licensing deals with publishers and author collectives, treating text as a licensed input rather than a free commons.
- The Output Test: As Gellis noted, the authors in the Anthropic case made “no allegations that any of the outputs of the LLM were infringing.” The next wave of lawsuits will almost certainly center on outputs. If an AI model generates content that directly supplants a copyrighted work’s market, the fair use defense weakens dramatically. That's when the battle truly begins.
Impact Analysis
- A precedent is being set that AI model training constitutes 'fair use' of copyrighted materials, changing the legal landscape for all AI companies.
- The ruling draws a critical distinction between legal acquisition of data and illegal piracy, making data sourcing methods the new key legal battleground.
- The financial penalty of $1.5 billion is framed as manageable against Anthropic's massive growth, suggesting fines may not deter corporate data practices if the legal precedent is favorable.
Originally published on XOOMAR. For more news and analysis, visit XOOMAR.
Top comments (0)