OpenAI has spent three years telling courts that training AI on copyrighted works is "transformative" â a new use that creates something fundamentally different from the original. That is the heart of fair use, and it is the legal theory that every foundation model depends on.
đ Read the full version with charts and embedded sources on ComputeLeap â
Then a federal court unsealed the emails.
In mid-September 2026, previously redacted filings in multiple copyright lawsuits against OpenAI and Microsoft became public. The documents paint a picture that is difficult to reconcile with the companies' courtroom posture. Microsoft Director of Applied Science Brent Hecht called the ingestion of copyrighted works "an astonishing theft of unprecedented proportions" and potentially "the largest theft of labor in human history." OpenAI's head of ChatGPT Nick Turley wrote that the company's products "are largely substitutive, period" and will only become more so. An internal note from August 2019 reads simply: "We trained GPT-3 on pirated stuff! No sharing that!"
âšī¸ "It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business." â Internal Microsoft document, unsealed September 2026
These are not the words of companies confident in their fair use defense. They are the words of companies that know â and have known for years â that what they built rests on a legal theory that their own employees do not believe.
The community reaction has been explosive. On Reddit's r/technology, the story hit 14,200 upvotes and 960 comments within 48 hours.
View original post on Reddit â
What the Unsealed Documents Reveal
The filings come from two parallel cases: the New York Times v. OpenAI and Microsoft (filed late 2023) and the Authors Guild v. OpenAI case involving 17 named authors including George R.R. Martin and John Grisham. Both sets of documents were ordered unsealed by the Southern District of New York, and both contain internal communications that the companies fought to keep secret.
The "Doom Loop" Memo
Brent Hecht's January 2024 internal presentation did not just use inflammatory language. It described a specific economic mechanism: Copilot was causing New York Times click-through rates to drop by up to 93% compared to traditional Bing search. Hecht warned this would create a "doom loop" â as AI substitutes for the original content, publishers lose traffic, revenue drops, content production declines, and the AI's training data degrades. The companies are, in Hecht's framing, sawing off the branch they sit on.
The LibGen Connection
The Authors Guild filing reveals that OpenAI knowingly trained GPT-3 on books scraped from Library Genesis (LibGen), a well-known piracy site. The company's research leadership was explicit about this:
- Dario Amodei (then OpenAI Research Director): "as a training set [LibGen is] a bit sketchier"
- Sam McCandlish (OpenAI Researcher): "I was just worried about optics â i.e. 'openai uses copyrighted data from sketchy russian website' showing up on [Hacker News] would be unfortunate"
- Ben Mann (OpenAI): Repeatedly described LibGen as "sketchy" in internal notes
Rather than acknowledge the source, the team published the GPT-3 paper using the deliberately vague labels "Books1" and "Books2" to obscure LibGen's role. When scrutiny increased, OpenAI VP of Research Bob McGrew launched "Project Clear" in June 2022 to quietly delete LibGen files: "Given how much OpenAI is in the news, now is the right time to excise Libgen from our systems and storage."
As a deep-dive analysis on Substack documented, the internal conversations show a pattern of knowing the source was problematic and choosing to obscure rather than fix it.
View original post on Substack â
The Substitution Confession
Fair use has four factors, and the most dangerous one for AI companies is market substitution â whether the new use competes with or replaces the original. OpenAI's own executives handed plaintiffs the evidence on this factor:
- Nick Turley (Head of ChatGPT): Products are "largely substitutive, period" and will become even more so
- Greg Brockman (OpenAI President): Called models "excellent at news" and responded "ah nice" to a researcher's report about circumventing the NYT paywall
- Jack Clark (OpenAI Policy Director, May 2020): "The better we do on GPT-X, the more worried genre fiction authors will become about us substituting for them on Amazon"
â ī¸ The one person who saw the full picture and said it plainly: OpenAI's own policy director Jack Clark wrote in May 2020 that "Our work on AI and Creativity is going to increasingly lead to us creating systems that substitute for the labor of people," adding: "Our work in this area will make people unemployed." He also predicted: "There will be a point where a bunch of artists express worry about what we're doing here and we'll likely ignore their concerns." Clark left OpenAI that same year.
The Scale of the Taking
The filings quantify what was scraped:
- 91,692+ copies of works from NYT, Daily News, and Center for Investigative Reporting in mid-training datasets
- 2+ million documents from nytimes.com in Common Crawl-derived datasets
- 160,903+ unique news publisher works in "Project Mango" dataset
- Copyright notices were deliberately stripped from training data
- Paywalled content was scraped while bypassing detection
The Government Weighs In â With a Conflict of Interest
Two weeks before these briefs were unsealed, the DOJ filed a 20-page statement of interest arguing that AI training constitutes fair use. Associate Attorney General Stanley Woodward announced the filing, framing it as a national security imperative: "Rules of law that make it significantly more difficult to develop a robust AI industry in the United States therefore threaten national security."
The brief distinguishes between training (which copies entire works but restricts public access) and outputs (which are publicly accessible but allegedly lack "substantial similarity" to source material). It connects the copyright question to President Trump's broader AI dominance policy.
What the brief does not mention: the federal government was simultaneously negotiating to acquire a 5% equity stake in OpenAI, worth approximately $42.6 billion. As legal publication Above the Law pointed out, the DOJ was filing to protect the legal interests of a company it was about to take ownership in â without disclosing that interest to the court.
What the Community Is Saying
On Hacker News, the TechCrunch version of the story pulled 955 points and 841 comments â one of the most active copyright threads in HN history. The debate crystallized around whether training an AI model on copyrighted content is fundamentally different from a human reading the same content.
User haritha-j captured the majority view: "One human learning from resources and being added to the labour pool is not the same as an infinitely copyable entity doing the same thing." User coffeefirst argued for mechanical royalty systems similar to what the music industry uses. Others drew analogies to surveillance: user ryandrake wrote, "It's perfectly OK for a police officer to observe a street corner... therefore, building a panopticon surveillance system is perfectly OK, too" â arguing that scale fundamentally changes the character of an act.
The contrarian camp pointed out that Judge William Alsup had already ruled (in the Anthropic/Claude case) that training on copyrighted books was "exceedingly transformative and was a fair use." If courts have already said yes, these emails are interesting color but not legally determinative.
â ī¸ Contrarian Corner: Techdirt's counter-argument has teeth: copyright infringement is not theft, period. The Supreme Court said so in Dowling v. U.S.: "Interference with copyright does not easily equate with theft... The infringer does not assume physical control over the copyright." A director-level employee's internal rhetoric does not change the four-factor legal test. And there is real irony here â Microsoft itself spent decades equating piracy with theft in its own anti-piracy campaigns. Now that language is being weaponized against them, but the underlying legal analysis has not changed. Courts apply statutes, not email vibes. Full Techdirt analysis â
Why This Is Legally Devastating (Despite the Counter-Arguments)
The Techdirt argument is legally correct on the narrow point. Copyright infringement and theft are different causes of action. Judges do apply the four-factor test, not rhetorical characterizations.
But here is what the contrarians miss: the four-factor test is exactly where these emails land hardest.
Factor 1 â Purpose and Character of Use (Transformative?): OpenAI's entire defense rests on arguing that training is "transformative" â it creates something fundamentally new. But when your head of product writes that the output is "largely substitutive, period," you have conceded the opposite of transformative. You have conceded replacement.
Factor 4 â Effect on the Market: This is historically the most important factor. When your applied science director documents a 93% click-through drop and describes a "doom loop" where your product destroys the market for the original works, you have built the plaintiff's case for them.
State of Mind: While not technically a factor in fair use analysis, courts consider whether infringement was willful when assessing damages. Internal notes saying "we trained GPT-3 on pirated stuff, no sharing that" and evidence of "Project Clear" to delete evidence of LibGen use could support a finding of willful infringement, which triples statutory damages from $150,000 to $450,000 per work.
The DOJ brief offers a lifeline by framing the question at the national security level. But a court in the Southern District of New York has no obligation to defer to the executive branch on a statutory question â especially when the executive branch has an undisclosed financial interest in the outcome.
What This Means for You
If you are building products on top of foundation models â and if you are reading ComputeLeap, you probably are â these filings should change how you think about legal risk.
1. Training Data Provenance Is Now a Liability Question. Every company building or fine-tuning models needs to know where the training data came from. If your model provider cannot certify clean provenance, you are inheriting their legal exposure. The authors suing OpenAI are not just going after OpenAI â they are establishing precedent that could affect every model trained on web-scraped data.
2. Licensing Is Coming, Ready or Not. The music industry went through this with Napster. The result was not the end of digital music â it was licensing frameworks (ASCAP, BMI, mechanical royalties). The AI industry will likely end up in a similar place. Companies that build licensing into their workflow now will be ahead of the curve. Some, like Reddit and the AP, have already started selling data access deals.
3. Watch the Southern District of New York. The hearing in the NYT case is scheduled for early 2027. Whatever Judge Stein decides will be the most consequential copyright ruling since Google v. Oracle. If fair use holds, the current model survives. If it does not, expect retroactive licensing demands, model retraining mandates, and a fundamental restructuring of how training data is sourced. The Authors Guild case, with its potentially tripled willful-infringement damages, adds a second front.
4. The DOJ Brief Is a Political Signal, Not a Legal Guarantee. Administrations change. Policies reverse. The current government's support for fair use reflects a specific political alignment that could shift. Building your business on the assumption that the government will always defend AI companies' right to train on unlicensed data is building on sand.
If you are interested in the broader liability landscape for AI agents beyond training data, or want to understand the growing public backlash against AI integration, we have covered both extensively.
What Comes Next
The legal machinery is now in motion. Summary judgment briefing in the NYT case continues through late 2026, with a hearing expected in early 2027. The Authors Guild case is on a parallel track. More publishers â over 30 newspapers now â have joined the litigation.
Meanwhile, the companies are not sitting still. OpenAI has signed data licensing deals with the AP, Vox Media, and others, suggesting it is hedging against a fair use loss. Microsoft's Satya Nadella testified in deposition that "anything that is paywalled should be licensed by anyone who wants to use it... for grounding or training" â a remarkable concession from the CEO of the company being sued for doing exactly the opposite.
The emails are out. The legal arguments are set. And the question that will shape the next decade of AI is headed to a courtroom in lower Manhattan.
đĄ The Bottom Line: These unsealed briefs are the strongest evidence yet that the companies at the center of the AI revolution knew their training practices were legally and ethically fraught â and did it anyway. Whether courts call that "transformative fair use" or "the largest theft of labor in human history" will determine the legal foundation of AI for a generation.
Originally published at ComputeLeap





Top comments (0)