A September 17 public filing in the consolidated publisher copyright case against OpenAI and Microsoft alleges industrial-scale copying and quotes a Microsoft executive warning that model data collection could be viewed as the largest theft of labor in human history. The legal news is the release of a party's evidence and argument, not a ruling that the companies infringed.
Key facts
- The News Plaintiffs' combined summary-judgment brief was publicly refiled September 17.
- It concerns The New York Times, Daily News publishers, Ziff Davis, the Center for Investigative Reporting and The Intercept, among others.
- The brief attributes the internal labor quote to Microsoft Partner and Director of Applied Science Brent Hecht.
- Plaintiffs allege more than 91,692 copies of works in mid-training datasets and more than two million documents from nytimes.com in a Common Crawl-derived set.
The document's most memorable phrase is not the most important evidence. The underlying claim is a proposed feedback loop: AI products ingest journalism, answer users without sending them to publishers, reduce traffic and revenue, and then starve the reporting ecosystem that supplies future information. The filing calls this a doom loop. Like a restaurant that uses a farm's produce to sell meals while depriving the farm of its customers, the alleged harm is not merely a copy being made; it is a product substituting for the source relationship.
The plaintiffs point to discovery involving Common Crawl-derived data, Bing-indexed content and projects called Taxi and Mango. They allege Project Mango contained at least 160,903 unique publisher works and cite internal discussions about paywall circumvention. These are litigants' allegations and selected discovery excerpts, not neutral fact findings. That distinction also applies to a cited comparison alleging Copilot could reduce click-through to The New York Times by as much as 93 percent relative to conventional Bing.
OpenAI's case page and its summary-judgment memorandum provide the essential countercase. OpenAI argues training learns statistical relationships rather than distributing expressive copies, that facts are not copyrightable, and that outputs are rarely verbatim. It also disputes market substitution with economist evidence. Those arguments are not erased by an internal employee's rhetorical warning.
Still, the filing changes the public texture of the debate by putting the labor-market mechanism beside copyright doctrine. The sophisticated question is whether a legally transformative training process can nevertheless create a commercially substitutive interface that undermines the market the law is meant to protect. That will be decided through contested evidence and law, not a viral quotation. Readers tracking the issue should distinguish training data attribution from legal liability: being able to trace a model's inputs does not itself answer fair use, market harm or consent.
Originally published on Ground Truth, where every claim is checked against the primary source.
Top comments (0)