Key Takeaways
- Microsoft’s Director of Applied Science, Brent Hecht, described AI data scraping as “an astonishing theft of unprecedented proportions” in court filings unsealed in September 2026, directly undermining the fair use defence Microsoft and OpenAI are mounting in the New York Times lawsuit.
- The US Copyright Office’s May 2025 report concluded that AI training on copyrighted works is not “categorically fair use,” and that training may constitute “prima facie infringement” where model outputs substitute for the original works.
- The $1.5 billion Bartz v. Anthropic settlement, approved in July 2026, set a per-work figure of roughly $3,000, four times the statutory minimum for ordinary infringement, that plaintiffs in other cases are already citing as a benchmark. Court filings unsealed in September 2026 reveal that Brent Hecht, Microsoft’s Director of Applied Science, described AI data scraping internally as “an astonishing theft of unprecedented proportions”, a characterisation that now sits in evidence against the fair use defence his own company is running in the New York Times lawsuit against OpenAI and Microsoft.
The Lawsuits Piling Up
The New York Times case, filed in December 2023 and still in discovery as of 2026, is the most closely watched AI copyright dispute in the United States. The newspaper alleges that OpenAI and Microsoft used millions of its articles to train ChatGPT without authorisation, and that the resulting outputs can substitute directly for original journalism, causing quantifiable market harm. In April 2025, Judge Sidney H. Stein largely denied the defendants’ motions to dismiss, allowing the core infringement claims to proceed. Separately unsealed documents show an OpenAI engineer discussed a “hack to get around nytimes paywall,” a detail that further complicates the fair use framing.
The Authors Guild and co-plaintiffs, including George R.R. Martin and John Grisham, filed a motion for summary judgment on September 4, 2026, in their consolidated class-action against the same defendants. Their filing argues the companies “pirated” works and passed them “between them as currency,” explicitly rejecting fair use. The authors characterise GPT models as “an existential threat” to writers, pointing to AI-generated books that compete directly with human-authored titles.
Disputes are widening across creative sectors. In July 2026, a coalition of publishers including Hachette and Elsevier sued Google, alleging unauthorised use of their works to train Gemini and deliberate alteration of copyright management information. Image creators have mounted parallel challenges: the English High Court in November 2025 found only limited trademark infringement in Getty Images’ case against Stability AI and dismissed secondary copyright claims because model weights did not store direct copies, but a US District Court in California in April 2026 allowed Getty’s trademark and false-designation claims to proceed, based on allegations that Stability AI models generated images with distorted Getty watermarks. Earlier, in February 2025, a US federal court found ROSS Intelligence liable for copyright infringement after it used Westlaw headnotes to train a legal AI, rejecting fair use. The range of developer exposure across these rulings is broader than any single case suggests.
What Fair Use Actually Covers
AI developers have consistently argued that their models learn from data in a transformative way rather than copying it, making training legally analogous to a human absorbing ideas from diverse sources. The US Copyright Office’s May 2025 report is the most detailed official response to that argument, and it does not endorse it. The Office concluded that while fair use may apply in some narrow scenarios, the unauthorised use of copyrighted material to train generative AI systems “may constitute prima facie infringement.” Whether training is “transformative” is, in the Office’s framing, “a matter of degree,” and the Office stated it is “less inclined to support this defense where models generate outputs that infringe copyrighted works.”
The July 2026 final approval of a $1.5 billion class action settlement in Bartz v. Anthropic gave that legal uncertainty a dollar figure. The agreement compensates authors for Anthropic’s acquisition and copying of their works for training through August 25, 2025. The per-work award of roughly $3,000 was noted in the settlement proceedings as four times the statutory minimum for ordinary infringement, a benchmark plaintiffs in other cases are already citing. The settlement preserves authors’ rights to sue over AI outputs and any future conduct, so the $1.5 billion closes only one chapter of the dispute. Our earlier coverage examined how the $3,100 per-work figure is being used as a reference point across the broader litigation landscape.
Regulation and Creator Compensation
The EU AI Act’s intellectual property provisions took effect on August 2, 2025. Recital 105 states that “any use of copyright protected content requires the authorisation of the rightsholder concerned unless relevant copyright exceptions and limitations apply,” and requires General-Purpose AI providers to establish policies that respect machine-readable opt-out signals such as robots.txt and the proposed llms.txt standard. The compliance burden is concrete: providers operating in the EU must now demonstrate they have systems in place to honour those signals, not merely assert that they do.
On the economic side, a February 2026 UNESCO report titled “Re|Shaping Policies for Creativity” projected income losses for artists, estimating music creators could see revenues fall 24% and those in the audiovisual sector by 21% by 2028, attributed to both intellectual property violations and AI-generated content competing in the same markets. These are projections rather than observed losses, and UNESCO does not control for the many other variables affecting creator income. They reflect, nonetheless, the scale of concern driving legislative action across multiple jurisdictions.
The creative community has broadly advocated for three principles: consent, credit and compensation, though no single named organisation has been credited with originating this framework.
Hecht’s description of scraping as potentially “the largest theft of labor in human history” was an internal assessment, not a public position. It is now a court exhibit. The gap between what AI companies say publicly about training data and what their own executives have written internally may prove as consequential as any judicial ruling on fair use. The broader legal picture on training data, authorship and infringement is tracked in our AI copyright law coverage.
Originally published at https://autonainews.com/microsoft-executives-internal-theft-quote-undermines-ai-copyright-defense/
Top comments (0)