“Authors Guild v OpenAI” is now a shorthand for a group of cases rather than a single one, and almost everything written about it describes allegations rather than rulings. This page states the procedural position: what was filed, when, where it sits now, and the short list of things a court has actually held.
What was filed, and by whom
The Authors Guild and a group of named fiction writers filed a putative class action against OpenAI in the United States District Court for the Southern District of New York on 19 September 2023, docketed as Authors Guild v OpenAI Inc., No. 1:23-cv-08292. The named plaintiffs included George R. R. Martin, John Grisham, Jodi Picoult, David Baldacci and Michael Connelly. The core allegation is direct copyright infringement in the reproduction of the plaintiffs’ books during the assembly of training corpora and in the training process itself, with contributory and vicarious theories pleaded alongside it.
A second SDNY action, Alter v OpenAI, No. 1:23-cv-10211, was filed on 21 November 2023 on behalf of non-fiction authors and named Microsoft as a defendant as well. Both are class actions, which matters for reading anything about them: until a class is certified, the only claims formally before the court are those of the named plaintiffs, and headline figures about “thousands of authors” describe a proposed class, not a party list.
This page is a summary of public procedural history and is not legal advice. If you are deciding whether your own training corpus, product or vendor contract exposes you to a claim of this kind, take advice on your own facts — the outcome turns on details, such as how the copies were acquired, that no summary of somebody else’s docket can tell you.
Consolidation into MDL 3143
On 3 April 2025 the United States Judicial Panel on Multidistrict Litigation ordered the OpenAI copyright cases centralised for coordinated pretrial proceedings as MDL No. 3143, In re: OpenAI, Inc., Copyright Infringement Litigation, in the Southern District of New York before Judge Sidney H. Stein. That order swept in the New York authors’ cases together with the news publishers’ actions and a set of cases that had been filed in the Northern District of California, so filings after that date are made on the MDL docket rather than on the original member-case numbers.
Centralisation is a case-management decision and nothing more. The JPML decides where pretrial proceedings happen; it does not decide any merits question, does not merge the claims into one claim, and does not survive the pretrial stage — member cases are in principle remanded to their home districts for trial. Reporting that treats the MDL order as a milestone in the substance of the dispute has misread it.
What has actually been decided
Very little, and less than the volume of coverage suggests. The rulings of substance to date have been on motions to dismiss and on discovery, not on liability.
- Pleading-stage rulings. Judge Stein’s April 2025 opinion on the motions to dismiss in the publishers’ member cases allowed direct and contributory infringement claims to proceed while dismissing certain claims under the Digital Millennium Copyright Act’s copyright-management-information provision, 17 U.S.C. § 1202(b). A denial of a motion to dismiss holds only that the complaint states a claim if everything in it is assumed true. It is not a finding that anything in it is true.
- No merits ruling. There has been no summary-judgment or trial ruling on infringement or on fair use in the authors’ cases. Anything you read describing “the Authors Guild fair use decision” is describing a different case — most often Bartz v Anthropic or Kadrey v Meta, both decided in the Northern District of California in June 2025 and both concerning different defendants and different corpora.
- No class certified. Certification under Federal Rule of Civil Procedure 23 is contested and is the point at which the practical size of the case is fixed.
The output-log preservation dispute
The most consequential order in the MDL for anyone building on the OpenAI API had nothing to do with copyright doctrine. In May 2025 Magistrate Judge Ona T. Wang directed OpenAI to preserve and segregate output log data that it would otherwise have deleted, on the plaintiffs’ theory that deleted conversations could contain evidence of infringing regurgitation. OpenAI objected publicly and procedurally, arguing the order conflicted with its own retention commitments to users, and the scope of the obligation was litigated through the second half of 2025 and subsequently narrowed.
The lesson generalises past this case: a litigation hold falls on the party that holds the data, and a provider’s published retention policy is a commercial commitment that a court can override. If your own compliance posture depends on a vendor deleting something within thirty days, that dependency has a failure mode that is not in the contract.
What the case has not decided
This is the section worth keeping. As matters stand, the litigation has decided none of the following, and it is not yet clear how any of them will come out.
- Whether training a generative model on lawfully acquired copyrighted books is fair use under 17 U.S.C. § 107. Two California district judges reached broadly favourable conclusions on that narrow question in 2025 on different reasoning; neither binds Judge Stein, and neither is appellate authority.
- Whether the acquisition of the copies is separable from the training. The distinction between a lawfully bought corpus and a pirated one has done more work in the decided cases than any argument about the model itself, and it is a factual question about provenance.
- Whether model outputs that reproduce protected expression create liability distinct from the training, and if so whose — the provider, the deployer, or the user who prompted it.
- Whether the model weights themselves are an infringing copy. This has been pleaded and has not been resolved anywhere.
- Any question of damages. The measure of damages in a class action of this shape is potentially enormous and entirely undetermined, which is why settlement pressure exists independently of the merits.
Checking the current posture yourself
Anything dated in this area goes stale quickly, so verify before you rely on it. The MDL docket is the authority; the JPML’s own site publishes the centralisation order and the current list of member cases, and the Judicial Panel on Multidistrict Litigation is where to start. Docket entries and filed opinions for the member cases are mirrored on CourtListener, which is free and searchable by case number.
When you read a new development, ask three questions before drawing any conclusion from it: which court, what procedural stage, and which defendant. A district court order is persuasive nowhere; an appellate ruling would change the landscape, and there is not yet one on the central question. For the doctrine the courts are applying, see the four-factor test as it has actually been applied to training, and for the parallel publisher litigation see the New York Times case and the broader training-data copyright picture.
Top comments (0)