Scott Turow and Five Publishers Sue Meta Over AI Training Data
A class action filed in the Southern District of New York alleges Meta knowingly trained its Llama models on millions of copyrighted books pulled from pirate sites. The lawsuit names Mark Zuckerberg personally and argues Meta bypassed legal licensing to gain an AI advantage.
Bestselling author Scott Turow and five major publishing houses sued Meta Platforms and its CEO Mark Zuckerberg this month in a class action that argues the company deliberately built its Llama generative-AI models on millions of copyrighted books, knowingly sourced from notorious pirate websites including LibGen and Anna's Archive. The complaint, filed May 5 in the U.S. District Court for the Southern District of New York, is among the most aggressive of the more than thirty AI copyright cases now moving through federal courts.
The plaintiffs are publishers Hachette, Macmillan, McGraw Hill, Elsevier, and Cengage, along with Turow and his company S.C.R.I.B.E. The complaint seeks statutory damages, a permanent injunction against further use of the works, and an order requiring Meta to destroy all infringing copies of the copyrighted materials.
The allegations
The complaint describes a specific decision-making process at Meta in early 2023, when, according to the plaintiffs, the company "briefly considered licensing deals with major publishers" but pivoted in April of that year to using unlicensed material pulled from pirate websites. The plaintiffs argue this was not a passive consequence of internet-scale scraping but a deliberate strategic choice to avoid the costs of paying publishers for training data.
The decision was made, the complaint alleges, with Zuckerberg's personal authorization, which is why he is named individually as a defendant alongside the company. Personally naming a corporate CEO is unusual in copyright litigation and signals the plaintiffs' theory that the conduct at issue rose to the level of willful infringement rather than ordinary corporate decision-making.
"All Americans should understand that the bold future promised by AI has been, to paraphrase the investigative writer Alex Reisner, created with stolen words," Turow said in a statement issued through his publisher. "It is all the more shameful that these violations of the law were undertaken by one of the richest corporations in the world."
The class the plaintiffs seek to represent is broad. The complaint describes it as "all legal or beneficial owners of registered copyrights, in whole or in part, for any book possessing an International Standard Book Number (ISBN) or journal article possessing a Digital Object Identifier (DOI) or International Standard Serial Number (ISSN)." If the class is certified at that breadth, the potential damages exposure for Meta would be substantial.
Meta's response
Meta has signaled it will mount a fair-use defense. In a statement responding to the complaint, a Meta public affairs director said: "AI is powering transformative innovations, productivity, and creativity for individuals and companies, and courts have rightly found that training AI on copyrighted material can qualify as fair use."
The reference to courts that have "rightly found" fair use is itself contested. The leading recent decisions on AI training fair use have been mixed. Some district courts have permitted AI training as transformative use, particularly where the AI outputs do not substantially reproduce the underlying copyrighted text. Others have allowed cases to proceed past motion to dismiss, finding that the fair-use analysis requires a more developed factual record about how training data is used and what it produces.
The Meta case is likely to turn on the factual record. If discovery confirms that Meta downloaded full copies of books from sites like LibGen and Anna's Archive, the fair-use analysis becomes harder for Meta to win at an early stage. The Supreme Court's fair-use framework requires courts to consider the nature of the copyrighted work, the amount taken, and the effect on the market for the original. Wholesale copying of full books, sourced from sites whose entire purpose is to bypass the publishing industry, is a less sympathetic factual setup than the more abstract "we trained on the open internet" framing that defendants in earlier AI copyright cases have offered.
The legal landscape
The Turow case joins a growing line of AI copyright suits. In 2023, The New York Times sued OpenAI and Microsoft over similar allegations involving the GPT models. The Authors Guild filed its own action against OpenAI on behalf of literary authors. Recording labels have sued various AI music companies. Visual artists have sued image generators. The cases are fragmented across districts, with different judges reaching different conclusions on similar legal questions, in part because each case turns on the specific facts of how the defendant collected, stored, and used the training data.
What distinguishes the Turow case is the specificity of the alleged source. LibGen and Anna's Archive are not gray-area aggregators; they are pirate sites whose entire model is unauthorized distribution of copyrighted books. If Meta did source training data from those specific sites, the company's position on the fair-use defense becomes meaningfully harder than if the training data had been collected from a more diffuse mix of web sources.
The complaint also positions the case against the broader policy debate over AI training. Plaintiffs note that publishers have made licensing deals with other AI companies. The Associated Press licensed its archive to OpenAI in 2023. Reddit licensed its data to Google. Several large publishers have negotiated similar arrangements. The argument the plaintiffs are making is, in essence, that licensing markets for AI training data exist, that Meta could have used them, and that the company's choice to take material from pirate sources instead amounts to willful infringement rather than a good-faith disagreement about the scope of fair use.
What comes next
Meta will likely file motions to dismiss in the coming months, raising fair use, statute of limitations, and class certification challenges. The pace of AI copyright litigation has been slow but accelerating. The earlier Times v. OpenAI case has been in discovery for more than a year, and a similar timeline is plausible for this case.
The bigger question, looming over all of these cases, is whether courts will ultimately settle the fair-use question for AI training data, or whether Congress will act first. Several legislative proposals would clarify the fair-use rules for AI training, with some preserving broad fair use for training and others creating compulsory licensing schemes. None has advanced to passage. In the absence of congressional action, the law of AI training will be made case by case, and the Turow complaint is one of the more aggressive attempts to set the terms of that case-by-case development.
Sources
- Hachette Book Group et al. v. Meta Platforms, Inc., No. 1:26-cv-XXXX (S.D.N.Y. filed May 5, 2026)
- NPR: "Scott Turow's latest real-life legal thriller: Suing Meta for copyright infringement" (May 5, 2026)
- Statement from Scott Turow, S.C.R.I.B.E. (May 5, 2026)
- Statement from Nkechi Nneji, Meta Public Affairs Director (May 5, 2026)
- The New York Times Company v. Microsoft Corp. and OpenAI, No. 1:23-cv-11195 (S.D.N.Y. filed December 27, 2023) (related case)