Newly released judicial documents from a lawsuit involving OpenAI, Microsoft, and authors' rights groups have revealed internal discussions at OpenAI about using pirated books from Library Genesis (LibGen) to train artificial intelligence models. Employees, including Tom Brown and Ben Mann, reportedly referred to LibGen as "fucking sketchy," while Dario Amodei raised concerns about the lack of transparency in naming training data sets "Books1" and "Books2" without disclosing their origin. These data sets, containing 12 billion and 55 billion tokens respectively, were used to train GPT-3 and GPT-3.5.
A federal court order from February 2026 indicated that an OpenAI employee had downloaded pirated books from LibGen as early as 2018. According to OpenAI's legal team, these data sets were deleted by mid-2022 and are no longer in use. However, the Authors Guild and prominent authors such as John Grisham and Jonathan Franzen are seeking a court ruling that OpenAI violated copyright laws, challenging the company's use of fair use as a legal defense. OpenAI argues that training generative models on protected works can fall under fair use, describing it as "transformative, analytical, and non-expressive."
The legal dispute has also highlighted the economic impact of generative AI on authors, with some OpenAI employees acknowledging that AI competition could lead to job losses in the writing industry. In a separate case, Bartz v. Anthropic, a $1.5 billion settlement was reached in July 2026. The court ruled that training AI models could fall under fair use but that downloading books from pirate libraries did not. This distinction may influence the ongoing legal battle involving OpenAI.
Other developments in the digital book space include Oxford University's collaboration with OpenAI to digitize its collections, concerns about AI-generated content in literary awards, and efforts by publishers to regulate AI usage through licensing platforms. Legal actions against piracy sites and the use of AI in audiobook production continue to shape the evolving landscape of digital publishing.
OpenAI's Use of Pirated Books for AI Training Under Scrutiny in Legal Dispute
AI-rewritten from original reportingHow it works
openaicopyrightfair-useai-traininglibgenlegal-dispute



