A group of authors, including renowned novelist George R.R. Martin, has filed a motion for summary judgment in a federal court in New York, seeking to hold OpenAI legally responsible for allegedly copying 194 books without permission. The motion, directed to Judge Sidney Stein, argues that OpenAI did not legally purchase the books it used for training its AI models but instead downloaded them from Library Genesis (LibGen), a website known for hosting pirated content. The U.S. commercial representative has classified LibGen as a notorious pirate market, raising concerns about the legality of OpenAI's data acquisition methods. Microsoft, which has a significant partnership with OpenAI, is said to have supported and benefited from this activity, despite having the ability to prevent it.
The motion also claims that OpenAI deliberately obscured the source of its training data by renaming its datasets "LibGen" to "Books1" and "Books2" in its research paper about GPT-3. These datasets were later deleted in 2022 due to legal concerns, and the motion highlights that these are the only two training corpora OpenAI has ever deleted. Additionally, the authors cited statements from Tarun Gogineni, a former OpenAI employee responsible for the editorial quality of GPT models, who reportedly expressed a desire in 2025 for GPT-5 to complete Martin’s unfinished saga, even if Martin were to pass away before finishing it.
In response, OpenAI filed its own motion, arguing that its use of the pirated books falls under the legal principle of "fair use." The Department of Justice (DOJ) also weighed in, supporting OpenAI in a Statement of Interest on September 1, 2026, citing national security concerns related to China. However, the DOJ clarified that its support was limited to the use of AI models trained on legally acquired data, not the initial acquisition of pirated material. The authors are challenging this distinction, pointing to the case Bartz v. Anthropic (June 2025), where a judge ruled that downloading books from pirate libraries was "inherently and irreparably infringing," even if the data was later used in a transformative way. Anthropic eventually settled the case for $1.5 billion, the largest copyright settlement in U.S. history.
If the legal principles from the Bartz case are applied in this New York court, the authors argue, OpenAI’s defense would crumble, as the company's use of pirated data would no longer be protected, regardless of how transformative the AI training process might be. The case highlights the growing legal and ethical debates surrounding AI development, copyright law, and the sources of data used to train powerful language models.
Authors Sue OpenAI Over Alleged Copyright Infringement Using Pirated Books
AI-rewritten from original reportingHow it works
copyrightopenailibgengeorge-rr-martinfair-useai-training



