Newly unredacted documents from a copyright lawsuit involving The New York Times, OpenAI, and Microsoft reveal that a top Microsoft executive privately called the companies’ AI training practices “theft.” The documents also show that OpenAI’s leadership described its AI models as an “existential threat” to publishers and journalists. The unsealed material details how the companies allegedly bypassed paywalls to scrape content from the internet, building massive training datasets without permission. Much of the new information comes from The New York Times’ own legal filing, though the underlying documents remain confidential. The quotes are presented without their original context, leaving some nuances unclear.
The unredacted filing is the latest development in a three-year-old lawsuit in which The New York Times accused OpenAI and Microsoft of violating copyright law by training their AI models on its content. The legal question of whether AI companies can use copyrighted material to train their systems remains unresolved. However, courts have generally supported AI companies’ arguments that such use constitutes "fair use," a legal principle allowing limited use of copyrighted material without permission in cases like parody, news reporting, or criticism. Earlier this month, the Trump administration submitted a brief defending OpenAI’s use of copyrighted material for training its large language models.
Microsoft’s internal data shows that its AI-powered Copilot feature significantly reduced traffic to The New York Times’ website. A January 2024 internal Microsoft presentation described this decline as a “doom loop” that could harm both AI model performance and the broader web. Microsoft CEO Satya Nadella testified that if he had known OpenAI had scraped content behind paywalls, he would have required the company to retrain its models. Meanwhile, OpenAI’s leadership expressed concerns about the impact of AI on journalism, with one executive describing AI chatbots as “largely substitutive” of traditional news sources.
The documents also highlight the massive scale of the alleged data scraping. OpenAI’s training datasets reportedly include over 91,692 copies of works from The New York Times, the Daily News, and the Center for Investigative Reporting. One Microsoft document warned that AI could “significantly disrupt” the employment of journalists and content creators. Internal Microsoft and OpenAI communications describe the scraping efforts as a way to avoid detection, including hacking paywalls and stripping copyright notices from training data to prevent AI from accidentally reproducing them. OpenAI and Microsoft have not commented on the allegations.
Legal Dispute Over AI Training Data Highlights Claims of Theft and Market Threats
AI-rewritten from original reportingHow it works
ai-trainingcopyright-theftmicrosoft-openaifair-usenews-publisherslegal-dispute



