The University of Oxford has shared digitized historical texts with OpenAI, a major American artificial intelligence company. Internal documents obtained by the Guardian newspaper show that these materials were used to train OpenAI's AI models. The collaboration, announced in March 2025, involved digitizing 125,000 images from ancient theses and other historical documents, including 10,000 "broadside ballads" — short, printed songs — and sets such as 18th-century Irish administrative archives and the notebooks of Dorothy Hodgkin, a Nobel Prize-winning chemist. However, the documents reviewed by the Guardian indicate that discussions about sharing some materials, like Hodgkin's notebooks, took place, but they were not actually sent to OpenAI. Oxford emphasized that the amount of material processed was "modest," only out-of-copyright documents were used, and the access granted to OpenAI was not exclusive. The Bodleian Library, Oxford's main library, keeps its rights to the digital reproductions and plans to publish them freely online, while the original physical documents remain in storage.
This practice is not unique to Oxford. The Bibliothèque nationale de France (BnF), France's national library, has also provided public domain texts to AI developers through the ArGiMi project. This project involves multiple partners, including Mistral AI, a French AI company. The BnF's texts are used to train an open-source AI model, with the goal of creating "digital commons" — shared digital resources — under open licenses. Additionally, the BnF offers commercial services to AI professionals, including access to free-of-rights data, digitization, and custom dataset creation. However, the BnF's partnership with Microsoft, which involves the digitization of 1,500 opera scores and costumes, does not explicitly state that these files will be used to train Microsoft's AI models.
In contrast to Oxford and the BnF, Anthropic, another AI company, has faced legal scrutiny over its acquisition and destruction of millions of printed books to create its training data. A U.S. court ruling in June 2025 detailed the purchase and destruction of these books, raising concerns about the environmental impact and the ethical implications of such practices. This contrasts sharply with Oxford's approach, where no historical volumes are destroyed, and the documents used are from the public domain.
The use of digitized texts by AI companies has sparked debates about digital sovereignty — the control of digital resources by a nation or group — copyright, and the ethical implications of training AI on cultural heritage. While some institutions, like the BnF, have established frameworks for licensing and usage, others face scrutiny over the potential misuse of AI-generated content. Additionally, the rise of AI-generated books and the impact on human authors have been topics of discussion, with some authors and publishers expressing concerns about the market's shift and the potential for AI-generated content to compete with human-created works.
Oxford Library's Digitized Texts Used in OpenAI Training, Raising Ethical and Legal Questions
AI-rewritten from original reportingHow it works
ai-trainingdigital-copyrightoxford-universityopenaianthropicbodleian-library



