The University of Oxford has formed a partnership with OpenAI, a company known for its advanced artificial intelligence models, to allow the firm to use digitised materials from the Bodleian Library. These materials include PhD theses from European and American universities from the 19th and 20th centuries, as well as a rare collection of 16th-century broadside ballads—short, often humorous songs printed on paper. By June 2025, 125,000 images of historical dissertations had been shared with OpenAI from the Bodleian collection. The university announced the partnership in March 2025, stating that the digitisation would make the content more widely available for students and researchers. However, the announcement did not explicitly say that the material would be used for training OpenAI's models.
Internal documents obtained through a freedom of information request suggest that the material has been used to "populate the OpenAI training set," which is a collection of data used to train AI models to understand and generate human-like text. An OpenAI spokesperson said the company was "proud" to help ensure that "the AI models of today preserve the world's historical knowledge for the future," adding that it is important for the technology to reflect different cultures, histories, and perspectives. Meeting minutes from the University of Oxford reveal that some staff, including members of the Bodleian governance committee, expressed concerns about the reputational risks of partnering with OpenAI and the environmental impact of the energy-intensive technology used in AI development.
Booksellers have reported a sudden increase in orders for obscure books, which some secondhand bookshop owners believe may be for fresh data to train AI models. The digitisation of these materials is part of a broader project called NextGenAI, under which OpenAI has struck similar agreements with US research libraries such as Boston Public Library, MIT, and the University of Michigan. Oxford is the only UK member of the project. The OpenAI contract with Oxford raises the prospect of the mass digitisation of the Bodleian's collection, which consists of 23 million items. The minutes also discussed the creation of an "Ask the Bod" chatbot, which could potentially help users access information from the library's collections.
A spokesperson for the University of Oxford said the amount of text being digitised was "modest in scale" and covered only out-of-copyright material. The Bodleian retains the rights to the scans and will begin publishing them openly online within months. The spokesperson rejected the suggestion that the machine-learning element had been hidden from the public and students, stating that digitisation was the university's primary interest, although staff had been open about the project's contribution to training data. The Bodleian's collections remain intact under the deal, unlike with secondhand book acquisitions elsewhere that are being pulped after scanning.
Anthropic, OpenAI's close rival, has spent tens of millions of dollars acquiring books and slicing off their spines so their contents can be scanned before having them pulped. Anthropic has said it does not buy and destroy rare and antiquarian books. A tech news site, 404 Media, also placed a tracking device inside a secondhand book order and traced it to an Amazon facility in the US, where the books were also dismantled and scanned. The Oxford spokesperson said the material digitised through the project with OpenAI is modest in scale, out of copyright, and OpenAI's use of the material is not exclusive. The Bodleian libraries also retain the rights to make the digitised material available themselves, and the library will begin to publish these materials openly online in the next few months, as they do with outputs from other digitisation partnerships. The project will allow the Bodleian to make the material more accessible to a wider number of people who might otherwise have found it difficult to access.
University of Oxford Partners with OpenAI for Digitisation of Historical Texts
AI-rewritten from original reportingHow it works
ai-trainingoxford-universityopenaidigitizationbodleian-library



