A pilot scheme, launched in 2025, allows OpenAI to use obscure materials from the Bodleian Library to train its models, documents seen by Cherwell show.
The University of Oxford announced its partnership with OpenAI, the company behind ChatGPT, last March. A key part of the ongoing collaboration, which includes access to new ChatGPT models for researchers and ChatGPT Edu for students, is the digitisation of texts from the Bodleian Libraries using OpenAI’s software. When announcing the partnership OpenAI said that it would make “centuries old knowledge searchable by scholars worldwide”.
However, neither party’s announcement mentioned that the digitised material would be used to train OpenAI’s models. AI companies need large volumes of human created data to train models to recognise patterns in language and ‘learn’ by processing large amounts of data to produce sentences and perform tasks.
Whilst websites become increasingly saturated with AI-generated language, developers require human-made data to train Large Language Models (LLMs), such as ChatGPT. Consequently, competing companies are turning to physical documents and book collections. Anthropic, a leading AI company behind the chatbot Claude, has spent tens of millions on ‘destructive scanning,’ a technique which slices the spines of books to make scanning easier.
Minutes from a meeting between the library and OpenAI in 2024 show that “populat[ing] the OpenAI training set with knowledge that uses under-represented data on the open web” is a key aim of the partnership between the Bodleian and OpenAI.
Data privacy, particularly for storage and later use in AI training, was a key concern when signing the University’s wider agreement with OpenAI. An internal audit report on AI said: “When staff or students use personal ChatGPT accounts for University work, they may unintentionally expose sensitive information […] This creates both legal and reputational risk, particularly if data is retained for model training or exposed to third parties.”
According to internal documents obtained by The Guardian and seen by Cherwell, the digitised texts cover only out-of-copyright material and the Bodleian retains the rights to the scans with the intention to publish them online, making them widely available to students, researchers, and the general public.
The project is digitising over 3,500 dissertations from 1498-1884, which the University said “demonstrates how the Bodleian is using AI to imagine a library for the future.” Meetings discussing further texts for digitisation proposed the personal archive of Nobel prize winning scientist Dorothy Hodgkin for inclusion in the Bodleian’s digital archive.
Similar deals have been made with US universities, including the University of Michigan, MIT, Caltech, and Boston Public Library, under an OpenAI consortium called NextGenAI. Oxford is the only UK university included in the initiative.
According to its website, NextGenAI brings together 15 institutions who are “dedicated to using AI to accelerate research breakthroughs and transform education”. Launched in March last year, the partnership involves $50 million in research grants and compute funding. It aims to empower AI fluency and strengthen connections between academia and industry. It is not clear whether all these deals include giving OpenAI training data.
The Digital Head of Boston Public Library, Jessica Chapel, has defended NextGenAI: “OpenAI had this interest in massive amounts of training data. We have an interesting in massive amounts of digital objects. So this is kind of just a case that things are aligning.”
Documents seen by Cherwell show that staff raised multiple “reputational concerns” surrounding Oxford’s partnership with OpenAI. Academics have previously voiced concerns about the kind of language, particularly racially outdated terms, included within the specific data.
A spokesperson from the University of Oxford said: “The Bodleian libraries also retain the rights to make the digitised material available themselves, and the library will begin to publish these materials openly online in the next few months, as we do with inputs from other digitisation partnerships. The project will in fact allow the Bodleian to make the material more accessible to a wider number of people, who might otherwise have found it difficult to access.”
After the project’s announcement last year, Bodley’s Librarian, Richard Ovenden, said: “My view is that libraries cannot ignore AI, and must engage seriously with the industry. We are testing the use of the technology in workflow design, in data exchange, and in testing automated processes against human processes. We will share the output of our project when it is complete.”

