Bodleian collection ended up in OpenAI's training set
When the University of Oxford and OpenAI announced their partnership in March 2025, the story was that OpenAI's software would digitise texts from the world-famous Bodleian Library, making the content more accessible to students and researchers. Internal model training was not mentioned. Yet internal documents, according to The Guardian's reporting on 26 September 2026, show that the digitised Bodleian material has been used to "populate OpenAI's training set". The case shines a light on how universities and libraries document — and fail to document — what they actually hand to companies hungry for fresh training data.
What was shared, and how much
The scope is partly documented. As of June 2025, 125,000 images scanned from historical dissertations had been shared with OpenAI from the Bodleian's collection, among them doctoral dissertations from European and American universities written in the 19th and 20th centuries. In addition, a rare collection of 10,000 16th-century broadside ballads was scanned — songs with lyrics and music that were once distributed on Tudor-era street corners (The Guardian, 26 September 2026).
A spokesperson for the University of Oxford, cited via The Guardian, described the digitisation as "modest in scale" and said it only covered material outside copyright. The Bodleian retains the rights to the scans, which are to begin being published openly online within a few months. The same details also appear in a Japanese industry report based on The Guardian, which, according to Oxford, limits the material to works whose copyright protection period has expired, with the rights remaining with the library (Hon no Kenkyusha, September 2026). The Japanese source is machine-translated and based on The Guardian, and must therefore be treated as confirmation of the same report, not as independent verification.
It is worth noting the precision that is missing: the sources document what was shared with OpenAI, but say little about how much of it was actually used for model training, as opposed to digitisation alone. That distinction is not documented in the available source material.
Internal resistance, documented via freedom of information
Meeting minutes from Oxford, obtained through a freedom of information request, show that staff — including members of the Bodleian's governance committee — expressed concern about the reputational risk of partnering with OpenAI, and about the effect on the university's environmental commitments of entering an agreement involving an energy-intensive technology. The minutes thus record the concerns raised as opinions, not as established harms.
The same minutes also discussed the creation of a chatbot called "Ask the Bod", and The Guardian reports that the OpenAI agreement raises the possibility of mass digitisation of the Bodleian's 23-million-item collection. That is a future prospect raised in internal work, not an adopted plan.
The pattern: NextGenAI and the hunt for paper
The Oxford agreement does not stand alone. OpenAI has entered similar agreements with American research libraries such as Boston Public Library, Caltech, MIT and the University of Michigan under the NextGenAI project, where Oxford is the only British member. The pattern points, as The Guardian frames it, toward a scarcity: model companies are hunting for fresh data, and what remains of unique material is largely located in physical collections that have never been scanned.
The same logic shapes a parallel report in The Guardian about antiquarian booksellers who have registered a wave of bulk orders for obscure titles — for example, a guide to 18th-century agricultural implements in Africa, or biographies of 1950s drivers. The booksellers have speculated that precisely these titles, which are rarely digitised online, represent fresh data that the next generation of AI models could consume. It is important to stress: this is the booksellers' speculation, not documented fact. The orders have not been confirmed as AI-related at this time.
What the parties say
An OpenAI spokesperson said, according to The Guardian, that the company was "proud" to ensure that "today's AI models preserve the world's historical knowledge for the future", and that with over a billion users it is important that the technology reflects diverse cultures, histories and perspectives. Oxford's spokesperson stressed, as noted, the modest scale and the absence of copyrighted material.
The reader should note how the two parties talk past each other: OpenAI speaks of preservation and cultural breadth; Oxford speaks of scale and copyright. Neither of the official statements addresses the disclosure gap itself — that training use was not mentioned in the original announcement.
Open questions
Several matters remain unresolved. First, the duration of the partnership is unclear: one source, a Japanese AI news summary, describes it as a "five-year partnership" (E&A Forlag, 27 September 2026), but the original report from The Guardian gives no duration, and the detail should therefore be treated as unconfirmed. Second, the split between scanned material and material actually used for model training is unclear. Third, there is the question of the transparency principle: if training use can be part of a library agreement without being mentioned in the public announcement, will other libraries considering similar agreements commit to more complete information — or follow Oxford's practice?
The Bodleian case is not a scandal in the classic sense: the material was out of copyright, the rights to the scans remain with the library, and the public is, according to Oxford, to be given access. But it illustrates a question that will become pressing as model companies turn their gaze toward physical archives: the value no longer lies in what is already digitally available, but in what exists only on paper — and it is the universities and libraries that must negotiate the terms, preferably before the announcement is written, not after.

