The University of Oxford allowed OpenAI to train its AI models on historical texts from its Bodleian Library, with digitized material used to "populate the OpenAI training set," according to internal documents reported by The Guardian on Saturday — a purpose that was never mentioned when the partnership was publicly announced.

Oxford unveiled the collaboration in March 2025 as a digitization project, saying that using OpenAI software to scan texts from one of the world's most famous libraries would make the content more widely available to students and researchers. What the announcement did not say was that the material would feed the training of OpenAI's commercial models. Internal documents reviewed by the Guardian make that purpose explicit, marking one of the clearest examples to date of a prestigious cultural institution quietly supplying raw material for the AI industry.

A partnership with an unstated purpose

Details of the arrangement emerged through a freedom of information request, which turned up university meeting minutes recording concerns from staff — including members of the Bodleian's own governance committee — about the reputational risk of partnering with the company behind ChatGPT. Staff also raised the effect a deal with an energy-intensive technology company would have on the university's environmental commitments.

An OpenAI spokesperson defended the program, saying the company was "proud" to ensure "the AI models of today preserve the world's historical knowledge for the future." "With more than a billion people using this technology in everyday life, it's important it reflects different cultures, histories and perspectives," the spokesperson added.

From Tudor ballads to doctoral theses

The scale of the digitization is substantial. By June 2025, 125,000 images scanned from historical dissertations had been shared with OpenAI from the Bodleian collection, including PhD theses from European and American universities written in the 19th and 20th centuries. Other material scanned includes a rare collection of 10,000 sixteenth-century "broadside ballads" — song lyrics and musical notes once circulated on Tudor street corners.

None of it is secret in the conventional sense; much of the Bodleian's historical holdings are public-domain works, and copyright law places few restrictions on training models on texts that old. But the distinction between digitizing a library for scholars and populating a commercial training set is the kind of distinction that institutions have historically been expected to state out loud. Oxford's didn't — and the omission matters less for what it says about legal rights than for what it says about transparency between a university and its own community.

Libraries as the new data frontier

The rush for library shelves has a simple economic driver: the open web is running out of useful data. As scraped websites become saturated with AI-generated material, training on that content gets circular and degraded, and developers have turned to physical, often historical, book collections as a source of fresh, human-written text.

The symptoms are visible on the high street. The Guardian reports that secondhand booksellers have seen a spate of orders for obscure titles — a guide to agricultural implements in 18th-century Africa, biographies of 1950s car drivers — precisely because such books are unlikely to exist in digitized form anywhere online. Booksellers speculated the purchases represent fresh data for the next generation of AI models.

OpenAI has pursued the same playbook across academia and cultural institutions. The company has struck agreements with the Boston Public Library, Caltech, MIT, and the University of Michigan under a project called NextGenAI, with Oxford as the project's only UK member. The Bodleian, one of the oldest libraries in Europe with a collection spanning millennia, is by some distance the most storied addition.

What it means for cultural institutions

The Oxford episode is likely to sharpen a debate already running through museums, archives, and libraries worldwide: when a tech company offers to digitize a collection for free, what is actually being exchanged? Digitization brings genuine public benefits — access, preservation, searchability. But the Oxford documents suggest the value flows both ways, and that the training-set use was material to OpenAI even if it was not material enough to announce.

For institutions facing tight budgets, the temptation is understandable: professional digitization is expensive, and companies like OpenAI are offering to do it at no cost. The Bodleian's governance committee members who worried about reputational risk appear to have intuited what the FOI documents later confirmed — that the university's brand was lending legitimacy to a data acquisition effort, whether or not that was the intent.

The question now is whether other institutions will insist on transparency clauses in their AI partnerships: explicit disclosure of training use, opt-outs for sensitive material, and public reporting on what was shared and why. Until they do, the world's great collections will keep becoming training data — one quiet digitization project at a time.

---

Stay Ahead of AI

The battle for AI training data is reshaping industries far beyond tech. Get the latest AI news, analysis, and breakthroughs — all in one place.

Read more AI news →