Oxford Shares Historic Bodleian Library Texts With OpenAI For AI Training

Oxford’s partnership with OpenAI, announced in March 2025 as a project to digitize Bodleian Library materials for students and researchers, also allows digitized texts to be used in OpenAI’s AI training sets, according to internal meeting records. By June, 125,000 images of historical dissertations had reportedly been shared, with other rare collections also discussed for digitization. Oxford says the materials are limited in scale, out of copyright and nonexclusive, and that the library retains the right to publish them; it plans to make the digitized materials available online. OpenAI says using historical texts can help its models reflect a wider range of cultures and histories. Some Oxford staff have raised concerns about reputational risks and the environmental impact of supporting energy-intensive AI infrastructure.
Oxford was the only UK institution in OpenAI’s NextGenAI project, which also included US institutions such as the Boston Public Library, Caltech, MIT and the University of Michigan.
The digitized material included 10,000 16th-century broadside ballads, with lyrics and musical notation. Discussions also covered digitizing 18th-century Irish state papers, novelist Maria Edgeworth’s private letters and Dorothy Hodgkin’s penicillin notebooks.
Oxford’s original March partnership announcement described digitization to improve access for students and researchers but did not explicitly say the texts would be used to train OpenAI’s models. Oxford later disputed that it had concealed the training use from staff or the public.
The Bodleian’s physical collections remain intact; the articles contrast this with some second-hand book acquisitions, where books may be removed from circulation or otherwise lost.
Publishers
6
Articles
5
Reach
11