Oxford lets OpenAI train its AI models on Bodleian Library

by | Sep 26, 2026 | Technology

Oxford lets OpenAI train its AI models on Bodleian Library

The University of Oxford entered into a partnership with OpenAI in March 2025 that includes digitizing historical texts from the Bodleian Library, one of the world’s oldest and most extensive academic libraries. The arrangement permits the company behind ChatGPT to incorporate the digitized material into its AI model training datasets. Internal documents indicate that material from the Bodleian has been used to “populate the OpenAI training set,” though this aspect was not emphasized in the initial public announcement of the partnership.

By June 2025, approximately 125,000 images scanned from historical dissertations had been shared with OpenAI, including doctoral theses from 19th and 20th century European and American universities. The digitized materials also encompass a collection of 10,000 16th-century broadside ballads featuring song lyrics and musical notation, as well as discussions about potentially digitizing 18th-century Irish state papers, letters from novelist Marie Edgeworth, and notebooks related to penicillin research. An OpenAI spokesperson characterized the initiative as a means of ensuring historical knowledge preservation within contemporary AI systems and promoting diverse cultural and historical perspectives in the technology.

The arrangement has prompted discussion among university staff about expansion possibilities, with the contract potentially opening pathways for mass digitization of the Bodleian’s 23 million items. Meeting minutes obtained through freedom of information requests documented concerns from governance committee members regarding reputational implications and environmental considerations related to energy-intensive AI technology. The partnership reflects a broader industry trend in which technology developers seek fresh training data from academic and historical sources, as internet-scraped content increasingly contains AI-generated material that has reduced utility for model development.

Oxford is the sole British institution participating in OpenAI’s NextGenAI project, which includes agreements with American research libraries including MIT, Caltech, Boston Public Library, and the University of Michigan. The university’s spokesperson emphasized that digitized material consists of out-of-copyright works on a “modest in scale” basis, with the Bodleian retaining rights to the scans and planning to publish them openly online in the coming months. Unlike practices by competing AI companies, the physical collections remain intact under the Oxford arrangement.

Article Attribution | Read More at Article Source

Article summary produced by Claude AI