Oxford lets OpenAI train its AI models on Bodleian Library

by | Sep 30, 2026 | Technology

Oxford lets OpenAI train its AI models on Bodleian Library

The University of Oxford has established a partnership with OpenAI that includes using digitized materials from the Bodleian Library to train the company’s artificial intelligence models. The arrangement was initially announced in March 2025 as a digitization initiative intended to increase accessibility of historical texts for students and researchers, though the initial public announcement did not specify that the material would be incorporated into AI training datasets.

Internal University of Oxford documents and meeting minutes obtained through freedom of information requests reveal that staff members, including members of the Bodleian governance committee, expressed concerns about the partnership. These concerns centered on potential reputational implications of collaborating with OpenAI and the environmental impact of engaging with energy-intensive technology. By June 2025, approximately 125,000 images from historical dissertations had been shared with OpenAI, including 19th and 20th-century doctoral theses from European and American universities, as well as rare 16th-century broadside ballads and other historical materials.

OpenAI operates the agreement as part of a broader initiative called NextGenAI, which includes partnerships with major research institutions such as Boston Public Library, Caltech, MIT, and the University of Michigan. Oxford represents the sole UK participant in this project. An OpenAI representative stated the company views the arrangement as a means to ensure contemporary AI systems preserve historical knowledge and represent diverse cultural and historical perspectives.

The University of Oxford’s spokespersons characterized the digitization effort as modest in scope, limited to materials in the public domain. The institution retains rights to all digitized scans and plans to publish them openly online within months. University officials noted that the materials remain intact and that OpenAI’s access rights are non-exclusive, distinguishing the arrangement from practices at other companies where physical books have been destroyed following scanning.

The partnership highlights a broader trend in which technology companies increasingly seek training data from academic and cultural institutions as publicly available internet sources become saturated with artificially generated content, reducing their utility for machine learning applications.

Article Attribution | Read More at Article Source

Article summary produced by Claude AI