Back to News
RSS feedwww.theguardian.com

Oxford Allows OpenAI to Train AI Models on Bodleian Library Materials

Summary

The University of Oxford has allowed OpenAI to use digitised material from the Bodleian Library to populate its AI training set. The arrangement grew out of a partnership announced in March 2025 in which OpenAI software was used to digitise historical texts, but the public announcement did not say that the material would train OpenAI models. Internal meeting minutes obtained through a freedom of information request record staff concerns about reputational risk and the environmental implications of an energy-intensive technology. By June 2025, 125,000 images from historical dissertations had been shared, including nineteenth- and twentieth-century theses, and the project also scanned 10,000 sixteenth-century broadside ballads. Oxford staff discussed possible digitisation of Irish state papers, Marie Edgeworth’s letters and Dorothy Hodgkin’s penicillin notebooks, while the contract could eventually support wider digitisation of the Bodleian’s 23 million items and an “Ask the Bod” chatbot. Oxford says the current amount is modest, limited to out-of-copyright material, non-exclusive, and primarily intended to improve access; the library retains rights to publish the scans online and plans to do so within months. OpenAI said the work would help its models reflect diverse historical knowledge. The Bodleian’s collections remain intact, unlike books reportedly dismantled for scanning in some other AI-data efforts. Oxford is the only UK member of OpenAI’s NextGenAI library project, which also includes several US research libraries.