Oxford shares Bodleian texts with OpenAI for model training
Oxford University has permitted OpenAI to use digitised texts from its prestigious Bodleian library to train the AI models behind ChatGPT, as technology companies increasingly seek fresh data from academic institutions. Whilst the partnership announced in March 2025 was framed around making historical texts more widely available to students and researchers, internal university documents reveal the material has been incorporated into OpenAI's training dataset—a detail not disclosed in the public announcement. The arrangement has raised concerns amongst Oxford staff about reputational risk and the environmental impact of partnering with an energy-intensive technology company.
By June 2025, over 125,000 images of historical dissertations had been shared with OpenAI, including 19th and 20th-century PhD theses and rare 16th-century broadside ballads. Oxford is the sole UK institution participating in OpenAI's NextGenAI project, which also includes major US research libraries such as MIT, Caltech and Boston Public Library. The University has stated that only out-of-copyright material is involved in the current digitisation effort, which it characterises as "modest in scale", and that the Bodleian retains rights to the scans, which will be published openly online within months. The potential scope of the arrangement is substantial, however, given the Bodleian's collection of 23 million items.
- Oxford allowed OpenAI to train ChatGPT on Bodleian library texts
- 125,000 historical images shared by June 2025; sole UK NextGenAI member
- Staff raised concerns; university says digitisation was primary objective
AI Art Business Companies Culture Technology
Read the full article at the source →
Originally published by The Guardian as “Oxford lets OpenAI train its AI models on Bodleian library”.