Oxford let OpenAI feed scanned material from the Bodleian Library into its AI training data.
That's according to internal Oxford documents reported on by Ethan Penny and Dan Milmo, whose story broke September 26, 2026. The papers show texts that OpenAI scanned on-site at the Bodleian were folded into the company's training data. That's a step beyond the Oxford-OpenAI partnership the university had already made public. The reporting doesn't specify which texts were scanned or how much material was used.
That gap matters less than the shift it points to. AI labs have spent years scraping the open web and fighting publishers in court over books and news archives. A university library handing over scanned material directly is a quieter, cleaner way to get training data - no lawsuits, no takedown notices, just an institutional agreement.
Worth asking: what did Oxford get in return, and did anyone ask the authors of those historical texts?