Wire
Oxford puts 125,000 Bodleian scans into OpenAI training
Oxford’s Bodleian Libraries have sent OpenAI at least 125,000 scanned historical-dissertation images, and internal minutes say the material was used to populate OpenAI’s training set. The university’s original collaboration announcement describes a five-year public-domain pilot covering 3,500 dissertations, a 500-user ChatGPT Edu test, rollout to 3,000 academics and staff, and OpenAI’s $50m NextGenAI research commitment; new reporting on the training use says Oxford retains scan rights, plans open publication, and describes the material as modest, non-exclusive, and out of copyright. For model builders, the provenance and statutory-exposure case is now a library-contract checklist: permission to digitise does not automatically settle training scope, exclusivity, or release rights.