Study Claims OpenAI Used Paywalled O’Reilly Books to Train AI Models
A recent study by the AI Disclosures Project alleges that OpenAI’s GPT-4o model was likely trained using paywalled books from O’Reilly Media without a licensing agreement. The research, co-authored by Tim O’Reilly, Ilan Strauss, and Sruly Rosenblat, employed a method called DE-COP to test whether GPT-4o had prior knowledge of restricted content. The findings indicate that GPT-4o shows significantly higher recognition of paywalled O’Reilly books compared to its predecessor, GPT-3.5 Turbo.
The paper suggests that OpenAI may have accessed this data through various means, including users copying and pasting content into ChatGPT. However, the researchers acknowledge that their method isn’t definitive proof of unauthorized usage. The study also highlights OpenAI’s broader search for high-quality training data, including hiring experts and entering licensing deals with publishers and media platforms.
OpenAI has not commented on the allegations, but the company has faced multiple lawsuits related to its data usage practices. While OpenAI offers opt-out mechanisms for copyright owners, concerns persist over how AI companies obtain and utilize copyrighted material for model training.
