Microsoft and OpenAI Employees Expressed Concern Over Using News Articles to Train AI
Newly unsealed court documents have revealed that employees at Microsoft and its partner OpenAI expressed significant concern about the potential harm their AI training practices could cause to the publishing industry. The documents show that as OpenAI advanced its artificial intelligence development, Microsoft staff internally debated whether the company’s use of millions of news articles to train large language models constituted what some called the ‘largest theft of labour in human history.’
The internal discussions highlighted fears that scraping vast quantities of journalistic content without proper compensation could create a ‘doom loop’ scenario. This term refers to a potential cycle where the degradation of the publishing industry’s economic viability could ultimately reduce the quality and availability of human-generated content, which in turn would threaten the quality of the AI models themselves that depend on this content for training.
These revelations come amid ongoing legal battles and public scrutiny over how major technology companies use copyrighted news content to develop their AI systems. The court documents provide rare insight into the ethical debates occurring within these companies as they race to build more powerful artificial intelligence. The concerns expressed by employees suggest that even within the companies driving AI advancement, there is recognition of the potential negative consequences their practices may have on journalism and the broader media ecosystem. This internal tension highlights the complex relationship between technological progress and the preservation of industries that produce the very content that makes AI systems possible.
