Microsoft called AI training on news 'the biggest theft of labor'
Declassified court materials in the case between American publishers, OpenAI and Microsoft showed that inside tech companies, risks of using millions of newspaper articles to train artificial intelligence have been discussed for years. One Microsoft executive called mass copying of content potentially "the largest theft of labor in human history," and inside OpenAI they warned that AI could become an "existential threat" to publishers.
The documents were released in the lawsuit filed by The New York Times and other American media companies against OpenAI and Microsoft. Publishers accuse tech companies of using copyrighted articles without permission and payment to create and train large language models. Microsoft and OpenAI deny wrongdoing and claim that training models falls under the US fair use doctrine - permissible use of protected works.
The sharpest assessments are contained in internal documents of Microsoft's director of applied sciences, Brent Hecht. In 2023, he wrote that millions of people might perceive the use of their works by large models as unprecedented theft, and he called the process potentially "the largest theft of labor in human history."
In another internal Microsoft presentation, Hecht warned about the risk of a kind of "vicious circle." In his opinion, AI services could reduce traffic and revenues of publishers, weakening the business of companies that create quality content needed by AI developers themselves to train future models. Microsoft recorded a drop in clicks to certain publishers' sites from Copilot by 83-93% compared to traditional Bing search, according to materials cited by plaintiffs.
Similar concerns existed inside OpenAI. ChatGPT head Nick Turley wrote in June 2023 that artificial intelligence poses an "existential threat" to publishers. In February 2024, he noted that AI products already largely replace primary sources and will do so increasingly as technology develops.
The scale of content used has also become one of the key arguments of media companies. According to court materials, a single dataset based on Common Crawl contained over 2 million documents from The New York Times domain. In other datasets for additional model training, plaintiffs counted tens of thousands of copies of materials from NYT, Daily News, and Center for Investigative Reporting.
A separate part of the claims concerns paywalled materials. Court documents cite internal correspondence in which an OpenAI employee informed co-founder and president Greg Brockman about a way to bypass The New York Times paywall. At the same time, the plaintiffs interpret these materials as evidence of deliberate circumvention of restrictions, but a significant portion of the original attachments to court documents remains closed.
Microsoft CEO Satya Nadella said in testimony that materials behind a paywall should be licensed for use in training or running AI. According to him, if he had known that OpenAI trained models on materials obtained this way, Microsoft could have demanded retraining of models.
Microsoft emphasized that Hecht's statements are the personal position of one employee, not a legal assessment or official company position. The corporation continues to insist that use of materials for AI is transformative and complies with copyright law, and Copilot is not a replacement for journalistic materials.
The case is being heard by the federal court for the Southern District of New York. Judge Sidney Stein is currently considering motions for summary judgment. For publishers, internal documents have become an important argument in trying to prove that AI products not only use their content but can also directly compete with primary sources. The court has not yet issued a final ruling on the legality of such use of materials.
Based on materials from: The New York Times, TechCrunch, Ars Technica