Microsoft employees raised serious internal concerns about OpenAI’s use of copyrighted news articles and other material to train its artificial intelligence models, according to newly unsealed court documents cited by The New York Times.
In one 2023 internal document, Microsoft Director of Applied Science Brent Hecht warned that the practice could amount to the “largest theft of labor in human history.”
“Millions of people around the world will soon consider large models ‘hoovering up’ all their work to be an astonishing theft of unprecedented proportions,” Hecht wrote.
Hecht also reportedly warned of a potential “doom loop” in which damage to publishers and other content producers could ultimately reduce the supply and quality of material available for training future large language models.
The documents also show Hecht raising concerns that OpenAI may have engaged in what he described as an “accidental cover-up” while attempting to identify material belonging to The New York Times and other publishers involved in litigation.
Microsoft spokesman Alex Haurek said Hecht’s internal memoranda represented his personal views rather than those of the company.
OpenAI executives themselves discussed the potentially disruptive effect of generative AI on the news industry, according to the documents.
Nick Turley, who led the team developing ChatGPT, described AI in a June 2023 memo as an “existential threat” to publishers. In another memo in February 2024, he wrote that AI products “will get more and more substitutive as they get better.”
An OpenAI engineer also acknowledged concerns that providing links back to original news sources might not necessarily preserve publishers’ traffic, writing: “No matter how prominently we show the links, users won’t click.”
The court records further show OpenAI President Greg Brockman responding “ah nice” after an employee referred to developing “a hack” to bypass The New York Times paywall.
Microsoft CEO Satya Nadella, meanwhile, said during a deposition that paywalled material “should be licensed.” He added that had he known OpenAI’s models were trained on paywalled content, he would have required the company to retrain them.
Since launching ChatGPT in 2022, OpenAI has entered into licensing agreements with numerous news organisations. The broader dispute, however, centres on whether and under what circumstances copyrighted material can lawfully be used to train generative AI systems, as well as concerns that increasingly capable AI products could reduce traffic to the publishers whose material contributes to their responses.
The statements contained in the unsealed documents form part of ongoing copyright litigation and do not themselves constitute a judicial finding that OpenAI or Microsoft unlawfully used copyrighted material.


