OpenAI and Microsoft have admitted to the very risk publishers have warned about for years. Their AI models are taking the work of journalists and putting the news business at risk. Newly unsealed court documents from a lawsuit by news publishers show both companies talked inside their own walls about how generative AI could break the business model of the very outlets it depends on.
Court filings reveal that OpenAI's training datasets included over 91,000 copies of works from The New York Times, Daily News, and the Center for Investigative Reporting, as well as more than 2 million documents from nytimes.com via Common Crawl.
OpenAI's top leaders did not hide how their systems pull content. When shown that their tech could get past publisher paywalls to grab protected articles, OpenAI President Greg Brockman seemed to approve. Company records show they knew about the so-called 'doom loop'-a cycle where AI eats up publisher content, slowly destroying the ecosystem it needs for training data.
Microsoft has tried to distance itself from Hecht's comments, saying he was hired to play the role of an AI skeptic and does not speak for the company. OpenAI would not comment to the Orlando Sentinel, one of the lawsuit's plaintiffs. But the legal filings say these admissions undercut any claim that the companies are acting under fair use or in good faith. Steven Lieberman, lawyer for the NY Daily News, says the documents show clear harm and theft, not just a mistake.
According to Reuters, the lawsuit was filed in Manhattan federal court in 2023 and accuses OpenAI and Microsoft of using millions of newspaper articles without permission to train ChatGPT. Plaintiffs include The New York Times, Ziff Davis, The Intercept, and the Center for Investigative Reporting.
For publishers, the stakes could not be higher. The lawsuit's unsealed documents leave no room to call the industry's concerns alarmist. They show that AI leaders knew their products could shake up the market they rely on. These facts only came out in court, not through public statements. That says a lot about the power gap between tech giants and content creators. If publishers can't win real protections or payment, the 'doom loop' may lock in, leaving them with little leverage and shrinking revenue as generative AI spreads.
As a New York Times report explains, OpenAI and Microsoft still argue that training on media content is covered by fair use. That legal fight is now playing out with evidence from inside the companies themselves.