AI Content Operations

OpenAI and Microsoft knew AI scraping could hurt publishers

OpenAI and Microsoft knew AI scraping could hurt publishers, lawsuit shows © fayfo.com
OpenAI and Microsoft knew AI scraping could hurt publishers, lawsuit shows © fayfo.com
Internal records show OpenAI and Microsoft saw the risk their AI posed to publishers. Executives talked about a 'doom loop' and the scale of content scraping. The lawsuit lays out how much the companies knew about the possible damage.

OpenAI and Microsoft have admitted to the very risk publishers have warned about for years. Their AI models are taking the work of journalists and putting the news business at risk. Newly unsealed court documents from a lawsuit by news publishers show both companies talked inside their own walls about how generative AI could break the business model of the very outlets it depends on.

One Microsoft document now made public says it's rare for a product to threaten its own main suppliers. But that's what these AI companies have done. The internal messages go further. Brent Hecht, Microsoft's Director of Applied Science, called the mass scraping of journalism "possibly the largest theft of labor in human history." These aren't outside critics. These are the words of executives at the heart of the story.

Court filings reveal that OpenAI's training datasets included over 91,000 copies of works from The New York Times, Daily News, and the Center for Investigative Reporting, as well as more than 2 million documents from nytimes.com via Common Crawl.

Reuters

OpenAI's top leaders did not hide how their systems pull content. When shown that their tech could get past publisher paywalls to grab protected articles, OpenAI President Greg Brockman seemed to approve. Company records show they knew about the so-called 'doom loop'-a cycle where AI eats up publisher content, slowly destroying the ecosystem it needs for training data.

Microsoft has tried to distance itself from Hecht's comments, saying he was hired to play the role of an AI skeptic and does not speak for the company. OpenAI would not comment to the Orlando Sentinel, one of the lawsuit's plaintiffs. But the legal filings say these admissions undercut any claim that the companies are acting under fair use or in good faith. Steven Lieberman, lawyer for the NY Daily News, says the documents show clear harm and theft, not just a mistake.

The lawsuit's details have shaken the publishing world. Publishers were already struggling with AI-driven content aggregation and falling ad revenue. The 'doom loop' is not just a theory. It's a real business problem that publishers now have to face. As reported earlier, even Google has started testing direct payments to publishers as AI-powered search cuts into old traffic sources.

According to Reuters, the lawsuit was filed in Manhattan federal court in 2023 and accuses OpenAI and Microsoft of using millions of newspaper articles without permission to train ChatGPT. Plaintiffs include The New York Times, Ziff Davis, The Intercept, and the Center for Investigative Reporting.

For publishers, the stakes could not be higher. The lawsuit's unsealed documents leave no room to call the industry's concerns alarmist. They show that AI leaders knew their products could shake up the market they rely on. These facts only came out in court, not through public statements. That says a lot about the power gap between tech giants and content creators. If publishers can't win real protections or payment, the 'doom loop' may lock in, leaving them with little leverage and shrinking revenue as generative AI spreads.

As a New York Times report explains, OpenAI and Microsoft still argue that training on media content is covered by fair use. That legal fight is now playing out with evidence from inside the companies themselves.

Ken Doctor Media analyst FAYFO Media
Media Analyst

Ken Doctor

An American media analyst, journalist, and publishing strategist