OpenAI staff discussed paywall workarounds, filing alleges

A newly unsealed legal filing in the consolidated copyright infringement case against OpenAI and Microsoft alleges that senior staff at both companies knew their practice of scraping news websites to train artificial intelligence models could damage or even destroy the publishers whose content they used. The filing, made public Thursday, cites internal company documents and executive testimony indicating the companies were aware their products were substituting for visits to original news sources.

Among the internal documents cited is one from Brent Hecht, Microsoft’s director of applied science, who wrote that the companies’ “AI content strategy has started a ‘doom loop’ that will hurt the performance of our models and the entire web at the same time.” He added: “It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its ‘content supply chain.’”

At OpenAI, ChatGPT leader Nick Turley used the word “substitutive” in remarks cited by the filing and said AI products “will get more and more substitutive as they get better.” Turley also wrote that publishers faced an “existential threat” from AI products. Internal OpenAI documents characterized ChatGPT as a “modern newsstand,” the filing alleges. Dario Amodei, then a top OpenAI researcher and now CEO of Anthropic, said in a presentation that “news generation” was a top skill of an earlier ChatGPT model, offering the sample query “What’s the NYT saying today?”

The filing also alleges that OpenAI employees discussed opportunities to skirt news sites’ paywalls during scraping. After an OpenAI researcher informed Brockman about a hack to get around the New York Times paywall, Brockman responded, “ah nice,” according to the filing.

In testimony, Microsoft CEO Satya Nadella said that “anything that is paywalled should be licensed by anyone who wants to use it” for purposes including AI training. Nadella added that he would have directed Microsoft to force OpenAI to retrain its models had he known the company was scraping and training on paywalled information. Nadella also testified that conversing with chatbots “substituted” for visiting websites, “versus needing to go to the underlying source.”

A Microsoft researcher told colleagues that compensating content creators was “in the best interest of my employer, of my country, and of many other groups I belong to,” according to the filing.

Microsoft and OpenAI have argued in court that their actions constituted fair use — which allows limited use of copyrighted material without permission — and that their use of the material was transformative. A Microsoft representative said Hecht’s comments “reflect one employee’s individual perspective, are not a legal analysis, and do not represent the company’s views.” OpenAI did not immediately respond to a request for comment. Anthropic also did not immediately comment.

The Microsoft representative added that the company lays out in court filings “why Copilot is not a substitute for publishers’ journalism.” The representative said Nadella had spoken about “broad changes in how people find information” and “those observations should not be confused with conclusions about copyright questions before the Court.”

The Justice Department recently filed a statement of interest in the suit, arguing that training on news content falls within the bounds of fair use under U.S. copyright law and that narrowing such carve-outs could hurt competition.

The New York Times sued Microsoft and OpenAI in December 2023, alleging that ChatGPT and Microsoft’s Copilot trained on millions of pieces of Times content and drew on that material to answer user queries. The Chicago Tribune, the New York Daily News and other newspapers owned by Alden Global Capital sued both companies in April 2024. Those and other cases have since been consolidated.

The filing comes as publishers pursue a dual strategy: striking content-licensing deals with some AI companies while suing others for copyright infringement. The Wall Street Journal parent News Corp has content deals with OpenAI and Meta, while two News Corp subsidiaries have sued another AI company, Perplexity.