New unredacted material in The New York Times' copyright lawsuit against OpenAI and Microsoft reveals a series of striking internal admissions — including a top Microsoft executive privately describing the companies' AI training practices as theft, and OpenAI's own leadership acknowledging that its chatbot poses an "existential threat" to the publishers whose work trained it.

The filing, reported by TechCrunch on Thursday, arrives three years after The Times first sued, alleging that the firms violated copyright law by training generative AI models on its content. The newly visible material details how the companies allegedly obtained and used that content: bypassing paywalls undetected, building training datasets through mass scraping, and deliberately stripping copyright notices from training data. For more context on this story, see our ongoing AI trends.

It is worth noting the caveat TechCrunch itself flagged: much of the new information comes from The Times' own brief rather than the underlying exhibits, which remain sealed, and the quotes are presented without their original context.

'A Doom Loop' for Publishers

Among the most consequential revelations is Microsoft's own data on what its Copilot "answer engine" did to traffic to The Times' website. According to the filing, Microsoft's internal figures show Copilot caused click-through rates for the nytimes.com domain to drop by as much as 93 percent compared with traditional Bing search.

An internal presentation written by Microsoft's director of Applied Science, Brent Hecht, in January 2024 described the decline as a "doom loop" that would "hurt the performance of our models and the entire web at the same time."

"It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its 'content supply chain,'" the Microsoft document reads, as quoted in the filing.

The admission cuts directly against one pillar of the fair-use defense that AI companies have relied on: that the use does not substitute for the original work or harm its market.

Executives' Own Words Enter the Record

The unredacted filings also capture candid assessments from the companies' most senior figures.

Microsoft CEO Satya Nadella testified in a deposition earlier this year that "anything that is paywalled should be licensed by anyone who wants to use it…for grounding or training." Nadella made clear that if he "had been made aware that OpenAI had scraped and trained on information that was behind a paywall," he would have "invoked [Microsoft's right to] require OpenAI to retrain its models."

Nadella also agreed under oath that conversing with chatbots "has substituted … giving you the information right there on the website on the AI platform versus needing to go to the underlying source."

On the OpenAI side, Nick Turley, the company's head of ChatGPT, wrote in internal communications that publishers face an "existential threat" from products like the chatbot, which are "largely substitutive" and "will get more and more substitutive as they get better." OpenAI President Greg Brockman described the models as "excellent at news."

And in a January 2023 internal memo, Hecht — the same Microsoft executive who later flagged the "doom loop" — called the scraping "an astonishing theft of unprecedented proportions" and "the largest theft of labor in human history."

The Sheer Scale of the Copying

The documents quantify the copying at a scale not previously disclosed. OpenAI's mid-training datasets alone, the filing says, contain more than 91,692 copies of works published by The New York Times, the New York Daily News, and the Center for Investigative Reporting.

A dataset derived from Common Crawl, the free open repository of web crawl data, included more than 2 million documents from nytimes.com alone. The filing also describes Project Mango, a Microsoft-OpenAI data initiative whose assembled training dataset contains copies of at least 160,903 unique works from the news publishers. OpenAI delivered its entire GPT-3 training dataset to Microsoft for evaluation, the filing says, and Microsoft similarly provided training data to OpenAI through initiatives called Project Taxi and Project Mango.

OpenAI employees also allegedly built early training datasets like WebText and WebText2 that disproportionately relied on scraped news content.

Paywall Hacks and Stripped Copyright Notices

Perhaps the most vivid details concern how the content was obtained. When OpenAI researcher Nick Ryder told Brockman about a "hack to get around nytimes paywall," Brockman — the filing says — replied: "ah nice."

The filings further describe deliberate efforts to strip copyright notices from training data before it reached the model, because researchers "wouldn't want" the model outputting "copyright notices" to users.

"The evidence revealed here for the first time shows that OpenAI and Microsoft knew that what they were doing was wrong," Steven Lieberman, counsel for the New York Daily News, said in a statement shared with TechCrunch.

What It Means for the Fair Use Fight

The question of whether AI firms can legally train on copyrighted material without a license remains unsettled, but judges have so far been largely favorable to AI companies' fair-use arguments. Earlier this month, the Trump administration contributed a brief in defense of OpenAI's unlicensed use of copyrighted material to train its large language models.

Yet several of the new admissions run counter to the fair-use test's market-harm requirement — the idea that the use must not substitute for or undercut the market for the original work. Internal statements that the products are "largely substitutive," combined with Microsoft's own data showing a 93 percent click-through collapse, give The Times' side its most concrete evidence yet of market substitution.

The filing is the latest escalation in a case that has become the defining legal battle over how AI companies acquire training data. Whatever the outcome, the unredacted record now shows that the companies' own executives understood the tension early — and put it in writing.

For continuing coverage of the legal fights shaping artificial intelligence, follow our latest AI news and in-depth analysis of the cases that will define the industry.

---

Stay Ahead of AI

Get the latest AI news, analysis, and breakthroughs — all in one place.

Read more AI news →