The New York Times, the New York Daily News and a coalition of media outlets are asking a federal judge to impose sanctions on OpenAI, accusing the company of deliberately concealing evidence central to a landmark copyright lawsuit that could reshape how AI companies are allowed to train their systems on the work of journalists.

The filing, submitted Thursday in a Manhattan federal courthouse, marks a sharp escalation in a legal battle that has come to symbolize the AI industry's fraught relationship with the news business. The plaintiffs allege OpenAI "chose obstruction" over transparency, withholding training datasets and ChatGPT logs that could reveal how its models consumed copyrighted news articles.

What the Newspapers Are Alleging

According to the Associated Press, the newspapers claim OpenAI engaged in "discovery misconduct" that could distort the evidence available at trial. At the heart of the dispute are two things: the massive datasets used to train ChatGPT, and the logs that record how the chatbot responded to real user prompts. The outlets argue these records are essential to proving that OpenAI's technology reproduces and competes with their journalism.

The motion asserts that a recent deposition of an OpenAI employee directly contradicts the company's earlier statements about what it could and could not search for within its systems. New York Daily News attorney Steven Lieberman said OpenAI has been "making misrepresentations" for two years about its ability to identify copyrighted content inside its AI training data and logs.

"This motion asks the court to punish OpenAI for hiding and destroying evidence showing how ChatGPT was trained on stolen journalism," Lieberman said, according to AP. He represents the Daily News and seven of its sister newspapers.

OpenAI's Response

OpenAI pushed back forcefully, framing the dispute as an attempt by a weakening plaintiff to pry into the private conversations of ordinary users. In a statement reported by AP, OpenAI spokesperson Drew Pusateri said the company had described its limitations in sharing ChatGPT logs as a measure to protect user privacy.

"As the Times' case weakens and they've been forced to drop claims against us, they're persisting with their efforts to invade the privacy of people who have nothing to do with this case, including by making these blatantly false allegations," Pusateri said. "We'll continue defending our users' privacy and the long-established principles of fair use."

The clash underscores a core tension in AI litigation: the technical records that would most clearly show how a model learned from copyrighted text are also the records companies are most reluctant to disclose, citing trade secrets and user privacy.

A Case Born From the ChatGPT Boom

The New York Times sued OpenAI and Microsoft in late 2023, roughly a year after ChatGPT's public debut ignited a commercial AI boom and began rewiring how people search for information online. The lawsuit argues that training generative AI on millions of copyrighted articles without permission or compensation amounts to infringement.

The threat to publishers grew more acute in 2024, when Google began placing AI-generated summaries at the top of search results. Those overviews can answer a user's question directly, cutting off the click-throughs that translate into advertising revenue for the original source — the economic lifeblood of news organizations.

The Times has since been joined by other publishers, including MediaNews Group-owned papers such as the Daily News and the Chicago Tribune, digital media company Ziff Davis, and the nonprofit Center for Investigative Reporting.

The Broader Fair-Use Showdown

The sanctions fight is one front in a much wider legal war. OpenAI and other AI developers maintain that training their systems on books, articles, and other text scraped from the internet is protected by the "fair use" doctrine of U.S. copyright law. That theory is being tested in dozens of lawsuits brought by visual artists, novelists, music labels, and other creative industries, with courts reaching mixed conclusions so far.

Legal analysts watching the case note that sanctions motions, even when denied, can be revealing. They force both sides to explain precisely what evidence exists, what has been preserved, and what may have been lost — details that often stay buried until trial. A judge's ruling on the sanctions request could also shape the rules of engagement for future AI copyright litigation, establishing how much transparency companies must offer about their training data.

The discovery fight also highlights a structural imbalance that publishers say favors well-funded technology companies. OpenAI, backed by Microsoft and reportedly preparing for a public offering, possesses the technical infrastructure and legal resources to litigate for years. The newspapers, many of them already operating on thin margins, argue that protracted disputes over evidence access are themselves a way of running out the clock.

What Comes Next

The federal judge overseeing the case will now weigh whether OpenAI's handling of the requested materials rises to the level of sanctionable conduct. Possible outcomes range from financial penalties and adverse inference instructions — which would allow the jury to presume that missing evidence was harmful to OpenAI — to stricter oversight of the company's future discovery obligations. A trial date for the underlying copyright claims has not yet been set.

For the news industry, already grappling with declining revenues and shrinking newsrooms, the stakes are existential. If AI chatbots can synthesize their reporting without attribution or payment, publishers argue, the incentive to fund original journalism erodes further.

Stay Ahead of AI

For ongoing coverage of the copyright battles, regulatory fights, and policy decisions reshaping the artificial intelligence landscape, follow AI Buzz Wire for reporting you can trust.

Read more AI policy news →