USA Today Co., Inc. and its affiliated newspapers sued OpenAI in federal court in New York on October 8, 2026, alleging the company trained its GPT models on hundreds of thousands of their copyrighted articles without authorization. The suit, first reported by Reuters, seeks damages in excess of $250 million and adds the largest name yet in local and national news publishing to the wave of copyright litigation against the maker of ChatGPT.
The case, docketed as USA Today Co., Inc. v. OpenAI Foundation in the U.S. District Court for the Southern District of New York, arrives as courts, regulators, and newsrooms continue to wrestle with how generative AI systems should treat copyrighted journalism. For readers tracking the latest AI developments, the complaint is one of the most detailed indictments yet of how a frontier lab's training corpus was assembled.
What the Lawsuit Alleges
According to reporting by Reuters, Forbes, and Bloomberg Law, the complaint alleges that OpenAI's large language models were trained on copyrighted material scraped from the internet without authorization, regardless of paywalls or other access restrictions. The plaintiffs also allege OpenAI used programs designed to strip away copyright management information, the ownership and rights data attached to published works.
Under the statutory framework the complaint cites, the plaintiffs may recover up to $150,000 for each willful copyright infringement, plus up to $25,000 per violation for the removal of copyright management information. The suit demands a jury trial.
The filings, described in detail by Unite.AI based on the court docket, were accompanied by an exhibit of copyright registrations and a second exhibit of GPT-5.6 output examples. Those examples, according to the complaint, show the model producing extensive multi-section summaries that paraphrase original articles and follow their structural organization when prompted to find a specific story by title.
The Numbers in the Complaint
The complaint states that content from the plaintiffs' publications comprises more than 160,000 entries in WebText, the internal corpus OpenAI built to train GPT-2, including 83,266 entries from usatoday.com and 12,994 entries from freep.com, the Detroit Free Press domain. It also states the publications' domains account for more than 122 million tokens in C4, a filtered English-language subset of a 2019 snapshot of the Common Crawl web archive.
To argue that OpenAI understood what its models contained, the complaint quotes internal communications. It states that co-founder Greg Brockman told colleagues the models were particularly good at predicting the text of news articles, and that a 2020 presentation by then-research leader Dario Amodei listed news generation among GPT-3's skills. The complaint also quotes an OpenAI vice president of research stating, "We train our networks to memorize the training data — that's their objective."
The complaint additionally cites written evidence OpenAI submitted to a British House of Lords inquiry in December 2023, in which the company stated that because copyright covers virtually every sort of human expression, limiting training data to public-domain works would not produce competitive AI systems.
Microsoft's Alleged Role
The filing also sweeps in allegations about Microsoft, OpenAI's close commercial partner. According to the complaint, Microsoft provided OpenAI with a copy of the Bing Index, a compilation of billions of scraped webpages that included the plaintiffs' content, under an initiative codenamed Project Taxi over a three-year period. It further alleges Microsoft separately developed and operated a crawler called Project Mango on OpenAI's behalf, for which OpenAI paid Microsoft.
The complaint alleges OpenAI's output filters did not suppress content from publishers that had not yet sued the company, an approach it says a Microsoft executive described internally in unflattering terms. The plaintiffs are pursuing OpenAI entities only; Microsoft is not a named defendant in this case, though it faces similar claims in earlier publisher suits.
Who Is Suing, and Who Is Being Sued
The plaintiffs, all owned by USA Today Co., hold copyrights in content published by 19 publications, including USA TODAY, The Tennessean, the Indy Star, The Bergen Record, The Enquirer, the Asbury Park Press, the Democrat & Chronicle, The Knoxville News-Sentinel, the Naples Daily News, The Oklahoman, the Milwaukee Journal Sentinel, The Columbus Dispatch, The Arizona Republic, The Courier-Journal, The Des Moines Register, the Detroit Free Press, The Detroit News, The Palm Beach Post, and the Star News. USA Today Co. is the former Gannett Co., Inc.
The defendants are seven OpenAI entities, including the OpenAI Foundation, the nonprofit that emerged holding equity in the for-profit after the company's October 2025 recapitalization, and OpenAI Group PBC, the public benefit corporation that now houses the commercial business. The models at issue span GPT-1 through GPT-6.1 and the GPT-OSS open-weight releases, including all Instant, Thinking, mini, nano, and Pro variants, according to the complaint.
Substitution Is the Core Theory
A central theme of the complaint is substitution. The plaintiffs allege OpenAI affirmatively post-trained its models to summarize copyrighted articles in lieu of returning the articles themselves, producing substitutes that serve the same informative purpose as the originals. The complaint quotes OpenAI's head of ChatGPT acknowledging that once the chatbot gives an answer, there is "no good reason to click" a link to the underlying source, and cites an OpenAI engineer's statement that users will not click links no matter how prominently they are displayed.
That theory echoes arguments in earlier cases. The New York Times sued OpenAI and Microsoft in December 2023, and outlets including the Seattle Times and Newsday have filed their own copyright suits. A statement of relatedness filed with the USA Today action asks that it be treated as related to the consolidated OpenAI copyright litigation already pending in the same court.
Why It Matters
The outcome of the consolidated cases could define the economics of AI-generated news summaries for years. Publishers argue that chatbot answers that recap their reporting destroy the traffic that funds it; AI companies argue that training on publicly available text is transformative fair use. OpenAI has not yet publicly responded to the USA Today complaint, and prior responses in similar suits have emphasized licensing partnerships and disputed the plaintiffs' reading of fair use.
For continuing coverage of this case and the breaking AI news shaping the industry, stay tuned as the litigation moves through the Southern District of New York.
---
Stay Ahead of AIGet the latest AI news, analysis, and breakthroughs — all in one place.
Read more AI news →