Two of America's largest publishers — Hachette Book Group and Cengage Group — have filed a proposed class action lawsuit against Google, accusing the technology giant of illegally copying millions of copyrighted works to train its Gemini AI service. The suit, filed on July 10, 2026, in federal court in New York and joined by bestselling author Scott Turow, marks a significant escalation in the publishing industry's fight against unauthorized AI training.
The lawsuit, first reported by Publishers Weekly and covered by Publishing Perspectives, alleges that Google scraped copyrighted works from multiple sources — including content behind paywalls, known pirate sources, and books previously submitted for its Google Books database, Google Play retail service, and Google Scholar platform. For ongoing coverage of AI policy and litigation, readers can follow AI Buzz Wire.
Accusations of Mass Copying
The complaint accuses Google of systematically acquiring copyrighted material through programs that were designed for entirely different purposes. According to the publishers, Google used copies of books from Google Books, Google Play Books, and Google Scholar — all of which were provided to Google under what the publishers describe as "scope-limited" arrangements.
The lawsuit also alleges that Google removed copyright identification to conceal its actions. The publishers argue that nothing in the landmark 2015 U.S. appeals court decision affirming the legality of Google Books authorizes the company to make copies of copyrighted works for the new purpose of training commercial AI models.
"The result is an AI system that competes directly with Plaintiffs' and the Class's works in the market," the filing states. The complaint describes AI-generated substitutes taking multiple forms, including verbatim and near-verbatim copies of portions of works, replacement chapters of academic textbooks, summaries of famous novels, and what the publishers characterize as "inferior knockoffs that copy creative elements of original works."
A Novel That Costs 39 Cents
One of the most striking claims in the lawsuit highlights the competitive threat that Gemini poses to traditional publishing. The filing notes that Gemini can generate "a 100-page murder mystery set in a quiet seaside town filled with secrets, that substitutes for an original copyrighted murder mystery on which Gemini trained" in approximately 20 minutes for just 39 cents.
"No publisher or author can compete with that," the complaint states. "Users are already touting Gemini's ability to generate books with ease, and the market is flooding with AI-generated substitutes. The scale and speed at which Gemini can create books and compete with human writers is unprecedented, and it can only do that because Google copied Plaintiffs' and the Class's works to train its AI."
Internal Communications Cited
The complaint cites several internal Google communications suggesting that the company's own officials acknowledged that using publisher-provided copyrighted books from Google Play Books to train AI was potentially "highly problematic for Google." This evidence, if admitted, could prove significant in establishing that Google was aware of potential legal issues with its training practices.
The Google Play Books angle may be particularly vulnerable. The complaint notes that Google Play is "a retail storefront through which publishers and authors sell books," with authors and publishers providing Google access to digital books for the "limited purpose of selling authorized ebooks."
A Broader Legal Campaign
The Google suit comes just two months after five major publishers — Elsevier, Cengage, Hachette, Macmillan, and McGraw Hill — filed their first AI lawsuit against Meta and CEO Mark Zuckerberg in New York, organized by the Association of American Publishers. The Google complaint is based on nearly identical claims of willful infringement of millions of works used for training large language models.
Notably, the publishers also withdrew a separate motion to intervene in an existing copyright case against Google in California — In Re Google Generative AI Copyright Litigation — which was first filed in 2023 and is before Judge Eumi K. Lee in the Northern District of California. Lee has yet to rule on the publishers' motion to intervene despite holding a hearing nearly five months ago on class certification.
The publishers explained that filing their own suit "aims to preserve the right to pursue all the claims that publishers and their authors have against Google, including important ones that fall outside the putative class in that case."
The Legal Landscape
Copyright lawsuits over AI development have delivered mixed results in U.S. courts. In June 2025, Judge William Alsup found that Anthropic's unauthorized use of copyrighted books to train its Claude AI system was fair use — but also ruled that the company's decision to keep millions of unauthorized downloads for a permanent research library was not. That finding led to a $1.5 billion settlement that is awaiting final court approval.
In a separate case, Judge Vincent Chhabria similarly found that AI training on copyrighted works was fair use but flagged concerns about "market dilution" — the possibility that AI-generated works could disrupt the marketplace for human-authored content. The publisher suits now feature market dilution claims prominently.
The publishing industry's filing adds to a growing list of more than 128 AI-related copyright actions in U.S. courts, according to legal blogger Edward Lee's running tally. The outcome of these cases will shape the legal framework governing how AI companies can use copyrighted content for training.
What the Publishers Want
The proposed class action seeks a court order declaring that Google's actions violated the Copyright Act, an injunction barring future infringement, monetary damages up to the maximum allowed by law, and the destruction of all infringing copies in Google's possession.
As the AI industry matures, the tension between technological innovation and intellectual property rights shows no signs of resolving. For authors and publishers, the lawsuits represent an existential defense of creative work. For AI companies, the ability to train on vast datasets is foundational to their products. The courts will ultimately decide where the boundaries lie.
Stay Ahead of AI
AI policy and copyright law are evolving rapidly. For daily updates on the legal battles shaping the future of artificial intelligence, bookmark AI Buzz Wire.
Read more AI news →

