Artificial intelligence companies are purchasing antique and out-of-print books by the hundreds of thousands, feeding their contents into training datasets, and then destroying the physical volumes—an industrial-scale practice that researchers and booksellers say is erasing cultural heritage even as it fuels the next generation of language models. The disclosures, first surfaced in court documents and now echoed across European bookselling communities, have opened a new front in the AI copyright war. For breaking AI news on the training-data controversies reshaping the industry, the details below trace how a hunt for clean text turned into a bonfire of rare books.
How the Practice Works
According to documents disclosed by Anthropic during a 2026 court proceeding and reported by The Washington Post, the process is deliberately industrialized. The filing describes how a scanning contractor's "hydraulic powered cutting machine" would "neatly cut" books apart so that the pages could be "scanned on high speed, high quality, production level scanners." Once digitized, the filing noted, the scanning company would "schedule with the recycling company to pick up the completed books."
In other words, the books are not returned. They are reduced to digital text and then pulped.
Why Companies Want Old, Printed Books
The appetite for physical books stems from a worsening problem for AI developers: the open internet is increasingly polluted with AI-generated text. As the legal blog Verfassungsblog explained in an analysis titled "AI Is Eating the Book World," training on synthetic content risks degrading model quality—a phenomenon researchers call model collapse. Books printed before 2022, by contrast, are guaranteed to be human-written.
There is also a legal calculation. Companies have reasoned that scanning physical books they own may fall under the "fair use" provision of U.S. copyright law, a defense they believe is stronger than scraping copyrighted web content. Whether courts accept that argument remains an open question, but it has not stopped the buying frenzy.
The Zoom Books Phenomenon
In Europe, attention has centered on a Canadian online bookstore called Zoom Books, which Verfassungsblog reports has been acquiring hundreds of thousands of nonfiction titles for which there is little remaining commercial demand. These volumes—obscure academic works, regional histories, and specialized technical manuals—are nearly worthless to used-book sellers but represent a rich vein of previously untapped training material for AI labs.
Futurism reported that the destruction is happening "at incredible scale, even if almost no copies remain," meaning that for some titles the scanned-and-shredded copy may have been among the last surviving editions in circulation.
Booksellers and Researchers Push Back
The reports have alarmed antiquarian booksellers, who say bulk buyers are distorting the secondhand market and removing rare editions faster than archives can preserve them. Anadolu Agency quoted a Turkish antiquarian bookseller expressing skepticism about the scope of the claims, while acknowledging growing concern across European book markets. Coverage in the Times of India and Boing Boing framed the trend as a direct consequence of AI firms trying to "escape their own AI slop" by retreating to pre-internet print.
Preservationists argue that the destruction is especially troubling because no central registry tracks which books have been scanned and destroyed. Once a low-survival title is shredded, the loss may be permanent—unless the company that scanned it chooses to share the digital copy, which it generally does not.
A New Stage of the Copyright War
The book-buying campaign sits at the intersection of two unresolved legal battles: whether training AI models on copyrighted works without permission constitutes fair use, and whether destroying the physical source after scanning changes the calculus. Courts have so far sent mixed signals, with some rulings favoring AI developers and others preserving creators' claims.
For AI companies, the calculus is straightforward. Printed books offer dense, high-quality, human-authored text that is increasingly difficult to find online. For libraries, archives, and the collectors who guard rare editions, the trade-off is far less favorable: irreplaceable physical artifacts reduced to training tokens and then recycled into pulp.
The practice also raises questions about transparency. Anthropic's disclosure came only because it was compelled in litigation. It remains unknown how many other companies are pursuing similar strategies, how many books have already been destroyed, or whether any industry-wide standard will emerge to preserve at-risk titles before they are scanned and discarded.
The Scale Problem Nobody Can Quantify
Part of what makes the trend difficult to police is its diffuseness. Unlike a single large archive or database, the book-buying campaign is distributed across thousands of individual transactions—at estate sales, library deaccession sales, and online marketplaces—where a bulk buyer purchasing a pallet of nonfiction titles raises no obvious red flag. Aggregated across hundreds of thousands of purchases, however, the cumulative effect is the quiet removal of vast swaths of niche, low-circulation knowledge from the physical record.
Librarians and archivists have warned that the loss is not evenly distributed. Mass-market bestsellers survive in countless copies, but the titles most attractive to AI trainers—specialized, information-dense works with small print runs—are precisely the ones least likely to be preserved elsewhere. Once destroyed, they may exist only as proprietary training data locked inside a company's servers, inaccessible to the public that originally produced and relied on them.
Stay Ahead of AI
The training-data wars are intensifying, and the stakes extend well beyond corporate balance sheets. To read more AI news and follow how courts, regulators, and preservationists respond, keep following AI Buzz Wire.
Read more AI news →