In one of the more inventive investigations into the AI industry's insatiable appetite for training data, journalists at 404 Media placed a tracking device inside a shipment of rare books they suspected would be bought for AI training — and followed it across the country to an Amazon facility in Las Vegas where, the outlet reports, books are scanned and destroyed.
The story, published Monday by 404 Media founding editor Emanuel Maiberg, identifies the destination as Amazon's VGT3 facility in Las Vegas. According to the investigation, Amazon has been buying massive quantities of books, scanning them for AI training data, and destroying the physical volumes in the process — an operation the outlet says had not been previously reported.
The findings add a major new name to a practice that has quietly troubled booksellers and archivists for months, and they arrive as public attention increasingly turns to where AI companies actually get their data. Follow our ongoing coverage of the training-data wars at AI Buzz Wire.
A Tracking Device in a Box of Books
The methodology behind the investigation is what makes it notable. Rather than relying on documents or interviews, 404 Media embedded a GPS tracker in a shipment of rare books that the outlet suspected would be acquired by an AI company, then monitored the package's journey to its final destination.
That destination — Amazon's VGT3 facility — is described in the report as the site where Amazon scans books for AI training data. The full investigation, which details the buying operation and its scale, is available to 404 Media subscribers.
The story spread quickly through developer and tech-policy circles after landing on the front page of Hacker News on Monday, where commenters debated the cultural and legal implications of destroying rare books to train commercial AI systems.
Amazon Joins a Growing List
The report is the latest chapter in a broader controversy over the fate of physical books in the AI era. Earlier this summer, booksellers and researchers documented how AI companies were purchasing antique and out-of-print volumes by the hundreds of thousands, feeding them into training datasets, and destroying the physical copies — a practice critics describe as the industrial-scale erasure of cultural heritage.
What distinguishes the new reporting is that it names a specific facility, a specific company, and a specific supply chain. Until now, much of the discussion relied on aggregate claims from booksellers and court filings. A tracked shipment offers a concrete data point: a box of rare books, a buyer, and a building.
Why Training Data Demand Keeps Escalating
The economics behind the practice are straightforward. Frontier AI models are widely reported to have consumed most of the high-quality text readily available on the open web, pushing companies to seek out harder-to-access sources: books, archives, and licensed corpora. Rare and out-of-print volumes are attractive precisely because they are unlikely to appear in standard web-scraped datasets.
For critics, the destruction of physical books to feed that demand has become a symbol of the industry's extractive relationship with human culture — data as a resource to be mined regardless of what is lost in the process. For the companies involved, digitized books represent a legal gray zone that courts and legislators are still sorting through, particularly as copyright litigation over training data works its way through multiple jurisdictions.
Neither scanning nor digitization itself requires destruction — the destruction, publishers and archivists note, appears to be a cost-saving measure, since storing physical volumes is more expensive than pulping them once their contents have been captured.
The Preservation Problem
Archivists have long argued that digitization is not preservation. A scan captures the text of a book but not its marginalia, its binding history, its material qualities — the features that make a rare volume a historical artifact rather than a container of words. Destroying the physical copy after scanning closes the door on future examination, and unlike digital files, a pulped book cannot be re-scanned as imaging technology improves.
Экономика книжной торговли усугубляет потери. Редкие и распроданные тома сохранились в небольшом количестве. Когда на рынок выходят покупатели с бюджетами в масштабах ИИ, они могут перебить цену библиотек и коллекционеров за оставшиеся копии, а когда эти копии уничтожаются после сканирования, количество выживших экземпляров данного издания может окончательно сократиться в течение нескольких месяцев.
Ассоциации книготорговцев в Европе были одними из самых ярых критиков этой практики, и расследование 404 Media дает их предупреждениям конкретные следы: предприятие, цепочку поставок и названного корпоративного покупателя.
Что будет дальше
Расследование, вероятно, усилит проверку роли Amazon в цепочке поставок данных искусственного интеллекта, в которой в противном случае доминирует ее облачный бизнес. Он также может возобновить призывы в Европе и США к введению правил, требующих сохранения культурных материалов, используемых в обучении ИИ. Эту идею выдвинули ассоциации книготорговцев и некоторые политики, поскольку масштабы покупки книг стали яснее.
На данный момент наиболее конкретным результатом расследования является информационный: указанное учреждение, задокументированное путешествие и публичный отчет о том, что предупреждения книжной торговли не были абстрактными. Редкие книги в отслеживаемой отправке завершили свое путешествие в Лас-Вегасе. Их содержимое, предположительно, сохраняется в обучающем корпусе.
Следите за войнами данных ИИ
Данные обучения, борьба за авторские права и этика ИИ в больших масштабах — ежедневно освещаются на AI Buzz Wire.
Читать больше новостей об искусственном интеллекте →