Skip to main content
Destroyed rare books, including a first edition, with a tracking device, highlighting Amazon's AI training practices.

Editorial illustration for Amazon Destroys Rare Books for AI Training, Tracking Device Shows

Amazon Destroys Rare Books for AI Training

4 min read

A tracking device tucked inside a rare book led reporters at 404 Media to an Amazon facility in Las Vegas called VGT3, marked outside with a symbol of a dinosaur clutching a book in its claws. The book that ended up there wasn't headed for resale. According to 404 Media's investigation, Amazon is buying up rare and out-of-print books, slicing off their spines, and scanning the pages, not to sell them but to feed the text into AI training pipelines.

The company confirmed it purchases books "through commercial channels to improve the products and services customers use," a description that leaves out the part where the books get destroyed in the process. The appeal of rare texts is specific: large language models have already scraped most of what's readily available online, and Anthropic's own training data included books pirated outright. Books published before 2022 carry something increasingly scarce online, text nobody can dispute was written by a human before generative AI existed.

That matters because models trained on their own AI-generated output can degrade in quality, a problem researchers call model collapse. Amazon's spine-cutting operation sits at the center of that scramble for clean data.

Amazon is buying tons of rare books, cutting off their spines, and scanning them for AI training, according to 404 Media, which placed a tracking device in a rare book that ultimately arrived at an Amazon facility in Las Vegas.

Why this matters

For anyone building with AI, this is a supply-chain story more than a bookseller quirk. Amazon needs training data scarce enough to matter and rare, out-of-print books fit that gap better than another scrape of the open web. Cutting the spine off a book to feed a scanner is a small, physical detail, but it tells you how far companies will go once text becomes the raw material for a model rather than a product to sell.

404 Media had to plant a tracker in a physical book to get this confirmation. Amazon's statement, that it "purchases books through commercial channels to improve the products and services customers use," is technically true and says almost nothing. Developers and founders relying on foundation models trained by Amazon, Anthropic, or anyone else should assume some fraction of that data came from sources like this, obtained legally but without much transparency about scale or selection.

Worth watching: whether other publishers or archives start tracking their own stock the way 404 Media did, and whether "commercial channels" becomes a phrase regulators start asking harder questions about.

Common Questions Answered

How did 404 Media discover that Amazon was destroying rare books for AI training?

404 Media placed a tracking device inside a rare book that was ultimately delivered to Amazon's VGT3 facility in Las Vegas, a warehouse marked with a dinosaur symbol clutching a book. This investigation revealed that Amazon was purchasing rare and out-of-print books, cutting off their spines, and scanning the pages to feed into AI training pipelines rather than reselling them.

Why does Amazon specifically target rare and out-of-print books instead of using other text sources for AI training?

Rare and out-of-print books provide training data that is scarce and valuable enough to improve AI models in ways that standard web scraping cannot achieve. As text becomes the raw material for AI models rather than a product to sell, companies like Amazon are willing to destroy physical books to access unique textual content that would otherwise be unavailable for training purposes.

What is the physical process Amazon uses to prepare rare books for AI training?

Amazon cuts off the spines of rare books and scans the individual pages to digitize the text content. This destructive process transforms the physical books into digital text data that can be fed into AI training pipelines, effectively destroying the original rare books in the process.

What does Amazon's book destruction practice reveal about how companies view text as a resource for AI development?

Amazon's willingness to destroy rare books demonstrates that companies now view text primarily as raw material for AI models rather than as products with inherent value to be preserved and sold. This shift in perspective shows how far companies will go to acquire training data once text becomes essential to building and improving artificial intelligence systems.

LIVE19:02Amazon Destroys Rare Books for AI Training, Tracking Device Shows