From Online Bookseller to Literary Pulping Mill: How Amazon Is Destroying Rare Texts to Feed Its AI
Thirty years ago, Amazon made its name by shipping cardboard boxes stuffed with paperbacks to doorsteps around the world. Today, the tech behemoth is buying up physical volumes in bulk, slicing off their spines with industrial guillotines, and feeding the pages into high-speed scanners. When I first read that an Apple AirTag dropped inside a bundle of rare books traced a path straight to an Amazon AI training hub in Las Vegas, the sheer historical irony hit like a sledgehammer. The company built on the backs of independent publishers and readers is now physically obliterating out-of-print literature to supply the raw token counts for its next generation of foundation models.
The Las Vegas Paper Mill
The trail was uncovered through an investigative effort by 404 Media. After an independent bookseller grew suspicious of anonymous bulk orders placed on Biblio—an online marketplace popular with rare and niche text dealers—journalists helped insert a tracker inside a 1,000-book shipment. The signal bounced across fulfillment networks before settling inside an Amazon facility in Las Vegas designated as VGT3.
Worker reports and internal forum chatter revealed the primary operation inside VGT3: receiving massive pallets of books, running them through industrial spine-cutters, feeding the loose sheets into document scanners, and dumping the physical debris. The practice, known inside tech circles as “destructive scanning,” strips the binding so high-throughput feeders can digitize hundreds of pages per minute without human intervention. In a dark touch of corporate humor, the internal team logo for the VGT3 facility reportedly features a T-Rex holding a book it’s about to devour.
When pressed on the findings, Amazon offered a characteristically clinical response: “Amazon purchases books through commercial channels to help develop and improve the products and services our customers use”.
Why AI Labs Are Hungry for Paper
To understand why a trillion-dollar company is buying rare physical texts just to shred them, you have to look at the severe data bottleneck facing the generative AI sector.
The Scraping Wall
The open internet has been scraped almost clean. Common Crawl datasets, Wikipedia, public forum threads, and open-access code repositories have already been ingested into models like Nova, GPT-4, and Claude. Finding fresh, high-quality prose online is becoming nearly impossible.
The Threat of Model Collapse
Because the web is now flooded with synthetic content generated by earlier LLM iterations, scraping current web pages runs the risk of feeding synthetic text back into new models. This recursive feedback loop leads directly to “model collapse”—a condition where output quality degrades over successive generations. Books printed before 2022 represent a goldmine of clean, uncorrupted, human-written language.
The Legal Loophole
Destructive scanning gained traction across the industry following court developments in high-profile copyright lawsuits against AI vendors like Anthropic. Some legal strategies rely on the argument that buying a physical copy, converting it into machine data, and destroying the physical original constitutes a “transformative” fair use process rather than digital piracy. By physically eliminating the book, the buyer claims they are not reproducing duplicate copies for distribution, but rather converting a single physical token stream into a mathematical one.
Why It Matters: The Erasure of Physical Provenance
For software developers, enterprise architects, and cultural historians, the implications of destructive scanning run deep.
+-------------------------------------------------------------------+ | TRADITIONAL vs. DESTRUCTIVE AI | +-------------------------------------------------------------------+ | Traditional Scraping | Destructive Scanning | +-----------------------------------+-------------------------------+ | - Digital text scraped from web | - Physical books bought bulk | | - High risk of AI-generated text | - Guaranteed human prose | | - Original web page remains | - Original book destroyed | | - Copyright gray zone | - "Fair Use" destruction loop| +-------------------------------------------------------------------+
When an AI lab buys up out-of-print or rare texts, it isn’t just acquiring data—it is actively reducing the surviving physical inventory of human knowledge. While these texts are not all museum-grade first editions, many are obscure historical records, specialized regional manuals, or niche literature with only a handful of physical copies remaining in circulation.
Once those pages pass through a spine-cutter, the digitized tokens end up trapped inside a proprietary weight matrix. The public cannot access the original text, verify the accuracy of the scan, or inspect the context. The physical artifact vanishes, replaced by a probabilistic representation sitting inside an AWS datacenter.
My Take: A Depressing New Era of Data Extraction
Is this legal? Thanks to loopholes in commercial acquisition and current interpretations of fair use, it very well might be. But it represents one of the most intellectually bankrupt practices in modern tech development.
There is a tragic absurdity in seeing the world’s most advanced artificial intelligence systems powered by the literal destruction of the printed word. We are watching tech giants treat centuries of bound human thought not as heritage to be preserved, but as biomass to be burned into a computational furnace.
If tech companies need high-quality human prose so badly that they are willing to set up dedicated shredding hubs in Nevada warehouses, they should be funding open-access digitization initiatives alongside libraries and archives. Instead, Amazon is using its massive capital to quietly buy up scarce physical media, extract the textual value for its proprietary models, and throw the remains into a recycling bin. It’s not innovation; it’s digital strip-mining.
Frequently Asked Questions
What is “destructive scanning” in the context of AI training?
Destructive scanning is a high-speed digitizing technique where the bound spine of a physical book is cut off so individual loose pages can pass through high-capacity scanners. The original physical book is ruined in the process.
Are these destroyed books copyright-protected?
Many of the books purchased in bulk carry active copyrights. However, AI companies argue that purchasing a physical copy through legitimate commercial sellers grants them ownership rights, and converting the text while destroying the physical copy qualifies as transformative use.
Why don’t AI companies just buy e-books or digital licensing?
Publishers generally do not sell bulk digital licenses that grant permission to ingest raw text directly into AI training corpora. Buying physical copies on secondhand or wholesale markets allows AI firms to bypass digital rights management (DRM) restrictions and licensing hurdles.
