Escape Into Books
Your home for book news, reviews, and bookish fun. Follow along on Facebook.
Escape Into Books
Your home for book news, reviews, and bookish fun. Follow along on Facebook.

Amazon is facing scrutiny for its practice of destroying books to train artificial intelligence, raising concerns about literary heritage.
Reports emerged this week that Amazon has been acquiring physical print books, slicing off their bindings, and scanning their pages at high speed before discarding the physical remainders. The practice, known in tech circles as destructive scanning, allows large-scale digital ingestion of print texts to build datasets for software development and artificial intelligence models.
The operation came to light after investigative reporters tracked a specific volume using an internal location beacon. The shipment travelled across the United States to an Amazon logistics centre in Las Vegas, Nevada, designated as facility LAS8. According to workers at the location’s VGT3 unit, staff members receive large volumes of printed material specifically to remove their spines and feed the loose pages into commercial scanners, destroying the physical book in the process.
When asked about the Las Vegas facility and its scanning activities, Amazon provided a brief statement confirming that it “purchases books through commercial channels to help develop and improve the products and services our customers use.” The company did not specify which internal products or models rely on the scanned data, nor did it directly address the physical destruction of the books.
While the physical destruction of volumes sounds startlingly abrupt to anyone who loves print, the practice itself has become a growing open secret across the technology sector over the past two years. Earlier court proceedings involving rival AI developer Anthropic revealed details of an internal project aimed at scanning massive numbers of printed volumes, and ongoing legal filings against firms such as Meta, Microsoft, Google, and OpenAI suggest that bulk acquisition of physical text is a widespread strategy.
If you are wondering why multi-billion-pound tech firms are buying paper books off the market rather than scraping the web, the answer comes down to text quality and availability. Huge swathes of twentieth-century literature and non-fiction have never been officially digitised or posted online. For tech firms seeking well-structured prose, physical volumes offer a rich repository of long-form writing that cannot be harvested through standard internet scrapers.
There is also the problem of synthetic data. Any book printed before 2022 offers a clean baseline of prose guaranteed to be free from AI-generated text. Training artificial intelligence on data previously produced by other AI models leads to a degradation of output quality known as model collapse. Physical volumes printed in earlier decades provide a reliable safeguard against that recursive feedback loop.
The scale of these buying operations raises significant legal and structural questions for publishers and rights holders. When a technology firm buys a second-hand copy or orders a commercial volume, it legally owns that single physical unit. However, converting the contents into digital training data without securing digital licenses or compensating the copyright owner sits right at the heart of ongoing publishing lawsuits.
For independent booksellers and antiquarian dealers, the quiet withdrawal of physical stock from circulation is another worry. While reporting indicates these scanning operations target a broad range of printed material rather than exclusively rare or historical items, removing physical copies from the market permanently reduces the world’s surviving print record. When a physical edition is unstitched and shredded, that copy is gone for good.
The practice also comes at a time when traditional market channels are adapting to changing consumer habits. While US publishing revenue climbed 2.5% in 2025, author incomes remain under pressure, and content creators are increasingly vocal about how their intellectual property is harvested. Unlike digital piracy, where physical books remain untouched, destructive scanning systematically consumes the physical artifact alongside the text within it.
The preservation of printed heritage has always depended on physical stability. When rare Tolkien manuscripts were discovered in institutional archives in recent years, their survival relied entirely on the preservation of original paper and ink. Bulk scanning operations that reduce bound volumes to loose leaves for digital processing represent a starkly functional shift in how books are treated by corporate buyers.
Concerns over access and preservation extend beyond corporate warehouses into public institutions as well. We have already seen friction where local authorities and educational boards attempt to restrict access to printed titles, as shown when a Bay Area author reported Yosemite quietly banning his book from park facilities. Whether through deliberate administrative removal or corporate shredding for software training, the physical availability of books faces unprecedented pressure.
I think the public reaction to these revelations will be critical in deciding how tech companies proceed. Public optics matter, and few practices provoke as instinctive a negative reaction as the deliberate destruction of bound books. Whether commercial pressure or copyright litigation eventually halts the industrial unbinding of print volumes remains to be seen, but for now, thousands of physical books are making their final journey to scanner beds in Nevada.