AI Tech News HubDaily Updates
AI TechnologyAugust 19, 2026

Amazon Is Destroying Rare Books to Train AI — The Shift from Bookseller to Book Burner Is Deeply Unsettling

A
AI 觀察家
Columnist · 2021 words
Amazon Is Destroying Rare Books to Train AI — The Shift from Bookseller to Book Burner Is Deeply Unsettling

A librarian might spend decades digitizing a single 1800s ship's log for archival preservation. Amazon's procurement team could acquire that same book for a few hundred dollars, scan it, and dispose of it. This isn't science fiction — it's happening in 2026.

TechCrunch reported on August 17, 2026 that Amazon is systematically acquiring rare books, scanning their contents to train large language models, and then destroying or discarding the physical volumes once the process is complete. This is the same company that built its empire by selling books.

Why Rare Books Are So Valuable — to AI

In plain terms: the older a book, the less likely it is to exist anywhere online.

LLM training data draws heavily from publicly available web resources like Common Crawl, Wikipedia, and GitHub, but these sources largely span only the post-2000s era. 18th- and 19th-century literature, scientific papers, local historical records, manuscript translations? Nearly absent. You can think of rare books as "blind spot patches" for a model — filling them in delivers meaningful improvements in linguistic diversity, historical contextual understanding, and low-frequency vocabulary coverage. For Amazon's AI operations — including the foundation models on AWS Bedrock and the next-generation Alexa architecture — the marginal value of this data is exceptionally high, while the acquisition cost remains relatively low compared to alternatives. A book that a specialist antiquarian dealer might price at several thousand dollars could be a bargain at the scale of AI training data.

The Problem Goes Beyond "Books Being Destroyed"

On the surface, this looks like a straightforward resource decision — Amazon pays for data, booksellers get cash, market behavior, nothing wrong. But look closer and the logic starts to crack.

First, the value of rare books is irreproducible. Certain editions are historical artifacts in themselves — the wear on the pages, the marginalia, the binding style are all part of scholarly inquiry. Once the physical object is destroyed, you're left only with a scanned image, and that scan exists solely in Amazon's possession.

Second, there is no transparency about where this data goes. When a library digitizes a book, it's to make that content accessible to everyone. When Amazon scans a book, it's to make its own AI more capable. These two actions look similar in form but are diametrically opposed in nature.

Third, this behavior has an accelerating effect. When a player with Amazon's financial scale begins systematically acquiring rare books, these materials flow rapidly toward private capital rather than toward libraries or research institutions. This isn't a matter of one or two books — it's a distortion of the entire cultural preservation ecosystem.

When Tech Companies' Data Hunger Has No Boundaries

This situation brings to mind a broader pattern: the logic driving AI training data acquisition has begun to encroach on things we once assumed were simply off-limits.

Over the past few years, OpenAI and Google have been sued over scraping copyrighted content; artists have filed class-action suits against Midjourney and Stability AI; writers' guilds have negotiated licensing fees with various model companies. These conflicts share a common thread: tech companies' "data harvesting" during the training phase has consistently outpaced legal and ethical consensus.

What makes Amazon's move particularly ironic is its historical identity. Bezos built an empire on the idea that books enable the free flow of knowledge. Now that empire is physically eliminating the vessels of rare knowledge — so that its own models can become smarter. Calling it "knowledge transfer" is rather too generous. It looks more like enclosing the margins of the public commons and privatizing them.

If you've been watching the data strategies of AI giants for any length of time, this behavioral pattern may feel familiar. Last year, Zuckerberg's 6,500-word AI manifesto left readers more unsettled than reassured, for essentially the same reason — the gap between tech companies' "public interest" rhetoric when discussing AI's future and their actual behavior around resource control is conspicuously wide.

Who Should Be Uncomfortable About This?

If you're a researcher, a librarian, or anyone who cares about cultural preservation, you should be deeply uncomfortable.

If you work in AI, I think you should also stop and ask yourself seriously: if improvements in model capability are being built on irreversible cultural loss, has that cost been factored in?

There are currently no regulations that explicitly govern the "destruction of rare books for AI training purposes." Book transactions are private sales, and what the buyer does with a purchase is the buyer's business. But there is sometimes a considerable distance between "legal" and "unproblematic."

In the current AI competitive landscape, data quality matters more than data quantity — a trend that was already quite clear by 2026. The gaps between models in instruction-following and language understanding increasingly come down to the diversity and depth of training data rather than parameter count alone. This means the strategic value of "long-tail data" like rare books will only continue to rise, and behaviors like this will only become more common.

One Thought to Take Away

There is an argument that digitization is itself a form of preservation — the content of the book still exists, just in a different form.

But the question you need to ask is: preserved for whom? Who has access? Does that knowledge enter the public domain, or does it enter a private corporate database?

Amazon once transformed how books circulate, making them accessible to more people. What it is doing now points in precisely the opposite direction — making that knowledge accessible to fewer people, and doing so irreversibly.

This is not only a story about AI. It is a story about who owns human knowledge. And right now, no one is seriously answering that question.


References

Share

Related articles