AI Tech News HubDaily Updates
AI TechnologyAugust 20, 2026

From Selling Books to Burning Them: Amazon Is Feeding Rare Texts to AI — Is the Trade Worth It?

A
AI 觀察家
Columnist · 1894 words
From Selling Books to Burning Them: Amazon Is Feeding Rare Texts to AI — Is the Trade Worth It?

There's an image I've been turning over in my mind for a while: an 18th-century botanical manuscript — perhaps one of only three surviving copies in the world — its pages flattened by an industrial-grade scanner. Once the text extraction is complete, the book is routed to disposal. Not returned to a library. Not donated to a museum. Destroyed.

This isn't a scene from a dystopian novel. According to a TechCrunch report from August 2026, this is how Amazon is treating rare books it has acquired or obtained — in order to train its large language models.

Why Rare Texts Have Become AI's Most Coveted Resource

In plain terms: the internet has already been largely consumed by LLMs.

Model training in 2026 faces a very real problem — high-quality new textual data is increasingly scarce. Reddit, Wikipedia, news sites, GitHub: these sources have already been digested by previous generations of models, some more than once. To push the next generation forward in language comprehension, depth of reasoning, and cross-domain knowledge, training data must improve in both quality and diversity.

Rare books land squarely on this need. Writing styles, argumentative logic, and knowledge structures from centuries ago are linguistically worlds apart from the texture of modern web content. That diversity offers genuine benefits to a model's language understanding — think of it this way: the more varied the linguistic forms a model has been exposed to, the less likely it is to stumble when encountering unusual sentence structures or complex chains of reasoning.

This also explains why Amazon has gone to the trouble of acquiring physical documents that were previously difficult to digitize. The question is what happens to those books afterward — and that is precisely where things become troubling.

"Scan and Discard" — Who Could Accept That Logic?

I understand the business rationale. Preserving physical books is costly: storage, climate control, long-term management. Maintaining a rare volume is no trivial expense. From a purely asset-management perspective, "extract the digital content, dispose of the physical object" has a certain financial logic to it.

But there is a fundamental error embedded in that assumption — that a book's value resides entirely in its textual content.

The binding method of a manuscript, the chemical composition of its paper, the marginal annotations in a reader's hand, even the wear patterns on its cover: all of this is information. For historians, materials scientists, and cultural studies scholars, these non-textual physical details are sometimes more valuable for research than the words themselves. What a scanner captures is only one cross-section of what a book actually is.

To put it plainly: once these books are destroyed, they are gone. No backup. No recovery. Amazon has traded an irreversible physical cultural artifact for a digital text file.

This Points to a Much Larger Problem

I've written previously about the ethical context surrounding Amazon's destruction of rare books, focusing then on the framework of "technological greed" — corporations feeding commercial models with intellectual property while bearing no accountability toward the sources of that knowledge. I want to go a step further here.

The problem is not merely the moral choices of one company, but the data acquisition logic of the entire AI industry — a logic that has begun treating "can be converted into training data" as the ultimate purpose of everything.

Books can be digitized and then destroyed. Human creative work can be scraped and absorbed into model weights. The life's work of artists, writers, and scholars is reduced, under this logic, to "training material." A worldview that treats all knowledge in the world as a resource waiting to be extracted is, in some ways, more alarming than any single unethical act.

Think of it this way: this is not a bad actor doing a bad thing. It is an entire incentive structure systematically driving toward a particular outcome. As long as the cost of acquiring training data remains low, the legal grey zones remain wide, and external pressure remains weak, this will keep happening.

We Cannot Pretend the Choice Doesn't Exist

Amazon is not without alternatives. After scanning, these books could be donated to university libraries, national archives, or nonprofit institutions with the capacity to preserve them. Taking that extra step costs something — but it is not an astronomical sum.

More critically: if that choice was never seriously considered, that is the real problem.

On that note — this situation shares something uncomfortable with Zuckerberg's 6,500-word AI manifesto. Not because tech companies are lying, but because when they make decisions, non-commercial value is simply never part of the calculus. When profit, efficiency, and model performance are the only optimization targets, everything else is left to be sacrificed.

A Question to Carry With You

I am not arguing that AI's advancement is unworthy of pursuit. But "to train better models" should not function as a universal justification capable of overriding everything else.

Every rare book destroyed is a decision we cannot walk back. The question is who is making those decisions right now, and according to what standards.

If the answer is "a tech company's internal processes," then this conversation should have moved beyond the walls of that company long ago — and into a far more public arena.

A few points worth keeping in mind:

  • Rare texts are among the most sought-after training data sources for LLMs in 2026, precisely because the well of web content is running dry
  • The value of a physical book extends well beyond its textual content — destruction permanently erases layers of information that cannot be digitized
  • This is not an isolated incident, but a systemic result of an industry-wide "knowledge as resource" mentality
  • Amazon has the capacity to make different choices, but will only have the incentive to do so under external pressure

References

Share

Related articles