Why Would Anyone Destroy a Rare Book After Paying For It?

Core Takeaway:

AI’s race for better data has created an unexpected controversy: companies may be transforming centuries of human knowledge into AI training material while rare physical books disappear. The debate is no longer just about technology — it is about whether progress should come at the cost of cultural preservation.

A post recently exploded on X, gaining more than 23 million views, with a claim that sounds like something from a dystopian novel:

AI companies are bulk-buying rare books, cutting their spines off with high-speed machines, scanning every page, and shredding the originals afterward.

For anyone who sees books as more than just containers of words, the idea feels almost impossible to accept.

Why would anyone spend thousands of dollars on a rare book — only to destroy it?

The answer is that the judge said it’s legal 

“We shred rare books and offer NDAs so nobody finds out is a legitimate business model in 2026. 

So why are AI companies buying millions of books only to shred them? How is this even legal, and why is the public so furious? Let’s walk through the story. 

Why Are AI Companies Buying Millions of Books?

Books Have Become Valuable AI Training Material

Modern AI models require enormous amounts of high-quality text to learn language, reasoning, writing styles, and factual knowledge.

Books are among the richest sources of training data because they contain:

  • Long-form explanations
  • Expert knowledge
  • Historical records
  • Human creativity and storytelling

Older books are especially valuable because they were created before the explosion of AI-generated content, making them less likely to contain machine-written material.

For AI companies competing to build more powerful models, millions of books represent something more than paper — they represent billions of examples of human knowledge.

Project Panama: Anthropic’s Race for More Books

The scale of this demand became clear in 2024, when Anthropic launched Project Panama, an effort revealed through court documents to acquire and destructively scan millions of books to train its AI models behind Claude. The company reportedly spent tens of millions of dollars buying books and worked with scanning vendors capable of converting 500,000 to 2 million books within six months by cutting off bindings and scanning individual pages.

The Part That Shocked People: Why Destroy the Original Books?

Digitizing a book non-destructively is painstaking work. A human operator must manually turn every page beneath a high-resolution overhead camera, taking hours per volume to ensure pages aren’t damaged or blurred.

For companies operating at industrial scale, manual scanning is far too slow. The fastest way to digitize millions of pages is destructive scanning:

An industrial paper cutter slices the glued or sewn binding directly off the spine.

The loose sheets are fed into high-speed auto-sheet-feeding scanners capable of capturing hundreds of pages per minute.

Once the high-resolution digital copy is generated, the loose, un-bound paper is discarded or recycled.

To engineers focusing purely on data ingestion, the digital file is the book. The paper was merely a temporary vehicle holding the code. 

But to archivists and historians, this logic collapses when applied to rare or out-of-print works. When a rare historical text exists in only a few dozen physical copies worldwide, destroying one permanent artifact for a proprietary digital snapshot represents an irreversible loss.

Why Would This Be Allowed?

The biggest reason this practice could expand is simple: Companies believe it is legal.

AI companies argue that using books to train AI models is a form of transformative use, similar to how humans learn by reading existing works.

Their argument:

  • AI does not simply reproduce entire books.
  • Training creates a new system.
  • Learning from existing knowledge drives innovation.

Some court decisions involving AI-related book digitization have supported parts of this argument under fair-use principles.

Is AI Learning From Books — Or Extracting Them?

There’re 2 different views:

The Industry Says This Is How All Knowledge Progresses

Supporters compare AI training to human learning.

A writer reads thousands of books before creating a novel.

A scientist studies previous research before making discoveries.

An engineer learns from existing technology before inventing something new.

From this perspective, AI is doing something similar:

absorbing human knowledge and creating new capabilities.

Critics Say AI Is Different Because of Scale

The problem, critics argue, is not learning.

It is industrial extraction.

A person may read thousands of books over a lifetime.

An AI company can process millions of books in months.

The criticism is:

Human creators spent decades producing knowledge.

Companies are converting that knowledge into commercial AI systems that may compete with those same creators.

The question becomes:

Is AI expanding human knowledge — or capturing its economic value?

Why NDAs Made the Controversy Worse

“Digital Preservation” or “Keeping It Quiet”?

The destruction of books is only one part of the controversy.

The other is secrecy.

Critics point to:

  • Private data agreements
  • Non-disclosure agreements
  • Limited transparency about training sources

A viral post argued that companies could describe the process as “digital preservation” while keeping details hidden from the public.

That created a deeper concern:

If this is beneficial for society, why does it need to happen behind closed doors?

A Possible Turning Point: Could AI Learn Without Destroying Books?

The controversy has already started changing the conversation inside the AI industry.

In response to the viral discussion, Elon Musk commented that he had asked the SpaceX AI team to take a different approach:

I’ve asked the SpaceXAI team to preserve any rare books in a library and scan them the hard way vs just cutting off the spine and scanning.”

The message suggests a compromise:

AI companies can still collect valuable training data — but rare and historically significant books should be treated as cultural artifacts, not disposable sources of information.

The difference is not about whether books should be digitized.

Almost everyone agrees digitization can preserve knowledge.

The question is how that digitization happens.

For ordinary books, destructive scanning may simply be a practical efficiency choice.

For one of the last surviving copies of an important historical work, preservation may matter more than speed.

Musk’s response highlights the larger challenge facing the AI industry:

Can companies build more powerful AI systems while respecting the physical history behind the data they consume?

The Bigger Question: Who Owns Human Knowledge?

The rare-book controversy represents a much larger AI debate.

Every AI model is built from things humans created:

  • Books
  • Articles
  • Music
  • Images
  • Software
  • Scientific research

The central question is no longer just:

Can AI learn from human knowledge?

It is:

Who should benefit when human knowledge becomes the foundation of artificial intelligence?

If the data’s undisclosed, how would people even know his/her work’s been used?

Grace Wilson
I'm — a storyteller who turns trending news into practical tips.
I read and test the latest blogs and apps from top tech and travel sites so you don't have to.... I write about tech, travel, and music to help everyday people save money, live smarter, and enjoy life more—without the fluff. Real advice, real simple.