Popular Posts

AI Giants Buy Priceless Books Only to Destroy Them

AI Companies Are Buying Rare Books by the Millions—Then Destroying Them to Train Chatbots

The Short Version

Imagine a library where every book gets its spine chopped off, run through a high-speed scanner, and then shredded—forever. That’s not science fiction. It’s happening right now. Major AI companies like Anthropic and Amazon are buying up physical books—including rare, antique, and out-of-print treasures—by the hundreds of thousands. They scan the pages to feed their large language models (LLMs), then destroy the originals.

Authors, archivists, historians, and book lovers are furious. They call it cultural vandalism. The tech companies call it "data acquisition." And remarkably, it’s all perfectly legal.


How We Found Out: The "Project Panama" Revelation

The scandal broke into public view through a lawsuit and some good old-fashioned investigative journalism.

1. The Lawsuit Receipts

In early 2026, court documents from a lawsuit against Anthropic (the company behind the AI assistant Claude) revealed an internal project code-named "Project Panama." The plan? Buy millions of physical books, hire contractors to slice off the bindings, feed pages into industrial scanners, and shred the remains.

Internal Anthropic Document:
"Project Panama is our effort to destructively scan all the books in the world. We don’t want it to be known that we are working on this."

2. The AirTag Sting

Not long after, tech news outlet 404 Media conducted a wild experiment. A rare bookseller received a suspicious order for nearly 1,000 antique volumes. The seller agreed to hide an Apple AirTag inside one book.

The tracker pinged a journey:

  • California → Milwaukee → Colorado Springs → Amazon warehouse "VGT3"

The building’s entrance reportedly sports a logo of a T-rex eating an open book. Employees describe "nice jobs where all we do is scan books all day."


Why Are They Doing This? (The "Good" Reasons, According to Them)

AI companies didn’t start buying physical books for fun. They hit two massive walls with digital data:

Problem Digital Approach (Shadow Libraries) Physical Book Approach
Legal Risk Massive copyright lawsuits from authors, publishers, artists First-Sale Doctrine: You own the book, you can destroy it
Data Quality Internet is now full of AI-generated text (garbage in, garbage out) Pre-2022 books = guaranteed human-written, edited, coherent, long-form
Availability Millions of niche/out-of-print books don’t exist digitally Physical copies sit in warehouses, estate sales, used bookshops

In plain English:
Training an AI on today’s internet is like teaching a child using only essays written by other children who copied from Wikipedia. Training on physical books from 1995? That’s like giving the child a curated library of Pulitzer winners.


The Legal Loophole: How Is This Allowed?

[!IMPORTANT]
Two legal pillars make this 100% legal (for now):

1. First-Sale Doctrine (Copyright Act)

"If you legally buy a physical book, you own that specific copy. You can resell it, lend it, annotate it—or cut it up and shred it."
Courts have upheld this for decades. It’s why used bookstores exist.

2. Transformative Fair Use

"Converting a lawfully owned print book into digital training data for an AI model is ‘transformative’—it creates a new purpose (statistical language modeling) rather than substituting for the original (reading)."
Several court rulings have leaned this way in early AI copyright cases.

Critics argue:
Companies are exploiting a loophole meant for personal use (lending a book to a friend) to build industrial-scale datasets—without paying a cent to authors or publishers.


How Many Books Are We Talking About?

Nobody knows the exact number. But the clues are terrifying.

Source Estimate
Anthropic (Project Panama) 500,000 – 2,000,000 books in a single 6-month vendor contract
404 Media sources Anonymous bulk orders of 1,000 to 1,000,000 books per purchase
Used booksellers (US, UK, EU) Routine bulk orders of hundreds to thousands of "random" non-fiction/out-of-print titles
Anthropic internal goal "Destructively scan all the books in the world"

Sales data shows a historic spike in pre-2022 physical book purchases starting roughly when LLM training went industrial. Booksellers quietly confirm: the mystery buyers are AI firms.


What Gets Destroyed? Not Just Bestsellers

This isn’t just extra copies of Harry Potter or The Da Vinci Code.

The chopping block includes:

  • Rare first editions (1800s, early 1900s)
  • Out-of-print academic monographs (only 200 copies ever printed)
  • Regional histories, local cookbooks, niche technical manuals
  • Illustrated volumes with plates, maps, fold-outsscanners often miss these details
  • Books in endangered languages with no digital backup

[!WARNING]
Once shredded, these physical artifacts are gone forever.
No library holds them. No archive preserved them. The only "copy" left is a statistical ghost inside a neural network.


The Human Reaction: "Evil Incarnate"

"Evil incarnate."Michael Burry (of The Big Short fame), on X (formerly Twitter)

Who’s angry?

  • Authors — their life’s work fed into a machine without consent or compensation
  • Archivists & Librarians — centuries of preservation culture violated
  • Bibliophiles & Historians — irreplaceable cultural memory turned into confetti
  • Legal scholars — warning that "legal ≠ ethical"

Tech defenders say:

  • "We’re preserving the information, not the object."
  • "Digital surrogates last longer than paper."
  • "This accelerates AI that helps everyone."

Critics reply:

  • A scan ≠ the book. (Try reading a 16th-century marginalia note on a glitchy OCR scan.)
  • "Preservation" doesn’t require destroying the original.
  • The benefit is private (corporate IP); the loss is public (shared heritage).

Summary: The Core Conflict in 3 Sentences

  1. AI companies need massive amounts of clean, human-written text to make their models smarter—and the open internet is now too polluted with AI garbage.
  2. They found a legal hack: buy physical books (protected by First-Sale Doctrine), destructively scan them (protected by transformative fair use), and build proprietary datasets without paying creators.
  3. The cost is irreversible cultural destruction: rare, antique, and unique books—many with no other copies in existence—are being turned into pulp to feed corporate AI models.

FAQ: Your Burning Questions, Answered Simply

Can’t they just scan the books without destroying them?

Yes. Non-destructive scanning exists (cradle scanners, overhead cameras). It’s slower and more expensive. Companies choose destructive scanning because it’s industrial-scale fast and cheap—high-speed sheet-fed scanners require loose pages.

Don’t libraries already have digital copies of most books?

No. Estimates suggest over 80% of 20th-century books are out of print and not digitized. Many rare/academic books exist in only a handful of physical copies worldwide. Once those are shredded, the physical artifact is extinct.

Is anyone trying to stop this legally?

Yes. Lawsuits are ongoing (e.g., authors vs. Anthropic). But current copyright law focuses on copying/distribution, not destruction of owned property. New legislation would be needed to ban destructive scanning of rare books—and that’s a steep political climb.

What happens to the scanned data?

It becomes training tokens inside proprietary LLMs (like Claude, potentially Amazon’s Olympus, others). The books are not made publicly available. The data becomes corporate intellectual property.

Can I do anything as a regular person?

  • Donate rare books to libraries/archives instead of selling to unknown bulk buyers.
  • Support legislation for "cultural heritage protection" that restricts destructive digitization of rare materials.
  • Ask your representatives why First-Sale Doctrine allows industrial destruction of unique cultural artifacts.
  • Spread awareness—this has flown under the radar because it looks like normal book buying.

Final Thought

We are watching the largest deliberate destruction of printed cultural heritage since the wartime bombings of European libraries—except this time, it’s voluntary, legal, and funded by the world’s richest companies.

The books don’t fight back. They don’t sue. They just turn to dust.

The question isn’t "Can they do this?"
The question is: "Should we let them?"


Leave a Reply

Your email address will not be published. Required fields are marked *