AI Companies Are Buying Rare Books Just to Destroy Them
AI Companies Are Buying and Destroying Rare Books to Train Their Models—Here’s Why It Matters
The Big Picture: What’s Going On?
Imagine walking into a library, buying every book on the shelves, ripping the covers off, scanning the pages into a computer, and then throwing the original books into a shredder. That’s essentially what several major AI companies are doing right now—at an industrial scale.
Authors, historians, librarians, and book lovers are furious. They’ve discovered that tech giants are purchasing millions of physical books—including rare, antique, and out-of-print titles—only to destroy them while creating digital copies for AI training.
Important Callout: This isn’t just about bestsellers. Investigations reveal companies are targeting irreplaceable historical texts that exist nowhere else in the world. Once shredded, these books are gone forever.
How We Found Out: The Evidence Trail
1. The Anthropic Lawsuit Receipts
Court documents from a lawsuit against AI company Anthropic revealed a secret operation called "Project Panama."
- Contractors sliced the spines off millions of books
- Pages were fed into high-speed scanners
- The original books were shredded
- Internal goal: "Destructively scan all the books in the world"
- Internal note: "We don’t want it to be known that we are working on this."
2. The Amazon AirTag Investigation
Tech news site 404 Media tracked a shipment of nearly 1,000 rare books:
- A bookseller hid an Apple AirTag inside one volume
- The package traveled: California → Milwaukee → Colorado Springs
- Final destination: An Amazon-owned warehouse coded "VGT3"
- The building entrance featured a T-rex eating a book logo
- Employees described jobs where "all we do is scan books all day"
3. The Bookseller Whispers
Used book dealers across the U.S., U.K., and Europe report:
- Sudden bulk orders for obscure, non-fiction, pre-2022 titles
- Orders ranging from hundreds to 1 million books per transaction
- Buyers using shell companies or anonymous accounts
Why Are AI Companies Doing This?
You might wonder: Why not just download books from the internet? Two big reasons:
Reason 1: Legal Safety
- Shadow libraries (pirated book websites) have triggered massive copyright lawsuits from authors, artists, and publishers
- Buying a physical book gives the owner legal rights to that specific copy—including the right to destroy it
- Courts have ruled that scanning a lawfully owned book for AI training counts as "transformative fair use"
Reason 2: Data Quality
- The internet is now flooded with AI-generated text (low quality, repetitive, sometimes wrong)
- Training new AI on AI-written text makes models dumber, not smarter (a problem called "model collapse")
- Physical books offer:
- Human-authored, professionally edited language
- Long-form reasoning and narrative structure
- Specialized knowledge not found online
- Guaranteed pre-AI content (books printed before 2022)
Is This Actually Legal?
Yes—technically. Here’s the legal toolkit companies are using:
| Legal Concept | What It Means (ELI5) |
|---|---|
| First-Sale Doctrine | If you buy a physical book, you own that copy. You can resell it, lend it, or even shred it. The copyright holder can’t stop you. |
| Transformative Fair Use | Converting a print book into digital training data for AI counts as creating something "new and different" — so it’s allowed under copyright law. |
The Loophole: Critics argue companies are exploiting physical ownership as a legal workaround to build massive digital datasets without paying creators a cent. The law may not have anticipated industrial-scale book destruction.
How Many Books Have Been Destroyed?
Nobody knows the full number—but the scale is staggering.
- Anthropic alone planned to scan 500,000 to 2 million books in just six months (per court docs)
- Internal docs stated goal: "Destructively scan all the books in the world"
- 404 Media sources report single orders of 1,000 to 1 million books
- Booksellers describe "routine" bulk purchases of hundreds to thousands of titles
That’s potentially millions of unique, irreplaceable volumes turned into confetti.
Why This Should Worry Everyone
- Cultural Erasure – Rare books contain marginalia, unique bindings, printing history, and provenance that scans cannot capture
- No Digital Backup Guarantee – Scanned files can be lost, corrupted, or locked behind corporate walls
- Monopoly on Knowledge – A few tech giants control the digital versions of humanity’s printed heritage
- Precedent Setting – If legal, what stops them from buying and destroying archives, letters, or artworks next?
Summary
- AI companies (Anthropic, Amazon, likely others) are buying millions of physical books—including rare/antique ones
- They cut off spines, scan pages, and shred the originals to create training data for large language models (LLMs)
- Legal basis: First-Sale Doctrine + "transformative fair use" rulings
- Motivation: Avoid copyright lawsuits + get high-quality, pre-AI human text
- Scale: Potentially millions of books destroyed; goal stated as "all books in the world"
- Outcry: Authors, archivists, and bibliophiles call it cultural vandalism and a legal loophole exploit
FAQ
What is a "Large Language Model" (LLM)?
Think of it as a massive text prediction engine. It reads enormous amounts of human writing to learn patterns, grammar, facts, and reasoning—then generates new text based on what it learned. Examples: ChatGPT, Claude, Gemini.
What are "shadow libraries"?
Websites like Z-Library or Library Genesis that offer free, unauthorized downloads of copyrighted books. AI companies used these for training data—until authors sued them for copyright infringement.
Can’t libraries just digitize books themselves?
Libraries do digitize—but slowly, carefully, and with preservation in mind. They follow strict standards, keep the originals, and often make scans publicly accessible. Corporate scanning is fast, destructive, and the digital copies stay private.
Why does "pre-2022" matter?
Late 2022 is when ChatGPT launched and AI-generated text exploded online. Books printed before then are guaranteed human-written—making them "clean" training data.
Is anyone trying to stop this?
Yes. Lawsuits are ongoing (e.g., authors vs. Anthropic). Advocacy groups are pushing for legal reforms to close the "destructive scanning" loophole. Public pressure has forced some transparency—but the practice continues.
Final Thought: We’re watching a race between preservation and extraction. The books being shredded today survived wars, fires, and centuries—only to meet their end in a corporate scanner. The question isn’t just is it legal? It’s is this the future we want for human knowledge?