Publishers Created the AI Book-Shredding Crisis. Now They’re Acting Surprised.

You probably imagined AI training data as something clean — digital archives, licensed databases, the gentle hum of servers. You didn’t picture a warehouse where someone feeds a 300-year-old botanical text into a paper shredder, because the printer’s ink is the only thing standing between a billion-dollar lawsuit and a cheaper scanner.

That’s the reality now. AI companies are physically destroying rare books to train their models. Not scanning them carefully. Not preserving them. Shredding them. And it’s perfectly legal — because the publishers who sued to protect their digital turf made this the only sane economic choice.

Publishers demanded a toll. AI companies built a bridge through the fire.

Here’s the chain of events that should make every bibliophile rage: Publishers sued AI companies for training on shadow library data. They wanted a cut of the AI gold rush. They expected negotiation — pay us, and we’ll let you use our books. Instead, AI companies discovered the “analog hole.” Buy a physical copy for $5, pay someone $25 to destructively scan it, and you’ve legally extracted the content without licensing fees. The physical book is gone. The data lives forever.

“You can reprint a bestseller. You can’t replace the last three copies of an 18th-century botanical text once someone shreds them for training data. And the judge said it’s legal. So it’s going to accelerate.”

That’s not a dystopian theory. It’s happening now. The comment thread on the original tweet names specific titles — but the real horror is that no one is tracking which books are being destroyed. Librarians who spent decades preserving knowledge are watching their life’s work become paper pulp for a statistical model that will never acknowledge the original.

Let’s be clear about who’s responsible here. The publishers. They could have created a reasonable digital licensing framework. They could have lobbied for updated copyright laws that protect physical artifacts. Instead, they saw a cash cow and sued. The result? Their own greed created the incentive to destroy the very books they were supposed to steward.

One commenter dismisses this as “too precious about old things” — arguing that the knowledge is preserved even if the paper isn’t. That’s a comfortable lie. A scanned book loses context: marginalia, binding, paper quality, the tactile history of use. We digitize to preserve content, but we destroy the artifact. And the AI companies don’t care about the artifact. They care about tokens.

This isn’t a story about technology. It’s a story about perverse incentives. The publishers wanted to charge a toll on the digital highway. The AI companies built a tunnel right through the library. And the books burned for fuel.

The real tragedy isn’t the destruction. It’s that we could have prevented it — and we chose not to.

So what now? Public pressure could force AI companies to certify they only use digitally licensed or ethically scanned data. But that requires transparency, and transparency is expensive. Until then, every time you see AI output that mentions a rare historical text, ask yourself: Did a book die for this sentence?

FAQ

Q: Why don't AI companies just license the digital rights from publishers?

A: Because publishers demanded exorbitant fees — effectively extorting the AI industry. The analog hole (buying a physical copy for $5 and destructively scanning it) is legally cheaper, even though it destroys the physical artifact.

Q: Does this only affect rare books under copyright?

A: Yes, primarily. Public domain works are already freely available online. But the most valuable training data for AI is often recent, copyrighted material — and those are the books being shredded. Rare older texts can also be caught in the crossfire if they're out of print but still under copyright.

Q: Is there a way to stop this without harming AI progress?

A: Yes: require AI companies to disclose their training data sources and certify that no physical artifacts were destroyed. If the public demands ethical sourcing, the cost of destroying books becomes a reputational liability. But that requires transparency regulation, which publishers initially opposed.

📎 Source: View Source