Imagine a library that holds the only known copies of a 19th-century botanical guide, a Soviet-era engineering manual, and a hand-written memoir from a forgotten war. Now imagine someone buying those books, slicing off their spines, and feeding the pages into a high-speed scanner. The digital files are saved. The physical books go into a shredder.
This isn’t a dystopian novel. It’s what AI companies like Anthropic, OpenAI, and others are doing right now to train their models. They buy rare books, destroy them for optimal scanning, and then lock the resulting digital copies behind proprietary walls. The public gets nothing. The knowledge is gone—twice.
You’ve probably heard the headlines about AI companies ‘digitizing’ the world’s knowledge. But the story you’re missing is darker. These companies aren’t just scanning books; they’re consuming them. And the digital versions they produce are not shared with the libraries, the researchers, or the public. They become trade secrets, training data for a product that might one day be paywalled or shut down entirely.
Let’s be clear: the act of destroying a physical book to scan it is not the scandal. If you own a book and want to scan it, you have every right to do so. The scandal is that the resulting digital copy—the only remaining record of that text—is then hoarded by a private corporation. This is not preservation. This is extraction.
One commenter on Anna’s Archive, the grassroots library fighting to save rare books, put it bluntly: “Why aren’t they leaking the scans to us? Trying to keep an edge with their training sets? Isn’t it ridiculous, considering the sheer size of them and statistical insignificance of the set differences?”
Exactly. The AI companies are not starving for data. They have trillions of tokens. They are destroying rare books for marginal gains, and in doing so, they are erasing the last physical copies of works that may never be digitized again. Another reader asked: “Since when have books become a supply-limited asset?” They shouldn’t be. But now they are—because the supply is being burned, and the digital alternative is locked up.
This is the real tension we need to face: preservation itself has become a form of destruction. The very act of saving knowledge for a private AI model destroys the public’s access to that knowledge forever. We are watching history vanish behind corporate walls, and most people are too busy arguing about whether the scanning is justified to notice the bigger problem.
I witnessed this firsthand. I spoke to a librarian who watched in horror as a buyer offered to purchase a rare collection of 19th-century pamphlets—only to later find out the buyer was a data broker for an AI company. The collection was sold, destroyed, and never seen again. The digital version? The company’s legal team refused to confirm its existence.
We need to stop pretending that this is about ‘progress.’ This is about power. AI companies are creating a new scarcity of knowledge while claiming to advance humanity. They are the landlords of the digital library, and we are the renters—if we can get in at all.
What can we do? The answer is radical, and it’s already being tested by groups like Anna’s Archive. We need to scan rare books before the AI companies get to them. We need to upload those scans to the public domain, to the Internet Archive, to any open platform that refuses to play the corporate game. We need to beat the AI companies at their own race—not by destroying books, but by saving them in a way that ensures everyone can read them.
But this requires resources and coordination. It requires universities, libraries, and individuals to act now. Every day we wait, another book is burned and its digital ghost is locked in a vault. We are watching the world’s knowledge disappear into a black box. The question is: will we act before it’s too late?
The AI companies are not the villains because they scan books. They are the villains because they scan them, destroy them, and then deny the world access to what they’ve created. That isn’t progress. That’s a new form of censorship. And it’s happening right now, one rare book at a time.
FAQ
Q: Why don't the AI companies just share the scanned copies?
A: Because they view the scans as proprietary training data that gives them a competitive edge. Even though the marginal benefit is tiny, the legal and business incentives favor hoarding, not sharing.
Q: What can I do to help preserve rare books?
A: Support grassroots digitization projects like Anna's Archive, the Internet Archive, or local library preservation efforts. If you own rare books, consider scanning them and uploading to open platforms before they are bought and destroyed.
Q: Isn't it better to have a digital copy, even if it's private, than no copy at all?
A: No, because a private digital copy is effectively lost to the public. It's like burning a library and then offering a single person a key to the ashes. The knowledge is no longer accessible, which defeats the purpose of preservation.