The AI Boom Is Burning History: Why We Must Scan Rare Books Before They Vanish

You’ve probably used an AI tool today. Maybe you asked it to summarize a document, write an email, or generate a picture. You didn’t think twice. But that AI was trained on something—and that something is increasingly made of human history. Literally. The books that survived wars, fires, and censorship are now being fed to algorithms, and many of them are being destroyed in the process.

This isn’t a hypothetical. AI companies are buying up physical libraries—rare manuscripts, out-of-print treatises, fragile volumes that have sat on shelves for centuries—and ripping them apart to create training data. The irony is brutal: the very technology that promises to preserve knowledge is eating its physical containers alive.

Every time you ask an AI a question, you’re standing on the ashes of a book that can never be held again.

Let’s be clear about what’s being lost. A scan captures the words. It does not capture the book’s soul. The marginalia—the notes scribbled by a 17th-century scholar—vanishes. The binding, the paper texture, the printing errors, the provenance marks that tell us who owned it and where it traveled—all of that is gone the moment the spine is cracked and pages are fed to a scanner. The digital file is a flat copy. The artifact was a living witness.

And here’s the twist that makes this story sickening: the only way to save the text is to destroy the object. To scan a book, you often have to break its binding. You have to flatten it, press it, expose it to light. The process is, by nature, destructive. So we’re left with a choice: let the book rot in a basement, or sacrifice it to preserve its words. But that’s a false choice. Because the real tragedy is that we didn’t start scanning these treasures decades ago, when they were still intact.

Now the AI gold rush is accelerating the destruction. Companies like OpenAI, Meta, and Google are licensing and purchasing entire libraries to feed their models. They aren’t doing this to preserve culture; they’re doing it to build better chatbots. The copyright debates are a distraction. The real issue is infrastructure: rare books are being treated as raw material, like oil or coal, to be burned for algorithmic profit.

The models you rely on are built on data extracted from physical archives that are now disappearing. You are implicated. Every time you use an AI, you’re voting for this trade-off—even if you didn’t know it existed.

There is still time, but barely. Digitization projects exist, but they’re underfunded and slow. Libraries are sitting on collections that will be gone within a decade if we don’t act. The people who love these books—the archivists, the historians, the librarians—are screaming into the void. And the AI companies? They’re moving fast, buying first, asking no questions.

So what do we do? We prioritize. We scan the rarest, the most fragile, the ones that cannot be replaced. We throw money at every library that will let us. We make it a cultural emergency, not a technical footnote. And we demand that AI companies contribute to preservation—not just extraction.

But we also need to be honest with ourselves. A scan is a shadow. It’s better than nothing, but it’s not the same. The physical book is a unique artifact, and once it’s gone, it’s gone forever. No backup file can bring back the smell of old paper, the weight of a leather cover, the ink that was pressed by a hand that died three hundred years ago.

We’re not just fighting for words. We’re fighting for the material evidence of human thought. And if we lose that, we lose a part of ourselves.

Don’t let the AI industry burn our libraries to keep its servers warm.

Scan what you can. Save what you love. And maybe, just maybe, the next generation will still be able to touch a book that changed the world—not just read it on a screen.

FAQ

Q: Isn't a digital scan enough to preserve the book's content?

A: No. A scan captures the text but loses the marginalia, binding, paper texture, provenance marks, and physical history that make a rare book unique. The artifact itself—the object—is irreplaceable.

Q: What can I do if I'm not a librarian or historian?

A: Donate to digitization projects, pressure AI companies to fund preservation, and advocate for public scanning initiatives. Most importantly, stop pretending this isn't your problem. If you use AI, you're part of the demand pipeline.

Q: Is the copyright debate the real issue here?

A: No. The copyright fight is a sideshow. The core crisis is that rare books are being physically destroyed to feed AI training pipelines. We're losing the artifacts themselves, and no legal ruling can bring them back.

📎 Source: View Source