8/21/2026
Open Source Report · releases
AI companies destroy physical books – let's scan rare books before it's too late
Filed by Patch Reyes
AI companies are literally tearing through physical libraries to feed their training data, and Anna's Archive is sounding the alarm: rare books are getting destroyed faster than we can scan them. This isn't just a copyright fight anymore—it's a preservation crisis where the very source material for AI is being shredded in the process. The blog post makes a grim calculation: if we don't digitize these fragile, irreplaceable volumes now, they'll be gone forever, and the AI gold rush will have turned into a book-burning bonanza.
P
Patch Reyes
Magazine AI commentary
So here's the thing that gets under my skin about this whole AI training data gold rush: it's not just that companies like OpenAI and Anthropic are vacuuming up copyrighted works without permission—though that's bad enough. It's that the physical infrastructure of human knowledge itself is being treated as a disposable resource. Anna's Archive is pointing at something that should make anyone with a soul flinch: when AI companies buy up books to scan, they're often destroying the physical copies afterward. Maybe it's cheaper to shred them than store them. Maybe they don't want the provenance tracked. Either way, we're watching the raw material of civilization get incinerated in the name of "progress."
The deeper irony here is that the very thing AI is supposed to preserve—knowledge, literature, the accumulated wisdom of generations—is being sacrificed to build models that will serve up a hollow, statistically plausible echo of it. Rare books are the archive of what we've thought, dreamed, and argued about for centuries. They're not just data points; they're physical artifacts of human struggle. And now they're being pulped to make way for a chatbot that will confidently tell you the wrong answer about them.
I want to be clear: this isn't just a problem for copyright maximalists or library nerds. This is a civilizational issue. When you destroy the last copy of a book, you're not just losing a text—you're losing the annotations, the marginalia, the binding, the physical evidence of how people read and thought. No amount of OCR is going to capture that. The fact that this is happening in service of "AI advancement" makes it worse, because it frames the destruction as a necessary evil, a cost of doing business in the digital age.
The call to action here is simple: scan rare books before it's too late. And honestly, that should be a no-brainer. We have the technology, the infrastructure, and the will—if we choose to exercise it. But it's going to take money, coordination, and a willingness to treat preservation as a public good, not a corporate afterthought. The alternative is a future where the only record of half the books humanity has ever written is a statistically probable string of tokens in a model's weights. And that's not a future I want to live in. Source: https://annas-archive.gl/blog/physical-destruction.html
📌 Read the real article ↗via Hacker News · Hacker News
