Anna's Archive argues that AI labs and their vendors are buying up rare physical books, guillotining their spines for sheet-fed scanning, and discarding the originals — a pipeline that's roughly 10x cheaper per page than non-destructive alternatives. They cite specific vendors advertising 'destructive scanning' tiers, rising prices on obscure academic titles, and invoices/photos as evidence. Their ask is narrow and actionable: donate to fund non-destructive scans of the single-digit-copy long ta
By submitting the Anna's Archive post to Hacker News, darccio amplified the argument that the AI training gold rush is creating an urgent preservation crisis for rare physical books. The 646-point score and 384 comments signal strong developer-community alignment that this destruction is a real problem worth surfacing.
The editorial pushes back on the developer reflex to shrug — 'books get scanned, LLMs get better, what's the loss?' — by arguing the loss is subtle but real along axes that matter to system builders, starting with provenance. Once the physical artifact is destroyed, the digital copy inside a training corpus becomes the sole record, with no way to re-verify against the original.
Anna's Archive, the shadow library that mirrors Library Genesis, Sci-Hub, and Z-Library, published a blunt post this week: AI companies are buying rare physical books at scale, slicing the spines off to run them through sheet-fed scanners, and discarding the remains. The pipeline is industrial. A book arrives, a guillotine trimmer cuts the binding, an automatic feeder pulls the loose pages through an optical scanner at hundreds of pages per minute, and the paper goes to recycling. The digital copy stays inside the buyer's training corpus. The physical artifact is gone.
The practice itself is not new — destructive scanning has been the cheap default in commercial digitization for a decade. What changed is the demand curve. Training-grade text is now valuable enough that AI labs and their vendors are competing with libraries and archivists for used-book inventory, and destructive scanning is roughly 10x cheaper per page than the non-destructive alternative. Anna's Archive points to specific vendors advertising "destructive scanning" as a service tier and to used-book marketplaces where obscure academic titles have jumped in price over the last 18 months. The post names names, links to invoices, and shows the trimmer photos.
The ask is narrow: fund non-destructive scans of the rare long tail — books that exist in single-digit copies worldwide — before the destructive pipelines get to them. Anna's Archive is soliciting donations to buy and scan targeted lists, cover-to-cover, with the spine intact and the book returned to circulation. They estimate the total cost to preserve the highest-risk tier of Western-language print at low single-digit millions of dollars.
The reflex reaction from developers is to shrug — books get scanned, LLMs get better, more knowledge ends up searchable, what's the loss? The loss is subtle but real, and it operates on three axes that matter to people who build systems.
First, provenance collapses. Once a book is chopped and its scan folded into a proprietary training set, the physical object that grounded the text is gone. You can't re-scan it at higher resolution, you can't verify a disputed passage, you can't check the marginalia, you can't do the paper-and-ink forensics that historians use to date and attribute works. For books that exist in a handful of copies — 19th-century technical manuals, small-press poetry, regional-language dictionaries, non-Anglophone scientific literature — the destructive scan becomes the source of truth. If the OCR mangled a formula or the scanner skipped a page, that error is now canonical. This is the same problem as losing the original of a photograph and keeping only a compressed JPEG, except the JPEG is inside a weights file you can't inspect.
Second, the incentive gradient is bad. Destructive scanning wins on price precisely because the externalized cost — the destroyed artifact — doesn't show up on the invoice. Libraries and university archives can't outbid AI-vendor procurement budgets in an open market for the same physical inventory, so the marginal rare book flows toward whoever is willing to pulp it. This is a straightforward tragedy-of-the-commons setup, and it resolves the way those always resolve unless someone deliberately funds the non-destructive path.
Third, and this is the part that should sting for anyone who's ever run a data pipeline: the labs doing this are, in effect, deleting their own future training data. A book scanned destructively today, at 2024-era OCR quality with 2024-era layout heuristics, is not the same asset as a book preserved and re-scannable at 2030 quality. The models will get better at reading footnotes, tables, diagrams, handwritten annotations, and multi-column layouts. The scans that exist right now will look, in five years, like the 2003 vintage of Google Books scans look today — full of column-break errors, missing plates, garbled math. Except in five years there will be no book to re-scan.
Community reaction has been sharp. The HN thread (646 points at time of writing) is unusually unified for that forum — most of the top comments are people asking where to donate, not debating the framing. A recurring subthread points out that the Internet Archive's non-destructive Scribe stations exist, are open-source, and cost roughly $10k to build; the bottleneck is operators and books, not technology. Another subthread notes that this problem is worse for non-English material, where a single destroyed copy can mean an entire minor literary tradition drops out of the digital record.
If you're building on top of a foundation model, the practical implication is that your training-data provenance story just got worse, not better. Any lab that runs a destructive-scan pipeline is producing corpus material that cannot be audited against a source, cannot be re-derived, and cannot be corrected. If you're on the compliance-and-copyright side of that conversation — building products in the EU AI Act's high-risk categories, or defending fair-use claims — the answer to "where did this training text come from and can we verify it?" quietly gets harder. Non-destructive scans, by contrast, leave a physical audit trail. Push your vendors on which pipeline they're buying from.
If you're running a RAG system on a technical corpus, the second-order lesson is about your own document ingestion. The cheap ingestion pipeline that throws away the source PDF after extraction is the same anti-pattern, at a smaller scale — you're optimizing for storage cost against a re-processing cost that you're going to pay later, at worse quality, with no way to recover the original. Keep the originals. Storage is cheap. Regret is expensive.
And if you have discretionary budget, this is one of the rare cases where a small donation buys a preservation outcome that is genuinely irreversible in the other direction. Anna's Archive publishes target lists. The Internet Archive runs Scribe stations. University libraries run their own programs. The dollar amounts to save a specific book from the trimmer are, on the scale of AI-industry spending, a rounding error on a rounding error.
The most likely outcome is that the destructive pipeline keeps running, most rare books get chopped, and in ten years we will have astonishingly capable models trained on a corpus whose physical substrate has been converted to recycled paper. A smaller but non-trivial fraction — the ones a handful of archivists and donors got to first — will be preserved intact and re-scannable forever. Which bucket a given book ends up in is being decided right now, one procurement invoice at a time, and the decision is essentially irreversible either way. That's the whole argument.
It is not a big deal. Since the invention of the printing press any important book has been duplicated by thousands, tens of thousands or even million of units.Just taking one of those and "destroying them"(it is not destroyed, a digital copy with the ability of doing millions of copies is
Nondestructive scanning can cost 10x as much. This is about cost. It is not about preservation. Google never destroyed the books it scanned. Amazon and Anthropic are attempting to save money. They are not considering whether or not a book is rare. They are treating books as a commodity. Rare books a
I was surprised to read that Anthropic (and probably other data / model companies) are doing this and it's extremely disappointing, as working towards the benefit of humanity is not an exclusive right / domain of theirs but rather, is a shared responsibility and mission carried by all
Big AI companies are leaving an easy opportunity on the table for establishing goodwill with the public.Just publicize a rare books vault where you put the older editions that aren’t in a lot of library catalogs. Use non-destructive scanning for those.Align yourself with the image of safeguarding so
Top 10 dev stories every morning at 8am UTC. AI-curated. Retro terminal HTML email.
I don't see any mention of Project Ocean - AKA Google books. Before AI they endevoured to digitize books in a massive online library. This inccluded rare and out of print books many which are archived at libraries. Because they had to preserve the books and return them in the condition they rec