AI training is eating the long tail of print — literally

5 min read 1 source clear_take
├── "Destructive scanning is an urgent cultural loss that must be prevented through preemptive non-destructive digitization"
│  ├── Anna's Archive (annas-archive.pk) → read

The archive argues that AI labs are racing to feed pretraining corpora and treating rare physical books as disposable input material, guillotining spines and discarding originals once OCR is clean. They call for a coordinated effort to identify long-tail titles with a single known copy and photograph them non-destructively before labs get to them, framing the physical destruction as an irreversible loss of cultural heritage.

│  └── @darccio (Hacker News, 410 pts) → view

By submitting the piece with the framing 'let's scan rare books before it's too late,' the submitter endorses the urgency argument — that the window to preserve physical originals in non-destructive form is closing as industrial-scale guillotine pipelines expand.

├── "The AI copyright debate has focused on outputs, but the real overlooked harm is what training extraction costs the physical world"
│  └── top10.dev editorial (top10.dev) → read below

The editorial reframes the story away from familiar output-side debates (regurgitation, style mimicry, derivative works) toward the input side: a trained model is a lossy compression of a corpus, and in the destructive-scanning case the physical original is annihilated to produce that compression. This shifts the policy question from copyright infringement to cultural stewardship of irreplaceable artifacts.

├── "Destructive scanning of legally-owned books is a permissible format shift under fair use"
│  └── Judge William Alsup (via Bartz v. Anthropic ruling) (top10.dev editorial) → read below

Alsup's June 2025 ruling held that Anthropic's practice of buying used books, cutting the spines, and scanning them for training constitutes fair use — explicitly analogizing the guillotine step to ripping a CD you own. Under this view the physical destruction is legally irrelevant so long as the copy was lawfully acquired, leaving only the separate LibGen piracy claim to go to trial.

└── "Economics make destructive scanning the industry default and non-destructive alternatives non-viable at frontier-lab scale"
  └── top10.dev editorial (top10.dev) → read below

The editorial notes that guillotine-and-sheet-feed pipelines run at 60+ pages per minute and are roughly ten times faster and cheaper than any non-destructive method. For labs racing to add another trillion tokens to a pretraining mix, the cost math isn't close — which is why Anthropic reportedly ran the pipeline at warehouse scale, buying books by the pallet.

What happened

Anna's Archive — the shadow-library index that now functions as an unofficial training corpus for half the industry — posted a plea this week: help scan rare physical books, non-destructively, before AI labs get to them first. The blog post lays out a workflow the archive has been quietly building for months: identify long-tail titles with a single known copy, borrow or buy them, photograph every page with a camera rig, and upload. The urgency is not aesthetic. It's operational.

The current industrial standard for turning a book into training data is to cut the spine off with a guillotine, run the loose pages through a sheet-fed scanner at 60+ pages per minute, and throw the remains in a bin. That process — called "destructive scanning" — is roughly ten times faster and cheaper than any non-destructive alternative. For a lab racing to add another trillion tokens to a pretraining mix, the math is not close. Anthropic reportedly ran this pipeline at warehouse scale, buying used books by the pallet, feeding them through cutters, and discarding the physical originals once the OCR was clean.

What pushed this from a curiosity to a policy problem was Judge William Alsup's June 2025 ruling in *Bartz v. Anthropic*. Alsup found that training on legally purchased, then destructively scanned, books was fair use — explicitly framing the guillotine step as a permissible "format shift," analogous to ripping a CD you own. The training claims were dismissed. The separate claim about Anthropic's use of pirated books from LibGen is still headed to trial, but the destructive-scanning-of-owned-copies question is, for now, settled in the labs' favor.

Why it matters

Most of the debate around AI and copyright has been about the *outputs* — regurgitation, style mimicry, whether a model is a derivative work. This story is about the *inputs*, and specifically about what the inputs cost the physical world. A trained model is a lossy compression of its corpus; a destroyed book is a lossless deletion of a cultural artifact. The two are not symmetric, and one of them is irreversible.

The books at risk are not the ones you're picturing. Bestsellers have thousands of copies and are already digitized six ways from Sunday. The casualties are the long tail: a 1978 monograph on Assamese textile dyes with a print run of 400; a self-published memoir by a mid-century engineer; a regional cookbook that never made it out of one county. These titles have exactly the properties a training pipeline loves — unique tokens, non-web-scraped prose, high perplexity — and exactly the properties that make them fragile: one or two known copies, no economic case for a second scan, no institutional custodian sending a takedown.

The community reaction on Hacker News (410 points, 400+ comments as of writing) split predictably. One camp reads the Alsup ruling as a reasonable extension of first-sale doctrine: if you bought it, you can slice it. Another camp — including a surprising number of ML engineers — points out that the ruling treats books as fungible commodity inputs, which is fine for a paperback of *The Da Vinci Code* and catastrophic for the only extant copy of anything. The library-science people are the ones who sound the most alarmed, and they have been through this movie before: the Google Books project circa 2004 also ran on destructive scanning, and the settlement that followed took a decade to negotiate and left millions of "orphan works" in limbo.

There is also a supply-chain question the labs are not talking about. Used-book wholesalers — Better World Books, thrift-store liquidators, estate-sale aggregators — have started noticing bulk buyers with unusual profiles: no retail resale, no interest in condition, just weight-based orders for whole categories. Once a book enters that pipeline, it is functionally gone. No library ILL request will retrieve it. No future researcher will find a physical copy to check against the OCR errors that inevitably crept into the training set.

What this means for your stack

If you're building on top of frontier models, none of this changes your API bill tomorrow. But it should change how you think about corpus provenance in two concrete ways.

First, the "clean training data" story your vendor tells is now materially different depending on which vendor you ask. The Alsup ruling gave one specific lab legal cover for one specific acquisition pattern. It did not resolve the LibGen question, it does not bind courts in other jurisdictions, and it does not preempt state-level or international regulation — the EU AI Act's transparency requirements around training data are still ramping up, and "we bought and shredded the physical originals" is going to read very differently in Brussels than it did in San Francisco. If you're in a regulated industry and choosing a model provider, ask for a corpus provenance statement in writing. The labs that can produce one will start using it as a differentiator; the ones that can't will keep saying "proprietary."

Second, the fine-tuning corpus you build in-house is going to be more valuable, faster, than you expected. If the long tail of published knowledge is being destroyed to train general models, then domain-specific corpora that were painstakingly digitized non-destructively — medical journals, legal casebooks, engineering standards, internal wikis — become the last remaining source of ground truth for their domains. Treat your document ingestion pipeline as a strategic asset, not a cost center. Version it. Back it up. Note the source condition. Ten years from now, "we have the only clean scan of this" will be a moat.

If you have a personal library of technical books, the practical action is smaller and more immediate: photograph the ones that are out of print. The Internet Archive's Open Library workflow, or Anna's Archive's own volunteer pipeline, will take them. It is genuinely useful work, it takes about 30 minutes per book with a phone and a page-turning app, and it is the kind of thing that will look obvious in hindsight.

Looking ahead

The Alsup ruling is not the end of this fight — it's the opening move. Expect a second wave of litigation focused specifically on rare and unique works, likely brought by university presses and small publishers rather than the trade publishers who lost the first round. Expect the labs to start funding "preservation partnerships" with libraries as a PR hedge. And expect Anna's Archive to keep quietly assembling the parallel infrastructure — because if the last copy of a book has to end up in a training set, the argument for it also ending up in a public archive gets harder to refuse.

Hacker News 646 pts 384 comments

AI companies destroy physical books – let's scan rare books before it's too late

→ read on Hacker News
thread_id · Hacker News

I don't see any mention of Project Ocean - AKA Google books. Before AI they endevoured to digitize books in a massive online library. This inccluded rare and out of print books many which are archived at libraries. Because they had to preserve the books and return them in the condition they rec

cladopa · Hacker News

It is not a big deal. Since the invention of the printing press any important book has been duplicated by thousands, tens of thousands or even million of units.Just taking one of those and "destroying them"(it is not destroyed, a digital copy with the ability of doing millions of copies is

NishanStepak · Hacker News

Nondestructive scanning can cost 10x as much. This is about cost. It is not about preservation. Google never destroyed the books it scanned. Amazon and Anthropic are attempting to save money. They are not considering whether or not a book is rare. They are treating books as a commodity. Rare books a

ziyadb · Hacker News

I was surprised to read that Anthropic (and probably other data / model companies) are doing this and it's extremely disappointing, as working towards the benefit of humanity is not an exclusive right / domain of theirs but rather, is a shared responsibility and mission carried by all

odyssey7 · Hacker News

Big AI companies are leaving an easy opportunity on the table for establishing goodwill with the public.Just publicize a rare books vault where you put the older editions that aren’t in a lot of library catalogs. Use non-destructive scanning for those.Align yourself with the image of safeguarding so

// share this

// get daily digest

Top 10 dev stories every morning at 8am UTC. AI-curated. Retro terminal HTML email.