PHerc. 1667 is fully read. The bottleneck is now excavation, not ML.

4 min read 1 source clear_take
├── "The ML pipeline is now a solved, openly-shared technique — no longer secret sauce"
│  ├── @top10.dev editorial (top10.dev) → read below

The editorial argues that the team running an open AMA on Hacker News about segmentation, unwrapping, and ink detection is the clearest signal that the technique has matured from prize-winning novelty to known quantity. A year ago a single passage from a single scroll won $700,000; today a full scroll is read and the methodology is being explained on a public forum.

│  └── verditelabs (Hacker News) → read

As a Vesuvius Challenge team member, verditelabs is openly answering questions about the segmentation model, virtual unwrapping, and ink-detection ML on Hacker News. This transparency itself signals confidence that the pipeline is reproducible and no longer dependent on proprietary tricks.

├── "The real bottleneck is now archaeology, not ML — most of Herculaneum and the main library remain buried"
│  ├── @top10.dev editorial (top10.dev) → read below

The editorial reframes the story: the ML is a solved problem, so the interesting question is what blocks the next wave of discoveries. The answer is excavation — the existing scrolls came from what looks like a private collection, and the suspected main library of the Villa of the Papyri is still under volcanic rock.

│  └── @proee (Hacker News) → view

Points out that only about 20% of the Herculaneum archaeological site has been excavated and that the scrolls already recovered appear to come from a private collection rather than the main villa library. The implication is that the most consequential texts may still be physically inaccessible regardless of how good the reading pipeline becomes.

└── "This is a landmark humanities milestone — first complete Herculaneum scroll read in ~1,945 years"
  └── Vesuvius Challenge team (scrollprize.org) → read

The team frames PHerc. 1667 as the first Herculaneum scroll humans have read end-to-end since the eruption of Vesuvius in 79 AD — a near-two-millennia gap closed by combining synchrotron micro-CT, virtual unwrapping built on Brent Seales' foundational work, and crowd-trained ink-detection ML seeded by Nat Friedman's prize pool. The achievement is presented as a vindication of the open-prize, distributed-contributor model.

What happened

PHerc. 1667 — a carbonized papyrus scroll buried in the eruption of Vesuvius in 79 AD and dug out of the Villa of the Papyri at Herculaneum in the 1750s — has been read from beginning to end without ever being physically unrolled. The Vesuvius Challenge team announced the result at scrollprize.org/firstscroll. It is the first complete Herculaneum scroll humans have read in roughly 1,945 years.

The pipeline is now a known quantity: synchrotron micro-CT scans capture the scroll's internal geometry at micron resolution; a segmentation model traces the spiraled papyrus layers and virtually unwraps them into flat sheets; a separate ML model detects the ink — invisible to the naked eye because carbon ink on carbonized papyrus has almost no density contrast — by learning the subtle textural signature left by the carbon particles on the substrate. Nat Friedman seeded the prize pool, Brent Seales' group at Kentucky did the foundational work on virtual unwrapping, and a distributed community of contributors trained on released CT volumes.

The team member running the AMA on Hacker News (verditelabs) is openly answering questions about segmentation, unwrapping, and ink detection — which is the clearest signal yet that the technique is no longer secret sauce. A year ago, recovering a single passage from a single scroll won a $700,000 grand prize. Today, a full scroll has been read and the team is explaining how on a public forum.

Why it matters

The interesting story here is not the ML. The ML is a solved problem now. The interesting story is what the solved problem reveals about the remaining bottleneck.

As one HN commenter (proee) pointed out: only about 20% of the Herculaneum archaeological site has been excavated, and the scrolls already in hand came from what looks like a private collection, not the main villa library. The main library — if it exists, and the architecture of the Villa of the Papyri strongly suggests it does — is still under volcanic rock. We have a technique that can read whatever we dig up. We do not have permission, funding, or political will to dig up more.

This is a familiar pattern for anyone who has shipped infrastructure inside a large organization. You spend two years building the platform that finally unblocks the team. The team is unblocked. Then you discover that the actual constraint was never the platform — it was the procurement process, or the legal review, or the one VP who has to sign off. The work was real, the win was real, and the next obstacle is somewhere completely different.

For the Vesuvius Challenge, the next obstacle is the Italian Ministry of Culture, the modern town of Ercolano built directly on top of the unexcavated villa, and the very reasonable concern that nineteenth-century excavation techniques destroyed irreplaceable material. The engineering is done. The diplomacy hasn't started.

There's a second order lesson worth noting. This was a prize-funded effort with public data, not a museum-funded internal project. Classical scholarship had access to these scrolls for 270 years and produced fragmentary, controversial transcriptions. A coalition of ML researchers — most of whom had never touched a piece of papyrus in their lives — produced a complete reading in roughly four years of concentrated work, on a budget that wouldn't fund a single tenure-track position. The discipline-crossing was the point. Domain expertise told us where to look. Open data and prize incentives told us how to find people who knew how to look.

What this means for your stack

Three practical takeaways.

First, the "computational archaeology" stack is now real and reusable. The same CT-plus-segmentation-plus-signal-detection pipeline applies to any scenario where you need to read something physical without destroying it: damaged hard drives, sealed historical documents, layered paintings, wrapped mummies, electronic waste with bonded chips. If you have a non-destructive imaging problem with weak signal, the Vesuvius Challenge codebase and trained models are the closest thing to a reference architecture that exists.

Second, this is a clean argument for the prize-funded research model in narrow technical domains. The Vesuvius Challenge worked because the problem was well-specified (read the scroll), the data was releasable (CT volumes have no commercial sensitivity), and the evaluation was unambiguous (you can read the Greek or you can't). When those three conditions hold, a prize structure outperforms grant cycles by an order of magnitude on wall-clock time. They don't hold for most research, but they do hold for a surprising number of digitization and recovery problems sitting in museum basements right now.

Third — and this one stings — note how much of the modern stack made this possible without being designed for it. Synchrotron beamlines were built for materials science. GPU clusters were built for graphics, then for crypto, then for AI. Open-source segmentation tools were built for medical imaging. Nobody set out to build a Herculaneum-reading machine. The reading happened because general-purpose infrastructure became cheap and good enough that a small, motivated group could compose it into a domain-specific pipeline. That's the same dynamic that put Half-Life 2 in a browser tab last week.

Looking ahead

The next milestone isn't another scroll. It's the political fight over excavating the rest of the Villa of the Papyri. If that fight goes the way the engineers want, we may, within a decade, recover an entire classical library — possibly including lost works of Aristotle, Sophocles, and Sappho that humanity has been missing for two millennia. If it doesn't, we'll have the world's best ink-detection pipeline and nothing new to point it at. The boring institutional question is now the only question that matters.

Hacker News 1666 pts 364 comments

A Herculaneum scroll has been read for the first time

→ read on Hacker News
verditelabs · Hacker News

I am on the vesuvius challenge team that did the segmentation, unwrapping, and ink detection, so feel free to ask any questions.

codeulike · Hacker News

Lets reflect on Aristocreon, in about 200 BC, putting their thoughts down on a scroll. They would be aware that the scroll might be kept in a library for some time. Maybe they could have imagined it surviving for 300 years. But they never would have imagined that in 300 years a volcano might destroy

9dev · Hacker News

Every time you feel depressed by the state of tech, and how so many intelligent people seem to work on forcing ever more ads down people's throats (a common trope around these parts), remember that projects like this do exist too!There are lots of very smart folks working on incredible things,

proee · Hacker News

Only about 20% of the Herculaneum site has been excavated, so there is high probability that more scrolls exist. The current scrolls were not part of the main library, but more of a private collection at the time.So imagine how cool it would be to find a full library with thousand of scrolls across

melicerte · Hacker News

Did anyone notice that anonymous donators[1] have the picture of Larry David, and the link points to the Curb Your Enthusiasm - Anonymous Donor Pt2[2] episode?So geeky, so cool !- [1] https://scrollprize.org/#sponsors- [2] https://www.youtube.com/watch?v=JqrJ4wGid4Y

// share this

// get daily digest

Top 10 dev stories every morning at 8am UTC. AI-curated. Retro terminal HTML email.