F3 ships a file format that carries its own decoders as WASM

5 min read 1 source explainer
├── "Embedding WASM decoders in files solves the encoding-ossification problem that's killing columnar formats"
│  ├── future-file-format/f3 (GitHub) → read

The F3 authors (from the Vortex/Spiral and CMU database lineage) argue that file formats die from frozen encoding registries, not bad design. By shipping a sandboxed WASM decoder keyed by content hash in the footer, readers can decode novel encodings without waiting for every reader in the stack to be forked and re-released.

│  └── @tosh (Hacker News, 385 pts) → view

By submitting F3 to HN where it hit 385 points, tosh surfaces the thesis that the Parquet/ORC duopoly's static encoding registry is the core barrier to adopting modern research (ALP, FSST, BtrBlocks, Vortex). The community's strong response signals broad agreement that the encoding-ossification framing names a real pain point.

├── "Parquet's bottleneck is now CPU decoding, not I/O — the format itself needs to change"
│  └── top10.dev editorial (top10.dev) → read below

The editorial cites converging measurements from DuckDB's Mark Raasveldt and the Vortex team showing Snappy decompression saturates a memory channel before it saturates a column. On modern hardware, a codec designed for spinning disks is the wrong abstraction, making the format — not the storage — the bottleneck.

└── "Inventing entirely new formats (Lance, Nimble, Vortex) is the wrong response — extensibility within a stable format is what's needed"
  └── future-file-format/f3 (GitHub) → read

F3's pitch is implicitly a critique of the format-proliferation strategy: rather than asking the ecosystem to migrate to yet another competing format, embed the decoder so a single forward-compatible format can absorb new encoding research as it ships. The content-hash-keyed footer lets one format evolve indefinitely without breaking existing readers.

What happened

The `future-file-format/f3` repo hit 385 on Hacker News, surfacing a research file format with an idea that sounds heretical to anyone who's spent time in the columnar storage trenches: ship the decoder inside the file. F3 (Future-proof File Format) embeds compiled WebAssembly modules as part of the file footer. When a reader encounters a column encoded with a scheme it doesn't natively understand, it instantiates the bundled WASM decoder, hands it the compressed bytes, and gets columnar output back.

The thesis is blunt: file formats die not from bad design but from frozen encodings. Parquet shipped in 2013. Its encoding registry has barely moved since — RLE, dictionary, delta, byte-stream-split, and a handful of compression codecs (Snappy, Zstd, LZ4) bolted on at the page level. Meanwhile the academic and industry literature has produced ALP for floating point, FSST for short strings, BtrBlocks' cascading schemes, and Vortex's whole compressed in-memory format — and almost none of it is reachable from a production Parquet writer without forking every reader in your stack.

F3's authors come out of the same lineage as Vortex (Spiral) and the CMU database group. The repo is small, the spec is preliminary, and the benchmarks are deliberately narrow. What it ships is a proof that you can put a sandboxed decoder in a file, key it by a content hash in the footer, and have a vectorized reader call into it fast enough to matter.

Why it matters

The Parquet/ORC duopoly has been wobbling for two years. DuckDB's Mark Raasveldt and the Vortex team at Spiral have both published the same uncomfortable measurement: on modern hardware, Parquet's decoding is now the bottleneck, not its I/O. Snappy decompression saturates a memory channel before it saturates a column. When your scan is CPU-bound on a codec designed for spinning disks, the format is the problem.

The industry's response so far has been to invent new formats — Lance, Nimble (Meta), Vortex, BtrBlocks — each with a strictly better encoding story and a strictly worse adoption curve. Every new format relives the same trauma: writers ship in eighteen months, readers ship in five years, and the long tail of "my Spark cluster is on 3.3" never finishes. That's why Parquet still wins despite being demonstrably slower: it's the format every reader already understands.

F3 attacks this directly. If the decoder travels with the data, the format can evolve without coordinating a planet-scale reader upgrade. A team that invents a better timestamp encoding ships a WASM module, writes files that reference it by hash, and any conforming reader executes it. The encoding registry becomes content-addressed instead of standards-committee-addressed.

The immediate skepticism on HN was predictable and not wrong. WASM is not free. Instantiation costs, memory copies across the sandbox boundary, and the lack of direct SIMD intrinsics (until WASM SIMD lands universally) all add overhead that a hand-tuned C++ decoder doesn't pay. The F3 authors' counter is that for cold data — the 95% of your warehouse that gets queried once a quarter — the difference between "slightly slower decode" and "unreadable" is the only one that matters. For hot data, you cache the decoder and amortize instantiation across millions of row groups; the marginal cost approaches a native call.

The second skepticism is security. Executing untrusted code from a file you downloaded is, historically, how computers get owned. WASM's sandbox is the strongest argument F3 has here — no syscalls, no filesystem, no network, deterministic memory bounds — but "sandbox escapes are rare" is not the same as "sandbox escapes don't happen," and a query engine that runs arbitrary WASM per file is a bigger attack surface than one that doesn't. Expect this to be the central fight if F3 picks up.

What this means for your stack

If you operate a data lake today, nothing changes this quarter. F3 is research-grade. There are no production writers, no Spark connector, no Iceberg integration. Treat this the way you treated Arrow in 2016: interesting, watch the spec, don't bet on it.

What *should* change is how you think about format lock-in. The reason your warehouse is on Parquet is path dependence, not technical superiority — and the gap between Parquet and the state of the art is now wide enough that vendors are going to start exploiting it. Databricks, Snowflake, and ClickHouse all have proprietary internal formats that are meaningfully faster than Parquet on their own engines. The open question is whether the next open standard looks like "Parquet v3 with a bigger encoding registry" (the conservative path) or "a format where encodings are user-defined code" (the F3 path).

For query engine authors the calculus is harder. Supporting embedded WASM decoders means shipping a runtime (Wasmtime, Wasmer, or a custom one), defining the ABI for vectorized batches, and accepting that your scan path now includes a JIT. That's a significant lift for DuckDB, Polars, or DataFusion, and it changes the performance profile in ways that benchmarks need to capture honestly. Expect early adopters to be the engines that already have plugin architectures — DataFusion's UDF system maps cleanly onto this; DuckDB's extension model less so.

For practitioners building data infrastructure today, the actionable take is narrower: if you're choosing a format for a greenfield system in 2026, the question to ask vendors is no longer "what compression do you support" but "how do I add a new encoding without forking your reader." That question has no good answer in Parquet. It has a research-grade answer in F3 and a proprietary answer in Lance and Vortex. None of those are ready to bet a Series B on, but the *question* is the one that matters.

Looking ahead

The history of file formats is the history of decisions that calcify. SequenceFile begat RCFile begat ORC; Thrift begat Parquet; each transition took roughly a decade and left a sediment of unreadable old data. F3's contribution, even if F3 itself never ships at scale, is to name the meta-problem: a format whose encodings are a closed set will eventually be the slowest thing in your stack. The next open format that matters — whether it's F3, Vortex going open-governance, or a hypothetical Parquet v3 with a WASM escape hatch — will be the one that solves forward compatibility, not the one with the cleverest single encoding. Bet on the architecture, not the codec.

Hacker News 524 pts 121 comments

F3

→ read on Hacker News
vouwfietsman · Hacker News

Not sure why this got so many upvotes, also the landing page is not great, its better to look at the paper (see link below).Seems to be a columnar storage format that addresses some shortcomings in parquet. Thing is, though, that of all these formats the real winning feature is compatibility, which

gavinray · Hacker News

This bit is quite genius, rather than depend on a language-specific SDK/lib for working with the formats you can fallback to exported WASM methods if none exist: > "Each self-describing F3 file includes both the data and meta-data, as well as WebAssembly (Wasm) binaries to decode the da

sph · Hacker News

I don’t know what are people commenting on. I see a README with little to no information about what this is, what problems it solves, just links to its Flatbuffer description and a directory full of source code.What context am I missing?

largbae · Hacker News

This could use a bit more "why".Shortcomings of Parquet are mentioned as overcome by this, which ones? Certainly not wide tool support...Why should one leave Parquet or ORC for this structure?

zerobees · Hacker News

Some folks described it as genius. I guess it's my turn to play the role of an annoying HN skeptic: I find it somewhat silly. Data compression formats are secondary to what you're planning to do with the data once decoded. An audio file is completely different than an SVG image. An embedde

// share this

// get daily digest

Top 10 dev stories every morning at 8am UTC. AI-curated. Retro terminal HTML email.