The repo author demonstrates that rendering code to a PNG and shipping it as a vision input cuts Fable transpilation costs by ~60% with no reported quality regression. The argument is that flat per-tile image pricing dramatically undercuts per-token text pricing when the tile is packed with dense monospaced source code.
Submitted the repo to HN framing the trick as a genuine cost win, driving it to 251 points. The framing treats the technique as a legitimate optimization worth broad developer attention rather than a curiosity.
Argues the interesting story isn't the engineering — rendering text to PNG is trivial — but what the numbers expose about provider pricing structures. A 768x768 tile can encode 1,500-2,500 tokens of dense monospaced code for the price of a few hundred text tokens, meaning vision billing is effectively arbitraged by anyone willing to render their prompts.
Multiple commenters report replicating the trick on their own workloads with savings in the 40-70% range depending on code density and target model. This suggests the effect isn't specific to Fable transpilation but is a general property of code-heavy prompts against tile-priced vision endpoints.
A GitHub repo called `teamchong/pxpipe` hit 251 on Hacker News with an unusually cheeky claim: you can cut your LLM bill by ~60% on code-heavy prompts by rendering the code to a PNG, sending the image, and letting the model OCR it back into tokens on the other side. The author reports the numbers against a Fable (F#-to-JS/TS/Python/Rust) transpilation workflow, where prompts are dominated by long source files.
The pipeline is almost embarrassingly simple. Take the source, render it to a monospaced image with a syntax-highlighted theme, ship the image as a vision input, and prompt the model to treat the pixels as code. The model's built-in OCR reconstructs the text internally, then answers the actual question — refactor this, port it to Rust, explain this diff — against the reconstructed text. On the billing side, you're charged the vision model's flat image tile rate instead of the input-token rate for the equivalent text.
On the Fable workload the author benchmarks, that swap lands at roughly a 60% cost reduction, with no reported quality regression on the transpilation task itself. The HN thread is largely people testing the same trick on their own workloads and reporting numbers in the same neighborhood — 40% to 70% depending on how code-dense the prompt is and which model is on the other end.
The interesting thing about pxpipe isn't the engineering — rendering text to a PNG is a weekend project. It's what the numbers reveal about how frontier providers price vision versus text.
Most vision-capable models bill images at a flat rate per tile (a fixed pixel region, typically 512×512 or 768×768). A single tile costs the equivalent of a few hundred to a couple thousand text tokens, depending on the provider. But a 768×768 tile of dense monospaced code can encode 1,500–2,500 tokens of actual source, because you can fit a lot of 12pt characters into 590,000 pixels. The result: on long, code-heavy inputs, the effective per-token cost of vision input can be a fraction of the text rate.
This is not a model-capability discovery. It's a pricing arbitrage — providers priced vision tiles assuming natural images, and code renders happen to be an unusually information-dense payload for the same tile. The trick works today because nobody at the API pricing meeting was thinking about developers screenshotting their monorepos.
The community reaction on HN split along predictable lines. One camp is delighted: this is exactly the kind of gremlin optimization that made early AWS billing so fun to game. Another camp is nervous about quality regressions that don't show up in Fable's narrow benchmark. OCR is not lossless. Similar-looking glyphs (`l` vs `1`, `O` vs `0`, `rn` vs `m`) get confused. Long identifiers with mid-word underscores drop characters. Whitespace-sensitive languages — Python, YAML, F# itself — are one indentation error away from silent bugs. Several commenters reported working results on TypeScript and Rust but noticeable degradation on Python indentation and on any code with heavy Unicode.
The deeper structural point: this is what happens when you have two pricing dimensions (text tokens, image tiles) that meter the same underlying compute (attention over a context window) at different rates. The market will always find the cheaper unit. We saw the same thing with S3 request pricing and Lambda cold-start gaming — every mispriced axis becomes a cottage industry.
Before you shove this into production, three things worth internalizing.
First, the savings scale with input density, not intelligence. If your prompts are already short — chat turns, small function edits, structured JSON — the tile overhead can make images *more* expensive than text. pxpipe pays off on Fable-style workloads because you're sending 3,000-line F# files and asking for a full transpile. For a typical Copilot-style completion, forget it.
Second, assume OCR will silently corrupt your prompt about 0.1–1% of the time, and design around that assumption. That means: never use this path for anything where a single-character error is catastrophic (SQL, regex, cryptographic material, exact identifier matching for a refactor). Do use it for tasks where the model's output is going to be reviewed anyway — summarization, explanation, first-draft refactors that a human eyeballs before merge. Consider a two-pass check: image input for the bulk generation, text input for the verification step on the output.
Third, and most important: this is a temporary arbitrage. The pricing gap that makes pxpipe work is not a natural law. Anthropic, OpenAI, and Google can (and probably will) reprice vision inputs the moment enough of their revenue starts routing through image-encoded text. Build the optimization in a way you can rip out in an afternoon when the numbers flip. A wrapper module, one config flag, one billing dashboard alert on effective $/token — not a rewrite of your prompt layer that assumes images forever.
The honest read on pxpipe is that it's a wonderful hack and a bad long-term architecture. It exploits a pricing inefficiency that providers will close within a release cycle or two — either by tightening the image-token accounting (already visible in some providers' 'text density' surcharges for OCR-heavy images) or by explicitly rate-limiting the pattern. The people getting the most value out of this today are batch pipelines with predictable code-transformation workloads and a healthy tolerance for occasional OCR noise. Everyone else should treat it as a fascinating diagnostic of how vision pricing actually works, file it under 'good to know when the AWS bill spikes,' and move on.
So, just be careful with this, it very likely is switching to other less capable model hence the cost reduction. So looks like Fable but isn’t. So you are doing extra work when you could just switch the model back to opus 4.8 instead.
I tried the same thing last year (with openai models), back then it worked to reduce prompt tokens, but you needed way more completion tokens, ultimately more expensive (and slower) https://pagewatch.ai/blog/post/llm-text-as-image-tokens/
Ahhh my eyes the vibe coded readme
This seems like a pricing hack that burns resources, that when the loophole gets closed the price of OCR will have to rise?
Top 10 dev stories every morning at 8am UTC. AI-curated. Retro terminal HTML email.
In Gemini at least, if you look at how they process PDFs, they do an OCR and then feed the text + image to the model, without charging you for the text tokens (I believe).So my guess is that Claude’s backend is doing the same — so this hack is probably more of a loophole in token accounting that mig