Zitron argues the scaling-law era has ended: benchmark gains across GPT-5, Claude Opus 4.5, Gemini 3, Llama 4, and Grok 4 have compressed to low single digits, with GPT-5's launch flopping so badly OpenAI restored GPT-4o within 48 hours. He frames the ~$400B in 2025-2026 hyperscaler capex against under $40B in real generative AI revenue as the largest speculative-infrastructure-to-demand ratio in computing history.
Submitted Zitron's essay to HN where it hit 243 points in under a day, signaling endorsement of the deceleration thesis. The submission's traction reflects a growing developer-community appetite for sober accounting of AI capex versus capability returns.
Agrees that headline benchmark gains have flattened but contests the framing of 'slowing down.' Argues the real 2024-2026 progress is a ~20x drop in inference cost (from $60/Mtok to $3/Mtok at equivalent quality), which is a legitimate form of progress — just not the form the hyperscaler capex was built to monetize.
Ed Zitron's *Where's Your Ed At* dropped a 12,000-word essay titled "AI Is Slowing Down" that hit 243 points on Hacker News in under a day. The thesis is unsentimental: the scaling laws that made 2022–2024 feel like a vertical takeoff have flattened into something closer to a normal product curve, and the capital expenditure schedule was set during the vertical phase.
Zitron pulls together a year of release notes — GPT-5, Claude Opus 4.5, Gemini 3, Llama 4, Grok 4 — and shows that the headline benchmark deltas have compressed from double-digit jumps to low single digits, often within margin of error on the harder evals. He points to specific cases: GPT-5's launch landing with a thud loud enough that OpenAI quietly re-enabled GPT-4o for paying users within 48 hours; Anthropic's 4.5 release notes leaning on coding and computer-use rather than general reasoning gains; Meta's Llama 4 missing internal targets badly enough to delay the flagship.
The second half of the essay is a balance-sheet argument. Microsoft, Google, Meta, and Amazon have together committed something north of $400B in AI capex for 2025–2026. Revenue attributable to generative AI across the same vendors — once you subtract the rebranded cloud services and the GPU resale to OpenAI — is, by Zitron's reading, under $40B. He calls this the largest ratio of speculative infrastructure to demonstrated demand in the history of computing, and he's not obviously wrong.
The HN comment thread is the interesting read. The top reply isn't from a hater — it's from a builder who agrees on the trend but contests the framing. Their argument: gains haven't stopped, they've moved from "raw capability" to "capability per dollar at inference." Models that were $60/Mtok in 2024 are $3/Mtok in 2026 with equivalent quality. That's a real form of progress, it's just not the form the capex was underwriting.
This is the crux. Scaling-pilled investors bought a story about emergent capability jumps unlocking new product categories every six months. What they got was a steady cost-down curve on capabilities that existed in 2024. Cost-down curves are great for users and terrible for the people who built $50B GPU clusters to monetize the next jump.
The pushback from the accelerationist camp is also worth steelmanning. They argue Zitron is reading a plateau into what's actually a measurement crisis — the benchmarks saturated, the new evals aren't standardized, and the actual capability frontier (agentic workflows, long-horizon tool use, code generation at repo scale) is moving faster than any score can capture. METR's task-completion-length metric, for example, is still roughly doubling every seven months. That's not slowing down by any reading.
Both sides are partially right, which is what makes the essay land. Zitron is wrong that *progress* has stopped. He is probably right that *the specific kind of progress the capex was sized for* has stopped. The industry sold sovereign-wealth funds and pension boards on a story where GPT-6 would be to GPT-5 what GPT-4 was to GPT-3.5 — a step-change that creates a new $100B market. Nothing in the last twelve months supports that prior.
The community reaction split predictably. Zvi Mowshowitz pushed back on the methodology in a long quote-tweet, arguing Zitron cherry-picks benchmarks. Gary Marcus boosted it as vindication. The most interesting take came from a former OpenAI researcher who replied that internally, the conversation shifted in mid-2025 from "what does the next model unlock" to "how do we make the current generation profitable" — which, if true, is the tell.
If you're shipping a product whose roadmap assumes the next model will close your capability gap, you need a new roadmap. The honest planning assumption for the next eighteen months is that the model you can call today is approximately the model you'll be calling in Q4 2027, just five to ten times cheaper. Build for that world. The teams who win are the ones who treat the current frontier as the permanent floor and compete on integration, eval discipline, and product surface — not on waiting for a capability bailout.
If you're an infra buyer, the calculus inverted. A year ago, locking in three-year GPU contracts at premium rates was the conservative move because supply was the bottleneck. Today, with inference prices halving every nine months and three new chip families (Trainium 3, TPU v7, MI355X) shipping in 2026, long-dated commitments are the risky position. Spot, short-term, and provider-agnostic abstraction layers are the conservative play.
If you're a developer using these tools, the practical news is good. The price-performance curve continuing to bend means your AI-assisted workflows get cheaper, your eval budgets stretch further, and the gap between "frontier" and "open-weight runnable locally" keeps narrowing. The 8B and 70B open models that ship in late 2026 will do what GPT-4 did in 2024, and you'll run them on a Mac mini. That's the actual story under Zitron's headline.
The interesting question isn't whether Zitron is right about the slowdown — the data mostly supports him — but whether the financial system can absorb the recalibration without something breaking. Three hyperscalers have bet roughly a Manhattan Project's worth of capital on a scaling story that the last four model releases failed to confirm, and the bill comes due across 2027–2028 depreciation schedules. If the revenue doesn't catch up to the spend in eighteen months, the correction will be the actual news, and Zitron's essay will read in hindsight as the first clear signal. If revenue does catch up — through agents, through enterprise deployments finally clearing procurement, through cost-down expanding the market — then this is just the trough before the next leg. Either way, the era of pricing AI products on the assumption that the model will save you is over.
Lots of dismissive comments ITT, very few tackling the substance of the article.> AI Cannot Afford To Slow Down — It Needs $3 Trillion Or More In Revenue By End Of 2030 To Sustain Its ExistenceIs this true? With the total 2024 wages being 11.7 trillion USD [0], and nonfarm payrolls totaling 158,0
Ed is an interesting character. His financial analysis of the AI industry makes logical sense to me (though I am not knowledgeable enough to actually know if it is correct.) However, he seems to be so angry at AI in general, that he misses the obvious areas where LLMs are actually changing the State
Today Apple launched its revamped AI offering. Judging by several reports, Apple pays Google a mere billion dollars a year to operate it. Essentially just licensing the IP. Google are (allegedly) happy to turn over the right to operate and distill their models for only a billion a year.Consumer reve
One of the "smells" that gives away a quacky ranter is they speak in impassioned, "Why doesn't everyone understand this?" tones, but in fact their argument just doesn't flow. If Zitron's argument were as solid as he keeps saying it is, you would read it and underst
Top 10 dev stories every morning at 8am UTC. AI-curated. Retro terminal HTML email.
Number of active users on ChatGPT is at an all-time high. Number of tokens consumed on OpenRouter is at an all-time high. I'm not seeing the plateau.