Tao: AI solving famous math problems misses the whole point

4 min read 2 sources clear_take
├── "Using famous math problems as AI benchmarks destroys their signal value and misaligns AI labs with the mathematical community"
│  ├── Terry Tao and collaborators (mathandai.org / Tao's blog) → read

Tao argues that famous open problems were lighthouses — their value was signaling that new ideas, abstractions, and methods existed, which the community would then digest through talks, simplifications, and eventual textbook presentations. When AI labs treat these problems as scoreboards to knock down for PR, the answer becomes the product instead of the byproduct, destroying the very signal function the problems were meant to serve.

│  └── @meredydd (Hacker News, 1194 pts) → view

By submitting the piece and driving it to #1 with 1,194 points, the submitter amplified Tao's framing that AI companies' benchmark-chasing goals are 'severely misaligned' with mathematics as a discipline. The framing treats this as part of broader alignment problems across scientific and creative professions.

└── "This problem isn't new — humans have already produced incomprehensible 'valid' proofs that damaged the field"
  └── @tmhn2 (Hacker News) → view

Points to Mochizuki's abc conjecture proof as a precedent: a human working in relative isolation dumped an incomprehensible manuscript on the community, technically claiming a landmark result while providing nothing the community could actually digest. If humans can already break the signal function, AI merely industrializes an existing failure mode rather than introducing a novel one.

What happened

On September 11, Terry Tao and a group of collaborators published *A Severe Misalignment of AI in Mathematics* at mathandai.org, cross-posted on Tao's blog. The piece hit #1 on Hacker News with 1,194 points and sparked a 500-comment argument that quickly outgrew mathematics.

The thesis is narrow and sharp. Over the last several months, LLM math capability has gone from party trick to something that can plausibly attack real research problems. AI labs have noticed, and are pouring resources into knocking down famous open problems as public benchmarks. Tao's argument is not that this is technically impossible, or even undesirable in the abstract. It's that the goals of the AI companies and the goals of the mathematical community are severely misaligned, and treating famous problems as scoreboards actively damages the thing those problems were supposed to measure.

The core claim: famous problems were never the point. They were lighthouses. Solving one signaled that a new set of ideas, abstractions, and methods existed — and the community's real work was the long, arduous process of talks, discussions, simplifications, and eventually a textbook presentation that any graduate student could learn from. The answer was a byproduct. The digestion was the product.

Why it matters

This is Goodhart's Law with a Fields Medal attached. When a measure becomes a target, it ceases to be a good measure — and "solve a famous open problem" was always a measure, standing in for "produce ideas worth teaching for the next fifty years." An AI system that dumps a valid but incomprehensible 800-page proof on arXiv technically clears the bar and destroys the signal at the same time.

The Hacker News thread surfaced the strongest counterarguments, and they're worth taking seriously. One mathematician (tmhn2) points out that Mochizuki's abc conjecture proof already did this — a human working in relative isolation dumped an incomprehensible manuscript on the community, and years later nobody agrees whether it's a proof. AI didn't invent that failure mode. Another commenter (jeremysalwen) makes a subtler cut: AI hasn't destroyed mathematicians' ability to develop and share understanding, it's destroyed the *yardstick* used to measure who's contributing to that understanding. Those are different problems with different remediations.

The chess analogy came up repeatedly and it's the wrong one. gwd on HN noted that people said computers would kill chess in the 90s, and today chess is more popular than ever. True — but chess is a game. Mathematics is a research profession whose output feeds every other science. When Stockfish crushes grandmasters, no bridge falls down and no cryptographic protocol breaks. The stakes of the metric-vs-substance question scale with what depends on the substance.

The deeper pattern here is not about mathematics at all. Every field with a legible benchmark and a fuzzy underlying craft is about to run this experiment, whether it wants to or not. SWE-bench for software engineering. MMLU for general knowledge. HumanEval for code. Each one started as a lighthouse — a proxy meant to tell you something about the territory. Each one is now a target that labs optimize directly, with training data curated to hit the number. The map replaces the territory, and then someone ships the map.

The piece also lands on something the AI safety discourse rarely names cleanly: this is an alignment failure between institutions, not models. The lab's incentive is a leaderboard win and a press cycle. The mathematician's incentive is a paper that generates a decade of follow-up work. Both are rational actors pursuing legitimate goals. The friction is that they've decided to share a scoreboard.

What this means for your stack

If you build with LLMs, the practical warning is that benchmark saturation is not the same as capability. A model that tops HumanEval and a model that reduces your on-call load are increasingly different animals, and the gap is widening in the direction the benchmarks don't measure. Contamination is part of it — models trained on solutions to the exact problems being scored — but the harder issue is that scoring well on legible tasks systematically fails to capture the illegible ones: taste, restraint, knowing what not to build, refactoring toward something a junior engineer can maintain.

The practical move is to build your own evals against your own code and your own failure modes. Every serious eng org running LLMs in production has quietly figured this out; the ones that trust the public leaderboards are the ones getting surprised in staging. If you're picking a model for a specific job, the benchmark rank is a filter, not a ranking. Run it on ten of your actual tickets and read the outputs. That's not folk wisdom, it's the only test that hasn't been contaminated by the training loop.

The broader lesson for anyone shipping software with AI in the loop: watch for the moment your team stops asking "did this help?" and starts asking "did the number go up?" That's the Goodhart transition. It's the same failure Tao is describing, just at smaller scale and with less prestigious problems.

Looking ahead

The mathematicians will survive this. They've survived worse — the Bourbaki reformation, the four-color theorem's computer-assisted proof, Mochizuki. What's new is the speed and the money. Labs will keep announcing famous-problem solves because the press cycle is real and the fundraising works. The mathematical community will keep doing the slow work of understanding, largely offscreen. The interesting question is whether other fields — engineering, science, medicine — notice the pattern early enough to build their own evaluation infrastructure before the same institutional misalignment eats them. Tao has given you the vocabulary. Whether the rest of us pick it up before our own lighthouses get repurposed as scoreboards is the actual open problem.

Hacker News 1202 pts 1185 comments

A misalignment of AI in mathematics

<a href="https:&#x2F;&#x2F;terrytao.wordpress.com&#x2F;2026&#x2F;09&#x2F;11&#x2F;a-severe-misalignment-of-ai-in-mathematics&#x2F;" rel="nofollow">https:&#x2F;&#x2F;terrytao.wordpress.com&#x2F;2026&#x2

→ read on Hacker News
Devblogs 87 pts 12 comments

A Severe Misalignment of AI in Mathematics

→ read on Devblogs
tmhn2 · Hacker News

As a mathematician maybe I am a little more optimistic than this declaration.I am thinking of Mochizuki&#x27;s abc conjecture: He worked in relative isolation, and dumped a huge incomprehensible proof on the community (to oversimplify a bit). That&#x27;s not totally unlike what might happen if AI ge

pks016 · Hacker News

I never expected this many people (on this thread) arguing semantics and what not. I know that not everyone has morality and ethics, but I didn&#x27;t realize it was this bad.I&#x27;m afraid of the ripple effect of the agenda pushed by AI companies will have. In future and even now, they say AI has

jeremysalwen · Hacker News

To me it doesn&#x27;t seem like what AI has destroyed is the ability for mathematicians to develop understanding and share it with each other, but rather it&#x27;s destroyed the yardstick (solving open problems) that has traditionally been used to measure how much they have contributed to that under

david-gpu · Hacker News

Tao&#x27;s critique of AI in the field of mathematics reminds me of what French art critic Charles Baudelaire said in the 19th century about photography [0].Baudelaire argued that photography became a haven for failed painters, the sorts of hacks that could not finish proper training. Photography, a

gwd · Hacker News

This sounds a lot to me like people in the 90&#x27;s complaining that computers were destroying chess. Thirty years later, chess is more popular than it ever was, and chess players are better than they ever have been. I wouldn&#x27;t be surprised if there are now more chess books now than there ever

// share this

// get daily digest

Top 10 dev stories every morning at 8am UTC. AI-curated. Retro terminal HTML email.