Tao and collaborators argue that AI labs treating landmark open problems as leaderboards to conquer fundamentally misunderstands what those problems are for. Historically, solving a famous problem was a signal that new insights, methods, and abstractions had been developed — insights the community would then digest, simplify, and teach. A proof produced by a machine that no human can meaningfully read delivers the trophy without the understanding that was the actual point.
In the mirrored post on his personal blog, Tao frames the situation as a 'severe misalignment' between AI company goals and mathematical community goals. He positions this not as a niche complaint but as part of broader alignment issues affecting all scientific and creative professions where legible benchmarks exist.
The editorial extends Tao's argument beyond mathematics, framing it as the cleanest statement of the AI-alignment problem in expert work. Every field with a legible benchmark — competitive programming, radiology, legal analysis, code review — is about to discover the benchmark was a proxy that only worked because humans could only produce it by also producing the underlying skill. Once the score and the taste decouple, the proxy stops measuring what it was implicitly measuring.
The submission surfaces the authors' own concession that LLM mathematical capability has 'improved dramatically' over recent months, to the point of solving major outstanding problems in many fields. The 1,166-point score suggests the HN community treats this capability claim as credible rather than hype — the debate is about consequences, not whether it is happening.
On September 11, Terence Tao and a group of co-authors published *A Severe Misalignment of AI in Mathematics* at mathandai.org (mirrored on Tao's blog). The Hacker News thread cleared 1,160 points within hours, which for a manifesto about the epistemology of proof is roughly the equivalent of going platinum.
The argument is short and unhedged. Over the last few months, LLMs have gone from struggling with olympiad problems to producing work that, in the authors' words, can "solve major outstanding problems in many fields of mathematics." That is not the complaint. The complaint is that AI labs are treating famous open problems as benchmarks to be knocked down, and that this optimization target is actively hostile to how mathematics actually accumulates knowledge.
Tao's framing is that research math is not a leaderboard. Landmark problems have served as *lighthouses* — solving one was historically a signal that some new idea, method, or abstraction had been developed, which the community would then spend years digesting, simplifying, and eventually teaching to undergraduates. A proof is a receipt for an insight. The authors argue that when a model produces a proof that no human can meaningfully read, the receipt arrives without the thing you were actually buying.
Strip away the domain and this is the cleanest statement of the AI-alignment problem in expert work anyone has written this year. It is not about paperclips. It is about Goodhart's law with a byline.
Every field that has a legible benchmark — competitive programming, radiology reads, legal issue-spotting, code review — is about to discover that the benchmark was a proxy for something the field never bothered to name, because humans could only produce the proxy by also producing the thing. Solving IMO problems used to require mathematical taste because there was no other way to do it. Now there is another way. The taste and the score have decoupled, and labs are racing on the score.
The community response is more interesting than the piece itself, because working mathematicians are not uniform on this. On HN, `tmhn2` points out that Mochizuki's abc conjecture proof already produced the exact failure mode Tao is warning about — a huge, incomprehensible artifact dumped on a community that couldn't absorb it — and no AI was involved. `jeremysalwen` sharpens the argument: what AI has destroyed is not mathematicians' ability to understand, but the *yardstick* they used to measure whether understanding had occurred. `gwd` pushes back with the chess analogy: engines were supposed to kill chess in the 1990s, and instead chess is bigger, players are stronger, and there are more books than ever. `david-gpu` reaches further back, to Baudelaire's 1859 essay complaining that photography was a refuge for failed painters — and photography, of course, went on to become its own art form while painting did just fine.
Both analogies undersell what Tao is actually claiming. Chess engines did not replace the *purpose* of chess, which was always the contest between minds. Mathematics has no comparable spectator layer — the entire point of the enterprise is the human understanding produced along the way, and there is no version of the field where the score matters and the understanding doesn't. Photography created a new medium alongside painting; a proof-generating model does not create a new mathematics alongside the old one, it produces artifacts inside the same one, competing for the same attention and the same prestige.
The deeper worry, only hinted at in the piece, is about students and ideas — what the authors call "the most precious resources of our profession." If the reward gradient in a PhD program tilts toward "prompt the model until it cracks the problem," the pipeline that produces the next generation of taste-makers atrophies. This is the alignment failure that actually scales: not the model, the incentive structure around the model.
If you're building AI features into any expert workflow, the practical read is this: your evals are lying to you in a way that will only become obvious when the domain experts leave.
Benchmarks in expert fields were almost always constructed under the assumption that the only way to score well was to have the underlying competence. That assumption is now false across the board. A code-generation model that passes your test suite is doing something categorically different from a senior engineer who would have written code that passes the same suite, and the difference does not show up in the test suite. It shows up two years later in the maintainability of the codebase, the onboarding time of new hires, and the number of incidents nobody can debug because the person who understood the system was a fine-tuned checkpoint.
Concretely, three moves worth making now. First, separate your *capability* evals from your *comprehensibility* evals — measure not just whether the output is correct but whether a human reviewer can explain *why* it's correct in a bounded time. Second, treat any benchmark that has been public for more than six months as contaminated by default and useful only for regression testing, not for capability claims. Third, if you employ domain experts, protect the parts of their job that produce taste — code review, design docs, mentoring juniors — even when a model could plausibly do the surface-level task faster. Those activities are the field's version of Tao's "long and arduous process of talks, discussions, simplifications." They look inefficient. They are how the discipline stays a discipline.
The interesting question is not whether Tao is right — he is, on the narrow point — but whether any field with a large benchmark surface can resist the gravitational pull of chasing it. Mathematics is the least commercial expert domain in existence and it's already having this argument in public; every other field will have it later, more quietly, and with worse incentives. The labs are not going to unilaterally stop targeting Millennium Prize problems, and the mathematicians are not going to stop noticing when a model gets one. What's left is the boring middle: building evaluation regimes that reward understanding over artifacts, and figuring out which parts of expert work were never really the parts we were measuring.
<a href="https://terrytao.wordpress.com/2026/09/11/a-severe-misalignment-of-ai-in-mathematics/" rel="nofollow">https://terrytao.wordpress.com/2026
→ read on Hacker NewsI never expected this many people (on this thread) arguing semantics and what not. I know that not everyone has morality and ethics, but I didn't realize it was this bad.I'm afraid of the ripple effect of the agenda pushed by AI companies will have. In future and even now, they say AI has
To me it doesn't seem like what AI has destroyed is the ability for mathematicians to develop understanding and share it with each other, but rather it's destroyed the yardstick (solving open problems) that has traditionally been used to measure how much they have contributed to that under
Tao's critique of AI in the field of mathematics reminds me of what French art critic Charles Baudelaire said in the 19th century about photography [0].Baudelaire argued that photography became a haven for failed painters, the sorts of hacks that could not finish proper training. Photography, a
This sounds a lot to me like people in the 90's complaining that computers were destroying chess. Thirty years later, chess is more popular than it ever was, and chess players are better than they ever have been. I wouldn't be surprised if there are now more chess books now than there ever
Top 10 dev stories every morning at 8am UTC. AI-curated. Retro terminal HTML email.
As a mathematician maybe I am a little more optimistic than this declaration.I am thinking of Mochizuki's abc conjecture: He worked in relative isolation, and dumped a huge incomprehensible proof on the community (to oversimplify a bit). That's not totally unlike what might happen if AI ge