Tao argues the gap between what frontier labs claim their models can do in math and what those models actually do unaided has become large enough to be a research-integrity issue. Reported numbers reflect a system — prompt scaffolds, external tool calls, best-of-K sampling with verifiers, and retrieval — not the model itself, and stripping those away yields dramatically lower raw performance.
The Economist frames OpenAI's announcement as a methodological scandal, reporting that the company claimed near-frontier performance on research-level problems without disclosing a substantial scaffolding harness and multiple independent samples per problem. The piece amplifies the outrage of top mathematicians who see the practice as misleading.
Characterizes the situation as the AI equivalent of a benchmark laundering scandal, arguing the same pattern of harness-inflated numbers appears across SWE-bench, GPQA, and other frontier benchmarks. The editorial treats this as a systemic disclosure failure rather than a math-specific dispute.
At least two named IMO problem-setters went on record saying they were not told their problems would be evaluated with scaffolding and multi-sample verification. They argue that if the same rules were applied to human contestants, the results would be disqualified, framing OpenAI's methodology as a violation of contest norms.
Submitted Tao's post to Hacker News, where it drew 777 points and 767 comments — signaling that a large developer/researcher audience treats this as a significant integrity issue, not a niche mathematical dispute. The submission's traction reflects broad concern about how frontier labs report capabilities.
On September 11, Terence Tao — Fields medalist, and probably the mathematician whose opinion the field weighs most heavily — published a blog post titled *A severe misalignment of AI in mathematics*. The post argues that the gap between what frontier labs claim their models can do in math and what those models actually do, unaided, has become large enough to constitute a research-integrity problem. The Economist picked it up the same day under the headline *Top mathematicians are outraged by OpenAI's methods*, and by Friday morning the HN thread was sitting at 777 points.
Tao's core complaint isn't that the models are bad at math — it's that the reported numbers describe a system, not a model. When OpenAI announces that its latest model solved N problems from the IMO or the Putnam, that number typically reflects a pipeline: a human-written prompt scaffold, external tool calls (Python, SymPy, Lean), best-of-K sampling with a verifier picking the winner, and in some cases retrieval over a corpus of solved problems. Strip those away and the raw model performance is dramatically lower. The reported number is real; it's just not a number about the model.
The specific incident that lit the fuse, per the Economist, was an OpenAI announcement claiming near-frontier performance on research-level problems, without disclosing that the runs used a substantial scaffolding harness and multiple independent samples per problem. Several named mathematicians — including at least two IMO problem-setters — went on record saying they had not been told their problems would be evaluated this way, and that if the same rules were applied to human contestants, the results would be disqualified.
This is the AI equivalent of a benchmark laundering scandal, and it matters beyond mathematics because the same pattern shows up everywhere frontier labs report capabilities. SWE-bench numbers depend heavily on the harness. GPQA numbers depend on whether the model gets to think for 30 seconds or 30 minutes. ARC-AGI numbers, famously, depend on whether you're allowed 1,024 samples with a verifier. The industry has quietly normalized reporting the ceiling of a scaffolded system as if it were the floor of a model's raw capability, and Tao is the first person with enough standing in a technical field to make the accusation stick.
The defense — which OpenAI and others make explicitly — is that scaffolding is part of the product. Nobody uses raw completions in production; everyone uses agents with tools. If the scaffolded system solves the problem, the system solved the problem. That's a defensible position for a product announcement. It is not a defensible position for a scientific claim about reasoning ability, which is what the math-community objection is really about. Tao's phrasing is careful here: he distinguishes between engineering achievement (real, impressive) and epistemic claim (misleading, and in some cases actively damaging to the mathematics community, which now has to spend time explaining to funders and administrators why an LLM did not, in fact, prove a novel theorem last week).
The deeper problem is a measurement one. Best-of-K with a verifier is a fundamentally different regime from pass@1, and the gap between them grows superlinearly for problems where verification is cheaper than generation — which describes most of formal math. A model that solves a problem 1 in 500 attempts, with an oracle picking the winning attempt, is not the same artifact as a model that solves it on the first try. Reporting the two under the same label is, to use Tao's word, misalignment. Every serious eval methodology paper of the last two years has said this out loud; every frontier-lab press release has ignored it.
Community reaction on HN split predictably. The 'this is obvious' camp pointed out that anyone reading benchmark papers carefully already knew scaffolding was doing most of the work. The 'this is a big deal' camp countered that 'anyone reading carefully' is roughly 200 people worldwide, and the other several billion — including the ones writing checks and setting policy — read the press releases. Both are correct. The Tao post matters not because the information is new but because the messenger changes what happens next.
If you're picking an LLM for a reasoning-heavy pipeline — code review, incident triage, contract analysis, anything where the model needs to actually chain steps rather than pattern-match — treat every published benchmark as a question, not an answer. Ask three things before you trust a number: what was the sampling budget, what tools did the model have, and what does raw pass@1 look like on the same task. If a vendor won't tell you, that's the answer.
For internal evals, this is a good week to audit your own methodology. If you're running best-of-N with a verifier in your test harness — and a lot of teams are, often without realizing it, because the eval framework does it by default — you're measuring the system, not the model. That's fine if the system is what ships. It's a disaster if you're using the number to decide which model to fine-tune, because you'll pick the model whose scaffolding is best-tuned to your verifier, not the one with the best underlying reasoning. The two diverge quickly.
The practical fix is boring and known: report pass@1, pass@k, and scaffolded-system numbers as three separate columns, and never let anyone in a slide deck collapse them. Internally, this takes a week. Externally, an entire industry has spent two years avoiding it.
Tao's post will not, by itself, change how OpenAI or Anthropic or Google DeepMind reports benchmarks. What it may do is give reviewers at NeurIPS and ICML enough political cover to start rejecting papers whose 'model X achieves Y%' claims don't include a scaffolding ablation. That's a slow lever, but it's the right one — and if it works, expect the next round of frontier-lab announcements to quietly start including the pass@1 numbers they've been leaving out. Watch for whether the next GPT or Claude release includes a raw-capability column. If it does, Tao won. If it doesn't, the misalignment gets worse, and the next scandal will be uglier.
<a href="https://terrytao.wordpress.com/2026/09/11/a-severe-misalignment-of-ai-in-mathematics/" rel="nofollow">https://terrytao.wordpress.com/2026
→ read on Hacker NewsI never expected this many people (on this thread) arguing semantics and what not. I know that not everyone has morality and ethics, but I didn't realize it was this bad.I'm afraid of the ripple effect of the agenda pushed by AI companies will have. In future and even now, they say AI has
To me it doesn't seem like what AI has destroyed is the ability for mathematicians to develop understanding and share it with each other, but rather it's destroyed the yardstick (solving open problems) that has traditionally been used to measure how much they have contributed to that under
Tao's critique of AI in the field of mathematics reminds me of what French art critic Charles Baudelaire said in the 19th century about photography [0].Baudelaire argued that photography became a haven for failed painters, the sorts of hacks that could not finish proper training. Photography, a
This sounds a lot to me like people in the 90's complaining that computers were destroying chess. Thirty years later, chess is more popular than it ever was, and chess players are better than they ever have been. I wouldn't be surprised if there are now more chess books now than there ever
Top 10 dev stories every morning at 8am UTC. AI-curated. Retro terminal HTML email.
As a mathematician maybe I am a little more optimistic than this declaration.I am thinking of Mochizuki's abc conjecture: He worked in relative isolation, and dumped a huge incomprehensible proof on the community (to oversimplify a bit). That's not totally unlike what might happen if AI ge