Tao argues that genuinely hard, well-posed open problems accumulate slowly over decades and require rare taste to formulate, while AI systems are now solving them at an accelerating pace. His concern is structural inventory depletion, not AI capability or cheating — every problem solved shrinks a pool that doesn't naturally regrow.
By submitting Tao's post with the framing 'Open math problems being non-renewably mined by AI,' the submitter amplifies and endorses the resource-depletion thesis. The 318-point score signals the framing resonated broadly with the HN audience.
The editorial extends Tao's argument to the broader AI evaluation ecosystem, noting that benchmarks like FrontierMath, Putnam, IMO, and Humanity's Last Exam all presume an inexhaustible upstream supply of hard problems. In math specifically, that presumption fails — saturating benchmarks doesn't just retire tests, it consumes the underlying research inventory.
On September 8, Terence Tao — Fields medalist, prolific problem-solver, and one of the few working mathematicians whose Mastodon posts routinely land on the Hacker News front page — published a short note on mathstodon.xyz arguing that the current AI-in-mathematics boom is quietly burning through a resource that doesn't grow back. His framing: open mathematical problems are a non-renewable resource, and AI systems are now mining them faster than the mathematical community can produce replacements.
The post landed at 318 points on Hacker News within hours, a signal that the framing struck a nerve well beyond the math community. Tao's argument is not that AI is bad at math, nor that it's cheating, nor that benchmarks are gamed in the usual sense. It's more structural: every time a well-known open problem gets solved — whether by a human, by an AI, or by a human using an AI — the pool of hard, well-posed, unsolved problems shrinks by one. Unlike code, which developers generate continuously as a byproduct of building products, genuinely hard open problems in mathematics accumulate slowly, over decades, and often require a specific kind of taste to pose in the first place.
Tao has spent the last two years being one of the more sober voices on AI-assisted mathematics, publicly experimenting with GPT-4, Lean, and Copilot-style tools on real research. That track record is what makes this post land differently than the usual "AI will/won't do math" discourse. He's not warning about capability. He's warning about inventory.
The standard mental model for AI progress in reasoning is that benchmarks get saturated, someone builds a harder benchmark, models catch up, repeat. FrontierMath, Putnam, IMO, Humanity's Last Exam — the treadmill assumes an inexhaustible supply of problems on the other end. Tao's point is that in mathematics specifically, the treadmill is finite, because the problems themselves are the treadmill.
Compare this to code. GitHub produces roughly 500 million new repositories a year. Every deployed service generates fresh bug reports, fresh edge cases, fresh integration failures. The training and evaluation substrate for coding models refreshes itself constantly, generated as exhaust from the world's actual work. Mathematics has no such exhaust stream. The Millennium Problems have been sitting there for a quarter century. The Erdős problems accumulate at a rate governed by how many Erdőses we happen to have alive at any given time — which, at present, is zero. New open problems get posed at conferences, in papers, on MathOverflow, but the rate is on the order of thousands per year of *interesting* ones, and only a small fraction are the kind of well-posed, self-contained, verifiable problems that make good AI benchmarks.
If a lab burns through a curated benchmark of 500 research-grade problems in one training cycle, the community cannot simply generate 500 more next quarter. The people qualified to pose them are the same people being asked to solve them, and there are maybe a few thousand of them worldwide across all subfields combined. Community reaction on Hacker News picked up on the second-order effect: benchmark contamination is already a chronic problem, and if labs start scraping problem lists faster than problems get published, the entire evaluation infrastructure for mathematical reasoning degrades. You end up with models that look like they're improving because the test set keeps regenerating from the same shrinking well.
There's also a subtler version of the argument that Tao gestures at but doesn't fully unpack: the *training* signal for mathematical reasoning may be even scarcer than the evaluation signal. Reinforcement learning on math problems requires ground-truth solutions, and the highest-quality ground-truth solutions — elegant, minimal, generalizable — are exactly the kind of artifact that takes a mathematician years to produce. Synthetic problem generation helps at the low end (arithmetic, algebra drills, competition problems in known formats) but degrades quickly as difficulty rises. You cannot generate Riemann Hypothesis-adjacent training data at scale, because if you could, you'd have solved the Riemann Hypothesis.
If you're building anything that touches formal reasoning — theorem provers, program verifiers, symbolic solvers, or any RAG pipeline that leans on mathematical or scientific correctness — Tao's framing has a couple of concrete implications.
First, treat benchmark scores on math-heavy evals with the same skepticism you'd apply to a training set that's been leaked. FrontierMath is closed-source specifically to slow this dynamic; most public math benchmarks are not, and their half-lives are shorter than they look. If your product decisions depend on "model X scores 87% on benchmark Y," you need to know when that benchmark was published, when the model's training data was cut off, and how many other models have already been optimized against it. "Held-out" is doing a lot of work in most eval reports.
Second, if your use case is genuinely novel mathematical reasoning — not competition-style problems, but real research-adjacent work — you probably need to invest in your own private eval set, curated by someone with domain taste, refreshed on a schedule you control. This is expensive and annoying, and it's also the only way to get signal that isn't downstream of whatever benchmark the labs happened to optimize against last quarter. The same logic applied to coding evals two years ago; it applies to math evals now, and to scientific reasoning next.
Third, and this is the uncomfortable one: the economics of "AI does math research" may be less about model capability and more about who owns access to the shrinking pool of unsolved problems. Labs that can lock down proprietary problem sets — through university partnerships, contest sponsorships, or by hiring the mathematicians who pose them — gain a durable moat that has nothing to do with training FLOPS. Watch who's writing checks to whom.
The optimistic read on Tao's post is that it forces a healthier conversation about what "solving math" actually means for AI. Mathematics doesn't run out — new fields, new formalisms, and new questions emerge from every solved problem. But the *readily-benchmarkable* subset is finite in exactly the way Tao describes, and the industry's evaluation infrastructure is built almost entirely on that subset. The next 18 months will make clear whether labs treat this as a resource-management problem to solve collectively, or as an extraction race to win individually. Given the current incentive structure, bet on extraction.
That’s not untrue. But it’s also a misstatement of mathematical history. Many leading mathematicians historically have been highly competitive — Gauss comes to mind. Woe betide the lesser intellect that sent Gauss some ideas. The Newton Leibniz controversy was very serious business at the time in th
Tao's central point seems to be:"In short, the indiscriminate use of powerful solution-extraction tools can achieve the immediate short-term goal of solving problems at hand, but at the cost of sustaining the ecosystem for the next wave of progress, or in understanding the progress already
I'm with @nilesh on this one, and not exactly sure how merely the existence of a solution precludes the advancement of human knowledge. If a problem is "solved" (say, symbolically verified) without any insights gained, it doesn't seem very interesting to the profession.Navier-Sto
It is now clear to me why the AI labs are sponsoring these mathathons: https://mathathonchallenge.com/. They are basically crowdsourcing human researcher data to get access to promising directions possibly later to scoop others.
Top 10 dev stories every morning at 8am UTC. AI-curated. Retro terminal HTML email.
From "Jokester" by Isaac Asimov 1956:"Early in the history of Multivac, it had become apparent that there was one big bottleneck: the questioning procedure. Multivac could answer the problems of humanity, all the problems, if -- if it were asked meaningful questions. But as knowledge