Originally published on The AI Prism
The Week an 87-Year-Old Conjecture Fell
On July 19, 2026, a problem mathematicians had chased since 1939 was finally settled. Not by a tenured professor. Not by a Fields Medalist. By Levent Alpöge, a mathematician who works at Anthropic, using the company’s Claude Fable 5 model to produce an explicit counterexample to the Jacobian conjecture in three dimensions.
Within 48 hours, Terence Tao had published a “digestion” of the counterexample on his blog, run a long ChatGPT Pro session hunting for a geometric explanation, and watched the Hacker News thread about it pull in 1,126 points and 635 comments — including a companion thread titled “Human mathematicians are being outcounterexampled.”
Five days later, Tao stood before the International Congress of Mathematicians 2026 and told his field the uncomfortable truth: “I believe we are entering a similarly turbulent period — a crisis in the foundations of mathematical values and practices.” That line is from his ICM public lecture, which compared the moment to the 1900-1930 crisis that forced mathematics to formalize its own foundations.
Here is the question nobody is asking: what does the world’s greatest living mathematician see that we don’t?
The World’s Greatest Living Mathematician Is Running a Public Experiment
Tao is not a casual AI observer. The “Mozart of Math” — a 2006 Fields Medalist routinely described as the finest mathematician alive — has spent four years publishing his AI experiments in real time on his blog and Mastodon. That public record is the closest thing we have to a controlled study of how frontier AI changes the work of an elite scientist.
The arc is unmistakable. In April 2023, Tao reported that GPT-4 had “saved me a significant amount of tedious work” for the first time. By June 2024, he told Scientific American: “I think in three years AI will become useful for mathematicians. It will be a great co-pilot.” By November 2025, he was documenting that “AI assistance is now becoming routine” on the Erdős problems website.
Every stage came with receipts: shared ChatGPT conversations, Lean formalizations on GitHub, detailed Mastodon threads. This is not commentary about AI. It is a lab notebook.
What Tao Sees: A Mediocre, But Not Completely Incompetent, Graduate Student
In September 2024, after testing OpenAI’s o1 reasoning model, Tao delivered the most-quoted verdict in AI mathematics: the experience was “roughly on par with trying to advise a mediocre, but not completely incompetent, graduate student.”
He later corrected the viral reading of that line. He was not comparing o1 to a graduate student in general — he was comparing it to a mediocre research assistant. It handles routine computation reliably but is “very unimaginative” at the clever step, and it lacks the one property that makes human students valuable: learning. “These models are static,” Tao told The Atlantic. “Humans have growth.”
He also gave the field its first honest efficiency metric. Producing useful output with the best models still costs 2x to 5x the effort of doing the work yourself. His stated tipping point: when that ratio falls below 1x — which he expects within a few years — adoption stops being a debate.
What the Numbers Say
The benchmark arc moves faster than most people can track. In July 2024, DeepMind’s AlphaProof and AlphaGeometry 2 solved four of six IMO 2024 problems for 28 of 42 points — silver-medal standard. The hardest problem had been solved by only 5 of 609 human contestants, and gold started at 29 points. The methodology behind AlphaProof was later published in Nature.
Then the goalposts moved. In November 2024, Epoch AI released FrontierMath: hundreds of original research-level problems written by more than 60 mathematicians. Leading models solved less than 2%. Tao called the problems “extremely challenging”; Timothy Gowers said they sit “at a different level of difficulty from IMO problems.” In December 2024, OpenAI’s o3 jumped to 25.2% — a leap that later revealed OpenAI had quietly funded FrontierMath’s creation, a transparency failure Epoch AI has since acknowledged.
In 2026 the frontier moved from benchmarks to open problems. First Proof, an independent assessment project, tested four AI harnesses against ten novel research problems on May 28, 2026: seven of ten were solved at publication-level quality, at compute costs of $10 to $1,000 per problem. In March 2026, a GPT-5.4 Pro-driven team became the first to solve a FrontierMath open problem — a Ramsey-theoretic construction Epoch estimates would take an expert human 1-3 months. In May 2026, DeepMind’s AlphaProof Nexus resolved 9 of 353 open Erdős problems and proved 44 of 492 OEIS sequence conjectures at a few hundred dollars per problem.
And then came the Jacobian counterexample: a degree-7 polynomial whose Jacobian cancellation involves 1,329 coefficients against only 120 degrees of freedom — what Tao called “a massive miracle” that brute force would never have found.
The Quiet Workhorse: Lean and the Formalization Pipeline
Generative models get the headlines, but Tao’s workflow runs on a quieter technology: Lean, an interactive theorem prover that checks proofs line by line. In October 2023, formalizing his own paper in Lean surfaced a small but non-trivial bug in an argument he had already published — an error no human referee had caught.
Lean also enabled the largest collaborative proof project in recent memory: the formalization of the Polynomial Freiman-Ruzsa (PFR) conjecture, where more than 20 mathematicians contributed pieces of one proof. “You don’t need to trust them, because they upload code and the Lean compiler verifies it,” Tao explained. “You can do much larger-scale mathematics than we do normally.”
Watch how routine this has become. In November 2025, on Erdős problem #367: a human contributor produced a disproof contingent on an unverified congruence identity; Tao handed the identity to Gemini DeepThink, which proved it in about ten minutes; Tao spent half an hour rewriting it into an elementary proof; and another mathematician formalized the result in Lean in two to three hours. Tao’s own summary: “AI assistance is now becoming routine.” A month earlier, an extended AI conversation helped him answer a MathOverflow question — a task he says he “would have been very unlikely to even attempt” unassisted.
The Erdős Wiki: Proof That AI Assistance Is Now Routine
The best evidence is a living document: the AI contributions to Erdős problems wiki, maintained by Tao’s project with 962 revisions and data through June 30, 2026. It logs dozens of AI attempts against Erdős’s open problems, with color-coded outcomes: full solutions, partial progress, incorrect proofs, and unverified candidates.
The list reads like a who’s who of frontier AI: GPT-5.5 Pro, Claude Fable 5 and Claude Mythos, Gemini 3 Pro, DeepMind prover agents, AlphaProof, Aristotle, Codex. Full solutions are recorded for problems #38, #90, #205, #457, #694, #960, #987, #990, #1014 and #1091, among others — several delivered in Lean, meaning they are machine-checked.
What makes the wiki credible is what it refuses to hide. It also records the incorrect proofs — the confident failures on #11, #51, #233, #616, #647, #888, #963, #1041 and #1044. The disclaimers are blunt: “This page is not a benchmark,” and success rates should not be inferred. That honesty is the difference between a marketing claim and a research log.
Where Machine Reasoning Hits Its Limits
Every serious observer now agrees on where AI math breaks down: without formal verification, an AI proof is just a confident story. Natural-language models hallucinate plausible-looking arguments — the entire point of the Lean pipeline is that a checker, not a vibe, decides correctness.
But verification is not the only bottleneck. Tao’s ICM lecture called out what he terms “proof indigestion”: the Erdős problems site already holds “dozens of AI-generated proof submissions. Many are likely to be correct, but no human expert has yet volunteered to verify and vouch for them.” Some submitters have declared themselves unqualified to check their own AI’s output. Could we get a verified proof of a major result that no human can explain? Tao thinks the question is live.
Then there are the softer limits. AI exposition “dwells at length on trivialities, while passing very briefly through the most interesting and novel portions of the argument.” AI knowledge is frozen at training time — the same week the Jacobian counterexample went public, the models had to be told it existed, because their knowledge cut off before the discovery. And metrics corrupt: Tao invoked Goodhart’s law — when a measure becomes a target, it stops being a measure — and the FrontierMath funding episode showed how benchmark scores can be shaped by the companies being scored. Even the models’ training data is a separate battleground, as we explored in our piece on AI companies shredding rare books for training data.
What This Means for Science
Mathematics is the canary, but the pattern generalizes. Tao’s framing is the cleanest available: for centuries, mathematics ran on proof scarcity — the hard part was producing results. AI inverts the economics. The hard parts become verification, exposition, community acceptance, and what Tao calls canonicalization: a result only matters once it is digested, taught, and built into the theory that everyone else relies on. “We will transition from an era of proof scarcity to an era of proof abundance,” he warned.
His proposed guardrail is beautifully simple: if authors cannot convincingly give a clear, expert-level talk on their results, correctly attributed, the result should not be published. The Leiden declaration, referenced in his talk, pushes the same norms: disclose AI use, keep humans accountable. Meanwhile institutions are betting real money on the trend — DARPA’s ExpMath program funds AI-driven mathematics, and Tao himself has co-authored a philosophy-of-math paper, “Mathematical methods and human thought in the age of AI.”
Beyond pure math, the same machinery is quietly eating the verification economy: AlphaProof Nexus’s authors point to combinatorics, optimization and algebraic geometry, but the underlying capability — generating formally checkable proofs at a few hundred dollars each — is exactly what smart-contract auditing and zero-knowledge cryptography have been waiting for. If “the job description is changing,” as Tao told Nature, it is changing everywhere proof matters: mathematics, software, security, science itself.
What You Should Do About It
If you work in a reasoning-heavy field, the playbook is already visible in Tao’s workflow:
• Learn the verifier, not just the model. Lean (or Rocq, or HOL) is the difference between “the AI says so” and “it is so.” Tao’s own Lean journey began with GPT-4’s help, and open-source agents like Mistral’s Leanstral now lower the bar further.
• Use AI where output is checkable. Numerical searches, case verification, literature sweeps, formalization — Tao’s wins all share one property: a machine (or a 29-line Python script) can confirm them.
• Keep the “talk test.” If you cannot explain your AI-assisted result to an expert from memory, you do not own the result. Treat unexplained AI output as raw material, not a finding.
• Disclose AI use. Tao’s ICM slides carry a footnote admitting AI autocompleted text and generated diagrams. Normalize the disclosure, and you starve the covert-use scandals before they start.
The Bottom Line
Terence Tao’s real message is not that AI will solve mathematics. It is that AI is forcing mathematics to decide what it is for — and the same question is coming for every field that runs on verified reasoning. A genius sees this first because he has the strongest incentive: his entire craft is the production of trustworthy arguments, and the production half just got cheap.
The scarcity that remains — understanding, explanation, judgment, taste — is the part that was always human. The question is whether we treat it as the bottleneck or as the point. If the world’s greatest living mathematician is right, the mathematicians who thrive in the age of AI will not be the fastest provers. They will be the ones who know what a proof is for.
So here is the question we are leaving you with: when an AI produces a correct proof that no human alive can explain, is it mathematics — or is it just output?
References
• Jacobian conjecture — Wikipedia
• Terence Tao, “A digestion of the Jacobian conjecture counterexample” (July 21, 2026)
• Terence Tao’s ChatGPT conversation on the Jacobian counterexample
• HN thread: Terence Tao’s ChatGPT conversation about the Jacobian Conjecture counterexample
• HN thread: Human mathematicians are being outcounterexampled
• HN thread: Claude Fable produced a counterexample to the Jacobian Conjecture
• Terence Tao, “Mathematics in the age of AI,” ICM 2026 public lecture slides (July 24, 2026)
• HN thread: Terence Tao: Mathematics in the Age of AI
• The Atlantic, “We’re Entering Uncharted Territory for Math” (October 4, 2024)
• Scientific American, “AI Will Become Mathematicians’ ‘Co-Pilot'” (June 8, 2024)
• Terence Tao on GPT-4 (April 2023)
• Terence Tao on OpenAI o1 (September 2024)
• Terence Tao on the Lean4 formalization bug in his paper (October 2023)
• Tao et al., the formalized paper on arXiv (2310.05328)
• Terence Tao on Erdős problem #367: AI assistance becoming routine (November 2025)
• Terence Tao on the AI-assisted MathOverflow answer (October 2025)
• Google DeepMind, “AI achieves silver-medal standard solving IMO problems” (July 25, 2024)
• AlphaProof methodology paper, Nature (November 2025)
• Nature news: “Mathematicians put AI model AlphaProof to the test” (November 2025)
• First Proof Project — independent assessment of frontier AI in research mathematics
• teorth/erdosproblems wiki: AI contributions to Erdős problems (updated June 30, 2026)
• The Register, “DARPA to ‘radically’ rev up mathematics research. And yes, with AI” (April 2025)
• The Leiden Declaration on AI and mathematics
The post Terence Tao and the Mathematics of AI: What a Genius Sees That We Don’t appeared first on The AI Prism.
Cross-posted from theaiprism.com — Cutting Through the AI Noise 🧊
Top comments (0)