Anthropic says Claude formalized Fermat's Last Theorem in 11 days, working largely autonomously through the Prove2Me platform. The run generated 13 million lines of Lean code and proved 29,500 intermediate theorems, over five times the size of Mathlib. The proof uses only Lean's three standard axioms. Mathematicians had scoped formalization as a multi-year project. Claude did it in under two weeks.
But here's the thing that matters: Claude didn't prove Fermat's Last Theorem. Andrew Wiles did that in 1995. Claude turned Wiles's proof into code that a computer can verify. The artifact is enormous and genuine. The mathematics is 30-year-old mathematics.
This distinction matters because it points at what's actually hard about mathematical AI, and what isn't. Frontier models are becoming exceptional at a specific task: taking human-readable mathematical arguments and translating them into formal syntax that proof assistants can check. That's real work. It's labor-intensive. It requires coherence across 13 million lines of interdependent logic. Claude's run did something that humans had judged would take years. That's significant.
But it also tells you something about the ceiling. Claude is not doing original mathematical reasoning here. It's not discovering gaps in existing proofs or finding new insight into why Fermat's Last Theorem is true. It's performing a mechanical translation under machine guidance.
The infrastructure story is more interesting than the AI story. Prove2Me, an open-source tool from Columbia University, was added mid-run and made completion possible. Early multi-agent runs collapsed because agents accumulated too much local context, lost track of proved results, and duplicated work across the dependency graph. The breakthrough wasn't scaling up compute. It was adding the right tool. Anthropic's achievement rests on community-built infrastructure it did not create: Kevin Buzzard's Imperial College FLT project, Mathlib, and Columbia University's Prove2Me were all prerequisites for success.
This raises a real problem. If Anthropic couldn't do this without Prove2Me, and if the intermediate theorems, 29,500 of them, were originally generated to solve Anthropic's internal problem, then the question becomes: do those theorems belong to mathematical infrastructure or are they a one-off artifact? Watch whether Mathlib absorbs any material share of the 29,500 intermediate theorems in the coming months; real merges would show the artifact is reusable infrastructure, silence would suggest a self-contained one-off.
The model used was roughly on par with Claude Fable 5.1, the immediate predecessor of GPT-6 Astra. That's not the frontier. Claude ran several dozen agents that generated 6 billion tokens of output over the 11 days. That's expensive, coordinated compute. You can verify the result, Lean checked it. You cannot independently reproduce it unless you have access to the same model at the same scale.
The honest take: this is engineering. It's impressive engineering. It's the kind of thing that changes what's possible in formal mathematics. But it's not a preview of AI doing original mathematics. It's a preview of AI as a really expensive formalization engine, married to the right open-source tools, run by people who understand the target domain.
If you work on proof systems or formal mathematics, assume the volume of machine-generated code you're asked to trust is about to increase by orders of magnitude. If you're waiting for frontier AI to solve open problems in mathematics, this is not that story.
Top comments (0)