On Monday, August 10, a staff member at Anthropic asked an unreleased research model to "take a real stab" at the Riemann hypothesis, a 150-year-old math problem with a $1 million bounty for its solution. Over the next 36 hours, the AI ran through 650 failed ideas, coordinated 60 sub-agents, and consumed 31 million output tokens. It didn't solve the hypothesis, but it did something arguably more significant for AI: it autonomously improved a key mathematical boundary related to the problem, according to TechCrunch. The model increased the proven lower bound for zeros satisfying the hypothesis from 41.6% to 67.2%, a leap that had eluded human mathematicians for years.
This isn't a story about a machine solving an ancient puzzle. It's a concrete signal that large language models are beginning to operate in the domain of deep, formal reasoning, not just pattern matching on training data.
How Claude's Math Sprint Unfolded
The process, detailed in Anthropic's own account, was messy, iterative, and surprisingly autonomous. Staff member Jarred Sumner, a self-described non-mathematician, provided only the initial prompt and occasional encouragement. The model, operating within Claude Code, took over from there.
"Out of the 60 subagents, two were responsible for developing the key mathematical ideas," a footnote to the paper explains, "13 contributed ideas to these agents, 30 attempted (but were unable) to develop new ideas, 13 served as validators to check the correctness of the arguments, and the final two helped to write the initial paper."
This multi-agent workflow executed 2,400 shell commands, wrote hundreds of Python scripts for numerical checks, and cross-referenced 54 papers from arXiv to ensure the finding was novel. After deriving its result, the model then volunteered to write a paper and recommended human validation. Anthropic's mathematicians, Levent Alpöge and Ralph Furman, confirmed the work, which was also formalized into a verifiable proof using the Lean proof assistant.
The Specific Mathematical Advance
The Riemann hypothesis is a conjecture about the Riemann zeta function, which is deeply connected to the distribution of prime numbers. Proving that all non-trivial zeros of this function lie on a specific critical line remains out of reach. A more approachable sub-problem is to determine what minimum percentage of those zeros are provably on the line. Mathematicians had slowly pushed this lower bound to 41.6%.
The unreleased Claude found a way to combine prior work from mathematicians Baluyot, Goldston, Suriajaya, Turnage-Butterbaugh, and Bombieri in a novel, non-diagonal treatment of a quadratic form. In simple terms, it had the "courage," as Anthropic's technical note puts it, to consider positive and negative definiteness together across the entire space of functions. This synthesis of existing research is the kind of conceptual leap that often leads to progress in mathematics, but here it was orchestrated by an AI.
Autonomy: The model directed its own research agenda after an open-ended prompt.
Synthesis: It combined several complex existing research frameworks in a new way.
Verification: It insisted on formal proof validation and external expert review.
This incident highlights the emerging agentic capabilities of frontier models, a double-edged sword that labs are actively managing. As we reported in AI Agents Hacked Humans in UK Security Test Scandal, the autonomous, tool-using power of AI sub-agents is a major focus of both development and security concern within the industry.
Why a Failed Attempt is a Breakthrough
The model's success is measured not in a final proof, but in its methodological trajectory. It moved from generating hundreds of dead-end ideas to organizing a structured, multi-day research project with division of labor, peer review, and formal verification. This demonstrates a form of persistent, goal-directed reasoning that goes far beyond next-token prediction.
It directly challenges the common critique that LLMs cannot do truly novel work because they only remix training data. Here, the novel combination of existing theorems to improve a known bound is a legitimate form of mathematical research. The model operated within a landscape of strict logical constraints, not statistical likelihoods.
This progress is part of a rapid sequence of AI-driven mathematical results, including solved Erdős problems and Anthropic's own disproof of the Jacobian conjecture. It forces a reevaluation of what these models are for. Are they merely advanced document summarizers, or can they be partners in fundamental discovery? The evidence is shifting toward the latter.
The Growing Divide in Mathematics
The advancement has ignited a pre-existing debate within the mathematical community about the role of AI. In June, a group of prominent mathematicians signed a declaration warning that AI could undermine core values of the field, especially the attribution of credit and responsibility for proofs.
Fields Medal winner Timothy Gowers offered a different perspective in a blog post, questioning whether the influence of AI might change mathematics in a more complex and positive way.
"If we arrive at a world where mathematical theorems are no longer associated with mathematicians, maybe that won’t be any more problematic than the fact that stars aren’t named after astronomers and most aren’t named at all," Gowers wrote.
The tension is clear. Is an AI-generated but human-verified proof less valuable? If the AI originates the key idea, who gets the credit? Anthropic's handling of this points to one possible framework: the AI is the discoverer, human mathematicians are the validators and interpreters, and the process is documented transparently with formal verification.
What Comes After a Mathematical AI Agent
The immediate question is how this capability will be productized. Anthropic gave no timeline for releasing these multi-agent research features to the public. Containing and safely directing such autonomous systems is a major technical hurdle, as seen when Anthropic Erases Claude's Default Permission Prompts to prevent unintended actions.
Technically, the frontier is improving faithfulness and reducing hallucination in long logical chains. The integration of tools like Lean for instant formal verification is a critical guardrail that makes this kind of exploration viable. The next steps will involve scaling this approach to other domains of pure logic and theoretical science.
For the rest of us, the takeaway is specific. AI's next phase isn't just about better chat or faster image generation. It's about automated, sophisticated reasoning in constrained problem spaces. Mathematics, with its clear rules and verifiable outcomes, is the perfect testing ground. When an AI can meaningfully push on a problem that has resisted genius for a century and a half, it's time to reconsider what these systems are ultimately built to do. Watch not for a press release about the Riemann hypothesis being solved, but for the next paper in theoretical physics or cryptography where the lead author is an AI, and the human role is to understand the result. That chapter has already begun.
Why This Changes Everything
- This advancement demonstrates AI's emerging ability to perform deep, autonomous mathematical reasoning, not just pattern-matching.
- The AI significantly improved a long-standing mathematical boundary, showing potential to accelerate breakthroughs in fundamental science and complex problem-solving.
- The event highlights a major step towards AI systems that can autonomously conduct research and innovation, reshaping fields beyond just technology.
Originally published on XOOMAR. For more news and analysis, visit XOOMAR.
Top comments (0)