DEV Community

Sunny Bhatnagar
Sunny Bhatnagar

Posted on Originally published at presentofai.com

AI Is Now Solving Problems Science Spent Decades Failing to Crack

Between May and July 2026, AI systems disproved an 87-year-old conjecture in mathematics, cracked the Erdos unit distance problem, achieved the first perfect score at the International Math Olympiad, predicted over a billion protein structures surpassing AlphaFold, and broke both weakened encryption algorithms and a NIST post-quantum cryptography candidate in controlled tests. The pace and breadth of these results, spanning pure mathematics, structural biology, and cryptography, suggest that AI capability gains are now compressing scientific timelines in ways that outrun the policy and security frameworks built around them.

The spring and summer of 2026 produced a cluster of AI-driven scientific results that would each have been remarkable in isolation. Taken together, they represent something qualitatively different: a compression of research timelines so sharp that frameworks built to manage scientific risk, from cryptographic standards to mathematical peer review, are struggling to keep pace. Within roughly ten weeks, AI systems disproved conjectures that had resisted human effort for decades, mapped more protein structures than all prior work combined, and broke weakened encryption algorithms and cracked a NIST post-quantum cryptography candidate in controlled tests.

What makes this moment distinct is not just the difficulty of the problems solved, but the breadth of domains involved and the speed at which results accumulated. Pure mathematics, structural biology, and cryptography each operate on different technical foundations and different institutional timelines. AI capability gains are now cutting across all three simultaneously, which means the second-order effects, in drug discovery, in standards bodies, in national security, are arriving faster than the institutions responsible for managing them can absorb.

The Mathematical Barrier Falls, Repeatedly

The sequence opened on May 20, 2026, when OpenAI's Reasoning Model Disproves the Erdős Conjecture, disproving the planar unit distance conjecture that Paul Erdős had posed in 1946. The model produced a verified proof presenting an infinite family of point arrangements with a polynomial improvement over prior constructions. External validation came from mathematician Timothy Gowers, establishing that this was not an artifact or a near-miss but a genuine result in discrete geometry.

Two months later, the pace accelerated. On July 20, Claude Fable 5 Disproves the Jacobian Conjecture, an unsolved problem that had stood for 87 years. On July 22, RedNote AI Achieves Perfect Score at the IMO, becoming the first AI to score a flawless 42 out of 42 at the International Mathematical Olympiad, surpassing the 35 out of 42 that both Google DeepMind and OpenAI had achieved the previous year. Two days after that, on July 24, ChatGPT 5.6 Pro Solves a Decades-Old Problem in Four Prompts, finding a structured counterexample to a longstanding open problem with minimal human guidance.

The mechanism behind this acceleration matters. These are not systems retrieving known proofs or interpolating from training data. The Erdős result required constructing a novel infinite family of arrangements. The Jacobian disproof required autonomous reasoning over algebraic geometry. The IMO perfect score required solving six competition problems under conditions designed to defeat pattern-matching. The common thread is that general-purpose reasoning capability has crossed a threshold where it can operate productively at the frontier of human mathematical knowledge, not just below it.

Biology at Scale: The Protein Atlas Expands

On May 27, ESMFold2 Predicts 1.1 Billion Protein Structures, when Meta's Biohub released ESMFold2, an open-source model that generated a structural atlas covering 1.1 billion proteins. That figure is 800 million more entries than AlphaFold's database, making it the largest protein structure prediction effort ever completed.

The significance here is not purely academic. Protein structure determines function, and function determines drug target viability. A database of this scale, released openly, means that researchers working on rare diseases, antibiotic resistance, and novel therapeutics now have structural information for proteins that previously had none. The open-source release also means this capability is not gated behind a single institution, which accelerates downstream use but also removes centralized oversight of how the data is applied.

The structural biology result connects to the mathematics results through a shared mechanism: AI systems are now operating at a scale and speed that human researchers cannot match on the underlying computational task, whether that task is proof search or protein folding. The human role is shifting toward problem selection, validation, and application rather than primary discovery.

Cryptography: When the Standards Themselves Break

The most consequential results for near-term policy arrived in the final week of July. On July 28 and 29, Claude Mythos Preview Finds Novel Attacks on AES and Claude Mythos Breaks Weakened Encryption in Testing, identifying novel attack vectors against weakened versions of AES cryptographic algorithms in controlled testing, representing the first time an AI independently found weaknesses in widely-deployed encryption standards protecting financial transactions and private communications.

More alarming for long-term security planning, on July 29, Claude Mythos Cracks a NIST Post-Quantum Cipher in 60 Hours. The target was HAWK, a candidate in NIST's post-quantum cryptography standardization process. The model found a critical flaw in 60 hours, defeating two years of global expert review. The same model independently invented a novel AES-128 attack technique, subsequently named the Möbius Bridge.

Several qualifications apply here. The AES attacks targeted weakened versions of the algorithm, not production AES-128 or AES-256 as deployed. The HAWK result is significant precisely because HAWK was a candidate, not a finalized standard, and the flaw was found before deployment rather than after. These distinctions matter for immediate risk assessment. What they do not change is the structural implication: AI systems can now compress the cryptanalytic review cycle from years to days, which means the assumption that a candidate cipher surviving two years of expert review is safe needs to be revisited.

The NIST post-quantum standardization process was designed around human review timelines. If AI-assisted cryptanalysis can cover equivalent ground in 60 hours, the process needs new mechanisms, whether that means AI-assisted review on the defense side, longer candidate exposure periods, or both.

Policy Catches Up, Partially

On July 22, the same day as the IMO perfect score, The White House Announced the Genesis Mission, a national initiative committing more than 5 billion dollars to harnessing AI for scientific discovery, with a stated focus on medical research and positioning the United States as the global leader in AI-driven science.

The timing is striking but the framing is telling. A 5 billion dollar commitment to AI in science is a significant resource allocation. But the Genesis Mission announcement focused on opportunity, specifically medical research and geopolitical positioning, rather than on the governance questions raised by the cryptography results or the validation questions raised by autonomous mathematical proofs. The policy response is running behind the capability curve, addressing the upside while the downside risks accumulate in separate agency processes.

The remote real-time control of a nuclear reactor using AI and distributed HPC, demonstrated on July 14 by Idaho National Laboratory, UIUC, and Purdue, illustrates how rapidly AI is moving into high-consequence physical infrastructure. That result is a controlled demonstration, not a deployment, but it signals that the boundary between AI as a research tool and AI as an operational system in safety-critical environments is narrowing.

What to Watch

  • NIST's response to the HAWK result: Whether NIST accelerates review of remaining post-quantum candidates using AI-assisted cryptanalysis, and whether the Möbius Bridge technique generalizes to other lattice-based schemes, will determine how much of the current post-quantum roadmap needs revision.

  • Validation infrastructure for AI-generated proofs: The Erdős and Jacobian results were externally verified, but the pace of output is increasing. Watch whether the mathematics community develops formal verification pipelines fast enough to keep peer review meaningful as AI proof generation accelerates.

  • ESMFold2 downstream applications: With 1.1 billion protein structures now openly available, the first drug discovery programs built primarily on ESMFold2 data will enter preclinical stages within 12 to 18 months. Their success or failure will calibrate how much of the structural atlas is actionable.

  • Genesis Mission implementation details: The 5 billion dollar commitment needs a governance framework. Watch for agency-level guidance on how AI-generated scientific results will be validated, published, and acted upon, particularly in FDA-adjacent medical research contexts.

  • AI cryptanalysis on production standards: The July results targeted weakened algorithms and a candidate cipher. The next threshold to watch is whether similar techniques are attempted against production AES or finalized post-quantum standards, and whether those attempts are disclosed promptly or surface through breach investigations.


This piece was originally published on Present of AI, where we cover what AI is actually doing in the world, no hype. Read more or get it in your inbox.

Top comments (0)