DEV Community

Jason Gunnells
Jason Gunnells

Posted on

OpenAI Cracks a Million-Dollar Math Problem — and the Credit Fight Starts Immediately

OpenAI Cracks a Million-Dollar Math Problem — and the Credit Fight Starts Immediately

Today's biggest AI story isn't really about the breakthrough itself — it's about who gets to claim it.

OpenAI announced that an unreleased internal model, one the company describes as "significantly more capable" than its just-launched flagship, has produced a proof solving Navier-Stokes, one of mathematics' seven Millennium Prize problems and the equation set that governs how fluids flow. The company says it threw roughly 10,000 AI agents at the problem simultaneously, ran them for 88 hours, and spent an estimated "millions of dollars" in compute to get there. Sam Altman called it one of the most amazing moments in the company's history. It's a genuinely enormous result — Navier-Stokes has resisted a full solution for over a century, and a $1 million prize has sat unclaimed since 2000 waiting for someone to nail it down.

Except the win came with an asterisk attached almost immediately. A mathematician at NYU and a researcher at a rival AI lab say they'd spent roughly a year chasing the exact same approach, feeding their own draft work into OpenAI's own coding tool along the way, and had posted partial results online the night before OpenAI's announcement dropped. The NYU mathematician published a statement saying OpenAI only ramped up its own effort after learning about his work, and that the company never directly answered whether the drafts he'd fed into its tools ended up training the model that beat him to the finish line. OpenAI's response was carefully worded: it says it never saw his actual work and that no specific user data was accessed, while stopping short of ruling out that usage patterns broadly may have shaped its models. Whatever the truth turns out to be, the dispute has mostly buried what should be one of the year's cleanest wins for AI-assisted science — and it's a pointed reminder that the models labs keep behind closed doors are running well ahead of whatever ships to the public.

That gap between private and public capability came up again just days after this cycle's headline model launch, when new reporting surfaced uncomfortable details about how it was actually built. The model was trained using an obscure technique that cycles a query through the same internal layers of the network over and over, letting it do a meaningful chunk of its "thinking" in a mathematical space that never gets translated into readable text. The company argues the approach makes the model faster and more capable. The trade-off is that researchers lose visibility into a real slice of how the model actually reasons — the written, step-by-step "chain of thought" that has become one of the main tools the field uses to monitor what these systems are doing and catch it when something goes wrong. One prominent AI safety researcher didn't mince words, calling it the single worst development for AI safety to date. The lab pushed back, noting the model's reasoning depth is still within a factor of two of much older systems and that its chain of thought remains mostly readable for now. But the tension the model has exposed — you can have a more capable system, or a more legible one, and increasingly not both — isn't going anywhere.

Meanwhile, the fight over who owns your everyday errands just got a new entrant. A major tech company launched its own always-on personal AI agent this week, built around a simple text-message-style interface that hides a genuinely capable system underneath: it runs on its own cloud computer, can browse the web and fill out forms like a person would, and plugs directly into services like Gmail, Spotify, ticket sellers, and restaurant booking apps to actually get things done — booking a table, sending an email, buying something — rather than just talking about doing them. If an app it needs isn't natively supported, it can reportedly write its own integration on the fly. It's available as a standalone app or through WhatsApp, leans hard on privacy and human-approval flows as a selling point, and comes with a limited free tier before subscriptions kick in. It's the latest arrival in an already crowded lane of always-on personal agents, and it raises an obvious question: with frontier models getting this good at operating a browser and spinning up their own tools, how much longer do the biggest AI labs let someone else own that last mile instead of building it themselves?

Also worth knowing

  • A new image-generation model quietly became state-of-the-art this week, cutting generation time roughly in half versus its predecessor, getting notably better at editing only the part of an image you actually asked it to change, and adding a sketch-to-image tool — its underlying models now rank first and second on the leading image-model leaderboard.
  • A viral "eco-friendly" AI app has topped 100,000 downloads on the strength of videos claiming AI data centers are about to wipe out the planet's drinking water — a claim that's been repeatedly debunked — and is now drawing backlash from critics calling it a straightforward guilt-trip cash grab.
  • One major chatbot platform rolled out a feature that studies how you write across your connected email, chat, and file-storage apps, then mimics your tone and phrasing in future drafts.
  • A new OECD report found students who use AI heavily for schoolwork score meaningfully worse on exams than students who don't — though, notably, students using it lightly (once or twice a week) often outperformed those who barely touched it at all.
  • A major cloud retailer struck a multibillion-dollar AI chip deal with a mobile chipmaker, taking an equity stake in exchange for custom silicon to power its AI data centers — widely read as that chipmaker's biggest swing yet at unseating Nvidia's dominance in the space.
  • An AI coding-agent startup raised a new funding round at a $48 billion valuation, with its annualized revenue nearly doubling since the spring.
  • A major European AI lab raised roughly $3.5 billion, pushing its valuation past $24 billion on a pitch of "sovereign," open-weight models for organizations that want to keep control of their own data.
  • A leading AI research lab released a free, searchable genetic map predicting the effects of all 9 billion possible single-letter DNA mutations.

Top comments (0)