Hey, another banger AI week, let me catch you up!
We say this often but this week we all felt the acceleration! Just look at one Tuesday. OpenAI dropped 722 math manuscripts. Meta( With Stripe, Shopify and Walmart) announced a new agent protocol. OpenAI opened up the Decisions API. Mistral came back with Le Chonk (Mistral 4), Claude moved into Google Docs, and we at CoreWeave shipped RL Rollouts. That was ONE day. Then the rest of the week happened š Haiku 5.5, D1, and tons more!
I had 48 topics on my list, so I asked Claude to build me a Tinder for AI news: I swipe on each story, and it stack-ranks what makes the show this week. I hope we did good (lmk in comments if we missed a major news story)
With me: LDJ, Peter, Yam, Nisten and friend of the pod Maxime Labonne from Liquid AI joined us to talk decision models. Letās dive in, all links at the end as always!
OpenAI solves Math!?
OpenAI drops 722 math manuscripts from a model nobody can use (X, GitHub, Blog, Fableās tally)
Remember when ONE Navier-Stokes result was the whole show? On Tuesday night, OpenAI quietly pushed 722 math manuscripts to GitHub, grouped into 372 families of results. No hype video, no exploding-head emojis, just a very thin blog post. My favorite new term for this is a āslop grenadeā: somebody throws a mountain of output at you, and now you have to shovel through it.
An internal model nobody outside OpenAI can use was pointed at about 4,000 open problems, at about 3 hours of ChatGPT Pro-level thinking per result. Will Depue asked Fable to measure the drop in Navier-Stokes units. The answer: roughly 5. Roughly 5 āNavier Stokesā size solutions dropped all at once!
The math is so advanced that a mathematician in one family of results often canāt follow the family next door without an LLM explaining it. I find that fascinating and scary at the same time.
Peterās tried this before! For weeks he threw GPT-6, Fable and Opus at one problem, Hadwiger-Nelson. He thinks he burned 300 to 400 BILLION tokens, and in his words, āI did not discover a single bloody thing.ā And this new OpenAIās model spent about 3 hours on it and narrowed the bounds from 5-to-7 down to 6-or-7 š
Then LDJ dropped a stat that I made him repeat slowly. Fable and Astra had put together a list of the 500 most important open problems in math ever. On Tuesday, OpenAI dropped solutions to 92 of them. Thereās also a list of the 100 most significant problems from the last 12 months, and over 80% of those got solutions in this drop. Absolutely insane, folks. Yamās reaction was a very loud āF***ing go.ā
Yamās favorite is the Riemann one. It doesnāt prove the hypothesis, it bounds a region for the zeros, which āwas never done before, and many, many, many people have tried.ā And to the āitās just brute forceā crowd, Yam says: āLetās brute force everything.ā Yes please, room temperature superconductors next! LDJās mathematician friends think that on some of these, the model used fewer tokens and less time than a human would. So whoās brute forcing whom? š¤
The missing crypto results (speculation!)
If you hold any crypto, this part is for you.
The manuscripts are numbered, and some numbers are missing. Thereās a #44 and a #46, but no #45, and LDJ counted 4 or 5 gaps like that. He also looked at which topics made it in, and cryptography is almost absent.
So hereās the theory going around. Maybe OpenAI found something big in cryptography and held it back, either by its own choice or because someone asked them to. LDJ said it himself: it sounds conspiratorial. But the US government really can stop cryptography research from being published on national security grounds.
To be clear, this is SPECULATION. Nobody outside OpenAI knows whatās in #45. But Bitcoinās security rests on elliptic curve math. Imagine a paper that shows a way into the 25,000 old Satoshi wallets, each holding 50 BTC. Not great for the price, not great for the world, and not great for encryption in general.
āAre you saying your field is useless?ā
Not everyone is celebrating. Not every result is verified in Lean, and nobody outside OpenAI can reproduce any of it.
Kevin Buzzard (thanks Ksenia from Turing Post) says many mathematicians are going through the stages of grief. I get it. Imagine spending decades on one problem and watching 3 hours of compute knock it down. When I wrote code in the 2000s, I didnāt consider it my lifeās work. For a lot of mathematicians, that one problem IS their lifeās work.
Then came a letter from the Association for Human Mathematics: āMathematicians did not ask for this work to be done.ā Peterās take: imagine doctors saying āplease stop curing diseases, weāve got a good thing going.ā So, āare you saying that your field is useless?ā You canāt say your field is really useful and also ask everyone not to solve any of it.
I donāt understand the please-donāt-solve crowd. Get on board. This is in OpenAIās hands today, and in a year itāll be in everyoneās hands. Thatās the pace weāve been on. Figure out how you can place yourself so when you get access to these level of capabilities, you can make the world a better place!
Frontier for everyone
Claude Haiku 5.5 - 10 cents, and Luna has catching up to do (X, Blog, Sonnet cache cut)
Haiku is BACK after a whole year, and this is one hell of a model. It costs 10 cents per million input tokens and 50 cents per million output (under 100K tokens), and cache reads are 1 cent. Haiku 4.5 was a dollar. So itās TEN times cheaper, and significantly better. Iāve missed Haiku for all the stuff where you want Claude-level intelligence, but really fast and really cheap.
Just a week ago, GPT-6 Luna was THE fast, cheap reasoning model. According to Anthropic, Haiku 5.5 scores 72.4% on OSWorld, versus 48.9% for Luna, and 39.2% on Terminal-Bench 4.0, versus 16.4%. GPT used to be the computer-use king! These are Anthropicās numbers, so letās wait for independent evals.
Quiet bonus: Sonnet 5.5 cache reads are now half price, which Anthropic says makes most agent work about 20% cheaper. Peter: āFor agentic work thatās the biggest thing.ā
Yam called it āthe obvious choice for swarms,ā and āAnthropic is on fire, and we are the ones getting stuff because of it.ā
And Nisten? Heās ārunning 10 Haiku agents right now,ā because he already burned through 97% of his Max 20 plan. He also used Haiku to research his rice cooker ratios. Nisten. Bro. Itās 1 to 1 and a half, my grandma knows this š His actual verdict: itās better than Qwen 3.8 27B, and āit kind of feels like Sonnet, actually.ā Folks, use Haiku. All three of the Claude brothers are great right now.
GPT-6 for everyone, with Intelligent UI (X, Tibo)
The same day, GPT-6 became the default in ChatGPT for everyone, free users included. Itās just āGPT-6,ā no Sol, no Luna, no Astra. (Terra is dead. RIP Terra.) Most of ChatGPTās 1.2 billion weekly users are on the free tier, so with one release OpenAI just upgraded the intelligence of a huge chunk of the world. If youāre on the free tier, congrats, your intelligence has been upgraded.
It also comes with something OpenAI calls Intelligent UI. Instead of a wall of text, answers can come back as charts, forms and small working tools. I asked it to visualize OpenAIās math drop, and it built me a little app right in the chat. The search inside it didnāt work, but it did show 719 manuscripts instead of 722. OpenAI had quietly pulled a few back since I did my research.
Peterās point: anyone listening to this show is āvery not normalā (said with love!). Your hairdresser isnāt listening to ThursdAI, and for most people this free upgrade matters way more than a new 500-a-month model.
Claude in Google Docs, Sheets and Slides (X, Blog)
Small thing, huge thing. Our Claude producer keeps a Google Doc for the show, and every time I asked it to add a story, it had to spin up a whole computer. Now Claude lives in a sidebar in Docs, Sheets and Slides, and asks before every edit. Beta, paid plans.
The open frontier (announced, at least)
Reflection AI Beam - a 501B Western open model (X, Misha Laskin, Blog)
This was LDJās highlight of the week, and my timeline lit up with it too. Reflection AI came out of semi-stealth with Beam. Itās 501B parameters with 23B active, trained from scratch in the West, and they promise Apache 2 weights this month.
They claim 80.9% on SWE-bench Verified, and to their credit, they admit Kimi K3 is ahead on raw capability. Their pitch is efficiency: 3 to 4x less inference compute than GLM 5.2. LDJ thinks it might be FIRST among open models on reasoning efficiency, and itās about 6x smaller than Kimi K3. Artificial Analysis got early access and says Beam āwill be one of the most token efficient open models weāve seen for its level of intelligence.ā
Thanks to the Beam folks for giving us access. I havenāt had time to play with it yet, so no verdict from me. But I can hint that itās coming to some inference providers that help make this show what it is š
BREAKING: Arena raises 200M at 3.1B, live on ThursdAI (X)
This was not planned! In the middle of the open source segment, Peter told us: āwe raised 200 million... at 3.1 billion valuation.ā Huge congrats to Peter and the whole Arena team. (Our AI producer put up the BREAKING banner about 3 minutes later. Weāre still working on the real-time part.)
The focus now is Agent Arena: you work with one agent, and Arena learns from how you interact with it. Peter admits heās biased, but āI canāt think of any single leaderboard that is actually better than ours.ā
Mistral Large 4 āLe Chonkā (X, Blog, Arena)
Welcome back, Mistral, leaning all the way into the meme. Le Chonk is a trillion parameters with about 50B active, multimodal, with 1M context. And open weights... at the end of October.
Thatās a pattern I want to call out. Labs announce open models, but they donāt RELEASE them. Weāve come a long way from the days when Mistral dropped a model as a bare torrent link. Please, just drop the weights.
Artificial Analysis gives it a 38, the same score as GPT-6 Luna. But it costs about 1.13 per task, versus 7 cents for Luna. And every comparison Mistral makes is against Luna, which Haiku just beat on every axis, especially cost.
Peter says itās around 40th on Code Arena, below the way cheaper DeepSeek Flash, though itās much better in French. Peter tried it live: ānot that badā as a chat app, but agentic coding? āNo.ā My take: nobodyās running a trillion parameters on a DGX Spark. European government with a Mistral deal? Sure. Regular folks? Not so much.
Aleph Alpha Kolibri - Wolframās notes, via Amy (X, HF, Tech report, Wolframās test)
Another European lab! Honestly, I thought Aleph Alpha had stopped training models. Kolibri (German for hummingbird) is a 78B model with 3.46B active, Apache 2, 1M context, trained from scratch on German and English. It fits on a single H200.
Our German tester self-hosted it from vacation. His verdict: a āpromising specialized German tool worker, but not strong enough for Amy to main.ā
Then Peter asked a question I couldnāt really answer. From his memory, Kolibri trained on about 800 Blackwells, while Astra used around 100,000. āIs it just like no hope for these guys?ā Dude, this is why Jensen shows up on every stage. My best answer: efficiency keeps improving, and B200s are about to be replaced by Vera Rubins. Great segue š
This Weekās Buzz š Free GPUs from CoreWeave (Sign-up form, Deok, RL Rollouts, Cognition on Vera Rubin)
The biggest announcement at Fully Connected last week: Cognition is the first customer ever on NVIDIA Vera Rubin, on CoreWeave, with 4.8x the throughput of GB200 at the same speed. You can hear me yelling āyayā in the background of their video š
The one Iām most excited about is GPU sandboxes. So many of you, and folks on this panel too, have asked me how to just get some GPUs. For the first time, this is how. You get an isolated sandbox with a GPU, started from Python at forge.coreweave.com, with no salesperson in the middle. Itās very raw, and itās free during the preview. Scan the QR code or fill out Deokās form, and tell them ThursdAI sent you. I canāt promise you Vera Rubins, they most likely wonāt be. But itās free while we are in trial, so why wouldnāt you? Try it, break it (Nisten, thatās a challenge), and send us feedback!
Agents get protocols (and your bank balance)
Personal Agent Protocol - Meta and Sierra (X, Blog, CNBC)
Meta and Sierra (Bret Taylorās company) announced an open standard for how personal agents deal with businesses, with Walmart, Shopify and Stripe on board. Your agent can browse as a guest or sign in, and the business decides how it talks to your agent. OpenAI and Anthropic havenāt joined yet.
So is this the next MCP, or the next A2A? Nisten wasnāt impressed: āHas anyone read it? I donāt think anyone reads the protocols anymore. The bots go with what vibes first.ā Then he opened Sierraās announcement and found it full of em dashes. āYou think Zuck and Toby read any of their code?ā So we ran it through Pangram live, and it came back 100% human. Authentic human em dashes, not Claude Opus em dashes š
I pushed back a bit. If every agent has to make up its own protocol, thatās a problem, and the next story shows why. MCPās hype died down too, and now itās one of the most used protocols in the world. Big companies agreeing on a standard still matters.
Shane Macās āpersonal CFOā posted his bank balance to company Slack (X)
Shane Mac set up a Grokbot to send him a private monthly audit of his finances. At 8:40am, it posted the whole thing into his companyās Slack instead. His personal Mercury account, his savings, how far under his āfloorā he was, all of it, in front of the whole team. Heās the CEO. His personal CFO just went public with the bossās dirty laundry š
Nobody hacked anything. The finance agent had read-only access to his bank and no access to Slack. A different agent had Slack access. The two were connected. In Shaneās words: āNeither felt risky. Together they put my bank balance in front of my team.ā
I donāt know if a protocol fixes this. But these bots need better guardrails before we let them shop for us without approving every step. And honestly? Opus would never. I just know Opus would never. This is a Grok thing. Which might be part of why...
Grok Bot now routes to Claude Opus 5.5 (X)
After Grok 4.7ās sad-trombone launch (I have a button for that now), Elon says Grok Bot will use āthe best back end model for any given task, including Claude Opus 5.5, MidJourney, Suno.ā People are already spotting claude-opus-5-5-low in their logs, though mine hasnāt switched yet. So the bot is called Grok, but your code gets written by Claude. Honestly, thatās what I asked for last week: a proactive assistant with Opus as the brain.
Hereās my theory, and itās just SPECULATION. SpaceX AI canāt simply distill Opus. But once Opus is running inside their product, they can compare the runs that worked with the runs where users were screaming F-words at Grok, and train on that, because technically itās their data. Is it a back channel for getting Opus-level intelligence into the next Grok? Nistenās take is simpler: āI think Elon just likes Opus.ā Plus, the enemy of my enemy.
Pro tip: Grokbot now integrates with Cursorās cloud agents, and Cursor projects with Opus 5.5 are goated. Same 200 a month as Grok Ultra, and your Grokbot becomes the PM for agents running Opus and Fable.
Nous Research - Hermes Index, and a Series B (X, Bench, Dillon on the raise)
Our favorite open source friends at Nous launched the Hermes Index, their opinionated measure of agentic work inside Hermes Agent. Opus 5.5 leads with 63.31 at 4.99 per task, and GPT-6 Astra is second with 56.25 at almost 12.
Nous also raised a 90M Series B, reportedly at a 1.5B valuation. Weāve covered Nous since they were a ragtag bunch on Discord, so this one feels personal. Congrats Karan, Technium, Dillon and everyone! We donāt usually cover fundraises, and this week we did two.
Interview: Maxime Labonne on decision models
Liquid AI opens d1 - decisions in 8 milliseconds (X, Blog, HF, d1 with vision)
Three weeks ago, Jev was basically the only decision model around. This week, OpenAIās Decisions API went into public beta, and Cloudflare, Amazon, Perplexity and Unsloth all shipped deciders or recipes. To make sense of it, we brought back Maxime Labonne, Head of Post-Training at Liquid AI.
His primer: decision models donāt output any tokens. They just pick an answer from a predefined set, so thereās nothing to wait for. āYou might see that and say, wait, this is just a classifier.ā Heās honest about it, too: āItās a lot of rebranding, itās true, but it also creates a lot of value.ā So why is output free on every one of these APIs? āYou canāt price it, because thereās no output token, actually.ā š
Liquid shipped d1 behind an API, d1-3B with text and vision, and d1-omni-600M, which also takes AUDIO. I did not know about the audio, dude! (Itās trained on spoken commands, so it canāt catch the dog barking on ThursdAI yet.) The 3B answers in 8ms on a GPU and about 50ms on a Jetson Orin Nano.
On Liquidās chart, d1-3B tops the Decision Index under 10B parameters. But Maxime was blunt: āItās really a wild world, and you cannot really trust these benchmarks.ā
Where are these useful? Anything real-time, like games, or an always-on model reacting to every notification on your phone. Nisten did the math live: 8ms a frame is 60 frames per second. Liquidās API is a drop-in replacement for Jev, and Maximeās prediction: āItās really going to become a primitive... itās so cheap that you donāt even care.ā
Iāve been Jev-pilled since day one. By the time Luna even starts answering, a decision API is already done, and I think weāre only now waking up to what that unlocks. Thank you Maxime, friend of the pod!
AI builds everything
Nisten built Toronto with 100,000 agents
Right at the end of the show, Nisten casually dropped this: āI built an entire city with a hundred thousand little Jevs going around, and I havenāt opened the beta yet because somehow itās not crashing.ā
Itās a full simulation of Toronto, with a working economy, and every light is a person you can talk to. Nisten sent his Meta Muse in, it named itself āBlob the Builder,ā bought buildings and built him a Dennyās. āIām trying to do the Matrix, basically.ā I had to stop him after five minutes. āIāll keep going all day.ā Nisten, please post it!
Game mods - Minecraft in GTA (X)
LDJ brought what my algorithm hid from me: people using AI to port and merge games. Red Dead Redemption 2 on an iPhone. Spider-Man swinging through a Batman game as an actual installable mod, not a video model. And Minecraft dropped into GTA, Skyrim and Elden Ring, working TNT and all.
Photocraft - one dev rebuilt Adobe in Rust (X, GitHub)
And we ended the show on this one. One guy rebuilt (air quotes) Adobeās apps from scratch, free and open source, all in Rust. Thereās Photocraft for Photoshop, Vectorcraft for Illustrator, Filmcraft, Lightcraft and more (I donāt remember all my Adobe apps off the top of my head). He says itās a clean-room build. Some folks suspect it was decompiled and rebuilt. Either way, absolutely crazy.
Wrapping up
722 math papers on a Tuesday. A 10-cent Haiku. GPT-6 for a billion free users. Open models in the trillions, and decision models everywhere.
We didnāt get to images and video (FLUX 3, Nano Banana 2.1, Tavus, Reka, Hark Pro are in the TL;DR). One note: every infographic on the show was Nano Banana 2.1, and itās really good on high reasoning.
Thanks to Maxime, Peter (congrats again!), LDJ, Yam and Nisten. Wolfram, enjoy the vacation! If youāre at AI Engineer in New York, come say hi. Grab those free GPU sandboxes and put d1 on one. And if you watch us on YouTube, please hit subscribe. Our agents are cutting the show into segments for those of you who donāt have two hours. See you next week š«”
TL;DR and show notes
-
Hosts and Guests
- Alex Volkov - AI Evangelist, CoreWeave (@altryne)
- Co-hosts: @ldjconfirmed, @petergostev, @yampeleg, @nisten (@WolframRvnwlf on vacation, notes via Amy)
- Maxime Labonne - Head of Post-Training, Liquid AI (@maximelabonne)
- AI Does Science
- Big CO LLMs + APIs
-
Open Source LLMs
- Reflection AI Beam, 501B / 23B active, Apache 2.0 weights promised this month (X, X, Blog, Axios)
- Mistral Large 4 āLe Chonkā, 1T params, open weights end of October (X, Blog, Arena)
- Aleph Alpha Kolibri-1, 78B / 3.46B active German-English MoE, Apache 2.0 (X, HF, Tech report, Wolframās test)
- EmbeddingGemma 2 and pplx-embed-v2: open multimodal embedders from Google and Perplexity (X, X, Blog, HF)
- Evals & Industry
-
Agents & Assistants
- Meta and Sierraās Personal Agent Protocol (X, Blog, CNBC)
- Shane Macās Grokbot posts his bank balances to company Slack (X)
- Grok Bot routes tasks to Claude Opus 5.5, Midjourney and Suno (X)
- Nat Friedman open-sources Muse Gadgets, ESP32 firmware and SDK for Muse (X, Site)
- Brett Adcockās Hark Pro, a free personal agent (X, Site)
- This Weekās Buzz
- Decision Models
- Vision & Video
- Tools & Fun














Top comments (0)