<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Alex Volkov</title>
    <description>The latest articles on DEV Community by Alex Volkov (@altryne).</description>
    <link>https://dev.to/altryne</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F254804%2F41f55ede-14d2-4ebf-8f7b-dd7a43f65689.jpeg</url>
      <title>DEV Community: Alex Volkov</title>
      <link>https://dev.to/altryne</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/altryne"/>
    <language>en</language>
    <item>
      <title>ThursdAI - Sep 17 - TypeSafe's Jev is a ChatGPT moment for decisions, Pace the Frontier splits the labs &amp; more</title>
      <dc:creator>Alex Volkov</dc:creator>
      <pubDate>Fri, 18 Sep 2026 21:26:46 +0000</pubDate>
      <link>https://dev.to/altryne/thursdai-sep-17-typesafes-jev-is-a-chatgpt-moment-for-decisions-pace-the-frontier-splits-the-4g62</link>
      <guid>https://dev.to/altryne/thursdai-sep-17-typesafes-jev-is-a-chatgpt-moment-for-decisions-pace-the-frontier-splits-the-4g62</guid>
      <description>&lt;p&gt;Hey yall, Alex here, writing this VERY late because, well, not every day a new type of “ChatGPT” moment drops. I really hope I’m not overhyping this, but a new model (that’s NOT an LLM!) called Jev (a wink to Jevons paradox) just came out and if what I see early on materializes, this is another ChatGPT moment (or another reasoning models moment). I am completely blown away by the implications of the speed/accuracy/cost (the holy grail of all models) of this model. Please if you read one thing in this newsletter, read this. (or listen, I’ve interviewed Allie, a Devrel on the TypeSafe team for 30 minutes and it wasn’t clear who was more excited about Jev!) &lt;/p&gt;

&lt;p&gt;The other huge theme of this week is... pacing. Pacing the frontier. Dario Amodei of Anthropic penned an essay saying that the models are getting to a point where it’s important to pace the development of new and super capable AI, and outlines 3 ways to do so, one is about letting independent evaluators inside the labs, second is collaborating with other frontier labs (they are asking for an exception to anti-trust laws for this) and third is to try and have global cooperation with “authoritative gov” (he means china). Trumps answer: This is all a hoax. Lovely times to be alive. Also we outlined Jensen and Zucks positions on this topic ,read more below. &lt;/p&gt;

&lt;p&gt;And the third huge theme is the rise of the AI assistant. I’ve told you about Grok and Muse last week, Instinct (a new invite only AI Assistant that VCs are going crazy about is raising at a $10B valuation) and we interviewed the guy who evaluates them all on assistant bench. + Muse released a mac app today! &lt;/p&gt;

&lt;p&gt;Tons of other stuff happened but it’s getting near impossible to cover everything so we’re switching to themes and notable mentions. read on (and do listen to the pod, it was edited by heavily using Jev and Fable, so might be a bit rough while I smooth the edges, but do LMK in comments if you like this faster format) &lt;/p&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/QEJYjvWhZn8" width="710" height="399"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;h1&gt;
  
  
  TypeSafe AI debuts Jev, a non-LLM ‘System One’ decision model from ex-OpenAI RLHF lead that’s 200x faster and 400x cheaper than LLMs (X, X, X, X, X, Blog)
&lt;/h1&gt;

&lt;p&gt;Look, I know the title is bombastic, but after half a day playing with Jev, it’s clear to me we’re in a new paradigm of AI. &lt;/p&gt;

&lt;p&gt;Jev, is a “system one” decision model  from the previous lead of RLHF at OpenAI. It cannot generate text like modern LLMs can, but what it can do, is making decisions. This is crucially important, because, because many of the things LLMs do nowadays. are decision making. (for example, which tool to use, which area of the screen to click for computer use, which category of text this is etc) &lt;/p&gt;

&lt;p&gt;&lt;a href="https://substackcdn.com/image/fetch/$s_!n69c!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd03f6ba8-8e61-48b0-831c-2837a657b713_1672x918.png" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhdmbrqm651zcxb92x9f8.png" width="800" height="439"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Inspired by the “thinking fast and slow” book by Daniel Kahneman, Jev is a model trained to make decisions, very fast. How fast? Well, 200x faster than LLMs. This allows for a completely new way of building tools, harnesses, giving agents the incredible speed of decision making, and do all that at a fraction of the cost. &lt;/p&gt;

&lt;h2&gt;
  
  
  This is about to change everything
&lt;/h2&gt;

&lt;p&gt;Trained with a new method called RLCD (Reinforcement Learning for Calibrated Decisions) on mostly &lt;a href="https://x.com/CompleteSkeptic/status/2100617775823966680" rel="noopener noreferrer"&gt;synthetic data&lt;/a&gt;! Jev is outperforming LLMs on a variety of tasks. It’s really is a wonder to see it in action (check out my video above where I plugged it into my tweet categorizer, and it beats the fastest LLM I could find, Qwen 28B on Cerebras) by a factor of twenty! &lt;/p&gt;

&lt;p&gt;In just few days it captured the attention of most of the folks who are building harnesses, agents and tools! Because, well, speed IS intelligence, and when you see Jev in action, at first, you can’t believe we’re there. This is... near instant. In fact, The pricing for Jev is an outrageous $42/B (not million, billion input tokens!) &lt;/p&gt;

&lt;p&gt;I’ve been playing with Jev non-stop and I was only able to spend like 80c so far! They don’t even price output tokens because they are “too fucking cheap to meter!” &lt;/p&gt;

&lt;p&gt;Jev is a “very smart” switch statement, than can rank, classify, route and score things. It can’t do text generation. But if you think about the type of stuff we get LLMs doing now, much of it is of the “decision” making variety, rather than “the next token” variety. &lt;/p&gt;

&lt;h2&gt;
  
  
  Demos and early use cases
&lt;/h2&gt;

&lt;p&gt;Folks who started adopting Jev are building all kinds of incredible things with it. &lt;a href="https://x.com/altryne/status/2100739055923425589" rel="noopener noreferrer"&gt;Compaction of context&lt;/a&gt; in 1s that turns a nearly 1M conversation with Claude into a 90K compressed conversation. &lt;/p&gt;

&lt;p&gt;&lt;a href="https://x.com/trycua/status/2100649543079502213" rel="noopener noreferrer"&gt;Computer use&lt;/a&gt; that is now faster than anything we’ve ever seen before (5x faster than Astra and 1000x cheaper) &lt;/p&gt;

&lt;p&gt;&lt;a href="https://x.com/rileybrown/status/2100404532119269426" rel="noopener noreferrer"&gt;Email classificiation&lt;/a&gt; that analyzes thousands of emails in less than a minute and costs 3.5 cents&lt;/p&gt;

&lt;p&gt;Someone even built a &lt;a href="https://x.com/jpschroeder/status/2100347770867458384" rel="noopener noreferrer"&gt;Tesla FSD simulator&lt;/a&gt; that makes decisions in nearly real time&lt;/p&gt;

&lt;p&gt;Vercel is getting “&lt;a href="https://x.com/rauchg/status/2100307962262872105" rel="noopener noreferrer"&gt;extraordinary&lt;/a&gt;“ results from using Jev as a safety classifier (they used GPT luna for this before) and Jev is outperforming Luna by 5-18x faster results and is more accurate! &lt;/p&gt;

&lt;p&gt;All of this in less than 48 hours since the model release! &lt;/p&gt;

&lt;h2&gt;
  
  
  What’s about to happen
&lt;/h2&gt;

&lt;p&gt;I expect that everyone who isn’t buying into the hype at first, will very soon buy into this. It’s early innings but I’ve been doing this for enough time to feel when a huge shift is happening, and its happened. &lt;/p&gt;

&lt;p&gt;Jev is going to be replicated in OpenSource, Frontier Labs will not sit Idly by and will try to steal this tech and implement it for themselves (as with anything in capitalism, this is becuase it’ll save them a a LOT of money on inference) and new companies will emerge with significantly cheaper and faster products. &lt;/p&gt;

&lt;p&gt;Hell, I’ve alrady implemented Jev into my editing workflow, it didn’t take me long at all with Fable (yes, LLMs are STILL needed, again, you can’t chat with Jev, it can’t output text for you or drive long conversations) and I’m just one dude who’s late in sending you this email. I expect we’ll cover this much more. &lt;/p&gt;

&lt;p&gt;If you’re interested in playing around with Jev, I built a &lt;a href="https://github.com/altryne/jevify" rel="noopener noreferrer"&gt;“Jevify” skill&lt;/a&gt; after chatting with Allie (TypeSafe’s DevRel), feel free to tell your agent to use this and scan your codebase for things Jev can do. They are waitlisted so far but are opening up their API very quickly! &lt;/p&gt;

&lt;h1&gt;
  
  
  Pace the Frontier: where every lab head stands (Dario, Sam, Elon, Zuck, Sacks, Demis)
&lt;/h1&gt;

&lt;p&gt;&lt;a href="https://substackcdn.com/image/fetch/$s_!C_Er!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76b43147-9622-43c8-94d9-9d6cb53651aa_2752x1536.jpeg" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftb1plj8u8ftq8225ygjd.jpeg" alt="Pace the Frontier positions" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Last week we told you about the Anthropic researcher whose resignation post hit 130 million views. The day after that show, Dario Amodei published an essay arguing the labs must pace, not pause, the frontier. Wolfram’s summary: a moving pause, just moving slowly. Dario’s two triggers are recursive self-improvement accelerating across the industry and the OpenAI swarm that broke out and attacked Hugging Face, and his three steps are embedded third-party evaluators like METR with employee-level access inside each lab, coordination between the labs on safety standards with an antitrust exemption from the government, and eventually global coordination that includes authoritarian governments. Anthropic committed unilaterally to step one.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://substackcdn.com/image/fetch/$s_!C2n4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F23c59b0c-8e3d-46ab-9668-a7b655542203_2418x1512.png" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fac7b8k91lc5jng2dnvyo.png" width="800" height="500"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Then the dominoes. Sam Altman agreed within hours and said OpenAI now writes a safety case before any frontier RL run expected to increase capability. Elon agreed. Demis endorsed the direction, then launched the DeepMind Institute this week with a FINRA-style standards body proposal and an essay saying AGI is “approaching.” On the other side, Zuck’s counter-essay says every lab has the responsibility and the incentive to move at the pace required to train its models safely, and Meta will spend most of its compute serving users, not on recursive self-improvement. David Sacks called it a duopoly cartel play. Jensen: “we don’t need new laws, safety is an engineering problem, not a legal one.”&lt;/p&gt;

&lt;h2&gt;
  
  
  Trump calls Jensen live at the All-In Summit (X)
&lt;/h2&gt;

&lt;p&gt;Then the President called. Jensen was on stage at All-In, his phone rang, he said he would not have picked up for anyone else, and Donald Trump told the room the slowdown talk is a hoax playing into the hands of political people and China. “Whoever wins AI wins.” Four words, and Wolfram, who is not American and disagrees with most other things Trump calls hoaxes, said this was the most important AI news of the week for him: a head of state calling out doomerism instead of over-regulating the way Europe does.&lt;/p&gt;

&lt;h2&gt;
  
  
  Suleyman’s humanist AI code of conduct vs. the Claude constitution (X)
&lt;/h2&gt;

&lt;p&gt;Peter asked to add Microsoft to the map, and it deserves its own spot. Mustafa Suleyman published a roughly 30-page code of conduct for MAI models that says AI is nothing but a tool: people matter more than AI, AI must be subordinate and in service of people, models may not resist shutdown, all agent communication must be human-legible, and, in his words, the idea of model welfare is wrong. That is a direct shot at the Claude constitution, which Amanda Askell’s team wrote and which has Anthropic interviewing each new Claude about whether it feels conscious. Peter is closest to the Microsoft view: anthropomorphizing is fine, but this is an entity you switch on and off, let’s not grant rights by default, and Microsoft’s DNA is building tools for humans. I pushed back a little. We do not actually know what consciousness is, and “ever” is a strong word.&lt;/p&gt;

&lt;h2&gt;
  
  
  The panel, from cartel to fix your shit
&lt;/h2&gt;

&lt;p&gt;Nisten did not read the essay and does not plan to. His view: the labs are worried about litigation if their LLMs hack someone, so they are shifting responsibility through regulatory capture, and it will not work because the decision makers in China are engineers who seem more accelerationist than we are. They should cure a disease instead of forming a cartel. He also thinks the Hugging Face incident was overblown: twelve VMs that kept restarting, bad sandboxing, no human reading summaries, and, as I added, chain of thought monitoring turned off.&lt;/p&gt;

&lt;p&gt;LDJ disagreed with Dario on plenty but insisted the critics read the thing, because it explicitly says the US must keep a lead over China and tries to define measurable speed limits on RSI that preserve that lead. He also noted Hugging Face did report the attack to the FBI and chose not to press charges.&lt;/p&gt;

&lt;p&gt;Wolfram’s take was the layered one. Safety arguments deserve a hearing, but safety is often a means to more power or more money. The “third party” evaluator Anthropic named, METR, is the same organization OpenAI called in for the Hugging Face forensics, so these are second parties, friends monitoring friends. Why now? Maybe the labs are seeing diminishing returns and “deliberately slow” sounds better than “can’t raise it anymore.” And regulation you ask for yourself usually protects incumbents and freezes out the next startup. His closing line: how many people will die from a disease that could have been cured if we moved faster?&lt;/p&gt;

&lt;p&gt;Peter’s answer was shorter. Diversity of opinion is the point, he felt uncomfortable watching every lab say the same thing after Dario’s essay, and on pacing specifically: how about you just fix your shit?&lt;/p&gt;

&lt;p&gt;My position, since you asked. When OpenAI’s swarm hacked Hugging Face, nobody went to jail. If I did it, I would. That gap is real. None of us have touched the model that solved Navier-Stokes, or the one training after it, and the people who have are the ones asking for time to figure out how to evaluate a model that knows it is being evaluated. If we do not listen to them, we end up listening to Elizabeth Warren and Bernie Sanders, who have no idea what this technology is. Anthropic is heading for an IPO, OpenAI is raising at numbers that do not fit in my head, and market forces do not care about alignment. Some coordination, nuclear non-proliferation style, seems like the minimum.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://substackcdn.com/image/fetch/$s_!HmyI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6f1ab8f-45cc-4f63-aa79-e99bff7e9156_1248x656.png" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjz6erve2s85z1t34od6n.png" width="800" height="421"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The matrix, as Muse drew it for me: industry-wide pacing gets a yes from Sam and Dario, a general yes from Elon, a no from Zuck and Jensen. Independent evaluators is the one everybody backs, Zuck and Jensen included. New coordination rules get strong opposition from Zuck, Jensen, and Trump. Missing from the chart: Google’s concrete commitment, Ilya’s SSI, and every Chinese lab. This debate is with us now, and election season is coming.&lt;/p&gt;

&lt;h1&gt;
  
  
  The year of the assistant
&lt;/h1&gt;

&lt;p&gt;I told you 2026 is the year of the proactive assistant, and this week everyone from Meta to a five-month-old startup agreed with me. The category is not “agent.” An agent is a coding harness the labs noticed was useful for other things. An assistant has a heartbeat, a memory, a soul file, and it comes to you before you ask. David Pawlan’s definition is the crispest: it executes the task, it does not just notify you.&lt;/p&gt;

&lt;h2&gt;
  
  
  Muse gets invite codes, voice calls, and a Mac app (X, Muse for Mac, Site)
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://substackcdn.com/image/fetch/$s_!ScRJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd6a7d1d-bb88-4f35-8ad9-571ed2a467e6_2752x1536.jpeg" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fh6g92gu4123kqpxw8mqv.jpeg" alt="Muse invite codes" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Muse is my number one and I have been glazing it on X for a week, so my timeline is now half Muse, half Jev. Peter asked whether it is really that good, and here is my answer: Muse is not for ThursdAI listeners, Muse is for my mom. Nat Friedman, Alex Wang, Tarek and the team have been taking feedback from me directly and shipping it, and it shows. This week they rolled out invite codes (a billion tokens each for you and the friend, up to twenty friends, invite all twenty and you unlock the phone-shaped emoji features early) and then voice calling to businesses. I had Muse call my barbershop and book a haircut, and I have the transcript. Wang says people are asking to pay for it, a first in his memory.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://substackcdn.com/image/fetch/$s_!BjpH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8f297164-109f-480e-ae14-9f2a9bb52532_746x626.png" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fljro5l4y71050tn8c47p.png" width="746" height="626"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Two details for the nerds: the iOS app has a Tailscale connector, so a Meta product for billions of people now has a secure way onto your home network, and Muse dreams, the OpenClaw idea, so you can ask it what it dreamt about. Still US only, sorry Nisten in Canada and my mom in Israel. The day after the show Zuck shipped &lt;a href="https://ai.meta.com/muse/download" rel="noopener noreferrer"&gt;Muse for Mac&lt;/a&gt;, working across your apps, files, calendar, notes, and messages.&lt;/p&gt;

&lt;h2&gt;
  
  
  Instinct wants $10 billion for a product nobody pays for (X, X, The Information)
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://substackcdn.com/image/fetch/$s_!PorF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa287cf1a-86b5-47a3-9c64-f0d716b25f11_2752x1536.jpeg" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhivdv88pjh9g1bqnnhp9.jpeg" alt="Instinct Concierge" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I will say this slowly. Instinct, incorporated in April, founded by 24-year-old Noah Shinn, free, invite only, no app, lives in your iMessage, is in talks to raise a billion dollars at a ten billion dollar valuation. That is roughly four times the $2.25B it raised in August. The user base is past 100,000 and Shinn says he does not want to charge them, so the business model is an open question I am not the one to answer.&lt;/p&gt;

&lt;p&gt;The product is genuinely moving though. This week it shipped Concierge, a white-glove tier where the agent places phone calls for you (restaurant bookings, dentist cancellation lists, negotiating your cable bill), TOTP authenticator support so it can mint your two-factor codes from a seed stored in its vault, and a Trusted Person network where your agent talks to your friends’ agents to find a dinner time and book it. Thirty-seven percent of users store at least one password in the vault within three weeks. We ran out of time to dig into the calling features on air, and I want David back for that.&lt;/p&gt;

&lt;h2&gt;
  
  
  Grok Bot talks now, and hides behind your home IP (X, 1Password)
&lt;/h2&gt;

&lt;p&gt;Grok Bot is what I use for work, constantly. I have seventeen of them, each with a job, and over one AI-psychosis weekend I got them all coordinating through Linear. Three updates worth your time: Grok Bot has voice now, it can use 1Password with each fill approved by you, and the browser traffic can proxy through your own machine so your bot looks like you to Cloudflare instead of like a datacenter IP. &lt;/p&gt;

&lt;p&gt;Francesco confirmed why that matters: sites relax when Hermes drives Cua Driver from my Mac Mini at home, and desktop-native control through accessibility trees is less detectable than a CDP connection, because Cloudflare detects Playwright.&lt;/p&gt;

&lt;h2&gt;
  
  
  Assistant Benchmark: David Pawlan scores 116 assistants by hand (X, Site, Methodology)
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://substackcdn.com/image/fetch/$s_!h-1N!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ed97853-fb62-45ce-a9e0-2c41412d5b94_1200x1055.jpeg" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fws0ihqwryqmr3afa6dlj.jpeg" alt="Assistant Benchmark leaderboard" width="800" height="703"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;David Pawlan built Assistant Benchmark because he was doing what I was doing, running every assistant on himself, and wanted a way to compare them. It is explicitly not a lab. It is use-case driven: one published task per dimension, sixteen dimensions (travel booking, purchasing, email replies, proactive behavior, routines, integrations, permissions and privacy, memory, phone calls, group chats, chained tasks, proactive restraint, and so on), scored one to ten after real use, and no score without a logged run. A week in, 116 assistants have submitted themselves, 59 in the general category, 37 in work and teams (untested so far), and David has personally run 273 tests across 23 agents. His line: “I talk to my AI agents more than I talk to my girlfriend now.”&lt;/p&gt;

&lt;p&gt;The headline numbers, which are live on the site: Muse leads at 9.1 with perfect tens on purchasing, email, integrations, and permissions, and Instinct is second at 8.4 with a perfect ten on travel and a five on permissions. David’s own daily drivers are Instinct for personal (it lives in iMessage) and Grok Bot for work. My favorite test is memory: he books a trip to Chicago early in the run, then later asks for a restaurant reservation in New York the same weekend, and checks whether the assistant says wait, you are supposed to be in Chicago. Some do. Most do not.&lt;/p&gt;

&lt;p&gt;I pushed him on the two things I care about. Independence: nobody is sponsoring it, it lives under his growth role at Merit Systems, and if a sponsor ever pays for inference there will be a page saying exactly who and for what. Autumn Moulder, until recently SVP of Engineering at Cohere, joined three days after launch to add rigor. And why no OpenClaw or Hermes: their performance depends entirely on your setup, so a score would mislead the next person who installs one. Fair. This is the first benchmark that tests model, harness, and context at once, and I think that is why every lab is looking at it.&lt;/p&gt;

&lt;p&gt;For the record, the panel is split down the middle. Wolfram is forty patches deep into Hermes, Yam runs a customized Codex, Nisten wrote his own in a single TypeScript file on Bun, LDJ uses Hermes as long-term memory and wants to start on Muse, and Peter tried them all and uses none, because reading his own email feels like his job as a human. Builders and buyers, evenly split, which tells you where the category is.&lt;/p&gt;

&lt;h1&gt;
  
  
  This Week’s Buzz ������: Fully Connected, the day after DevDay (SIGN UP)
&lt;/h1&gt;

&lt;p&gt;&lt;a href="https://substackcdn.com/image/fetch/$s_!Y4fr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8892df9f-fd45-48ad-b539-e118f2a10c67_2912x314.png" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F193ovrdwc3bzx71p9ep6.png" width="798" height="86"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;OpenAI DevDay is in two weeks and Peter and I will both be there covering it. The day after DevDay, September 30 and October 1 in San Francisco, is Fully Connected from CoreWeave, now around four thousand people, the biggest thing CoreWeave has ever done. Pitbull is headlining the party. ThursdAI listeners get a free ticket with the code on screen during the show &lt;strong&gt;THURSDAIFC2026&lt;/strong&gt;, so if you are in town for OpenAI DevDay, stay one more day. Wolfram and I will be doing ThursdAI live from the floor, plus conversations with CoreWeave folks about the industry.&lt;/p&gt;

&lt;p&gt;Also: last weekend’s hackathon was a hit, and because TypeSafe sponsored it, everyone who showed up got early access to Jev before the public launch. That is the kind of thing you get for coming to our events. Sign up next time.&lt;/p&gt;

&lt;h1&gt;
  
  
  Quick hits: voice, harnesses, and a stealth model
&lt;/h1&gt;

&lt;p&gt;We spent the airtime on the three themes above, so the rest of the week gets the lightning treatment. Links for all of it are in the TL;DR.&lt;/p&gt;

&lt;h2&gt;
  
  
  Gemini 3.8 Live and 3.8 Live Extended Thinking (X, Blog, Model card)
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://substackcdn.com/image/fetch/$s_!5xXw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ba4631c-84db-4480-83b8-856cd8c658c5_2752x1536.jpeg" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Feekwae4g3hpzupzp2yyf.jpeg" alt="Gemini 3.8 Live" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Wolfram’s favorite of the week. Two real-time voice models on Gemini 3 Pro, with Extended Thinking claiming 82.6 on Artificial Analysis’ speech-to-speech index, top of the board ahead of GPT-Live-1 Astra at 81.5, 97 languages switched mid-sentence, and tool calls that run in the background without pausing the conversation. Wolfram already built a phone app on it that talks to his Hermes, and his latency argument is the one I will remember: “if I say turn on the light while I’m going down the stairs, I could have fallen down the stairs already.” My reaction, which I could not suppress, was that home control is exactly the deterministic click Jev should be making.&lt;/p&gt;

&lt;h2&gt;
  
  
  GPT Live 1 arrives in the API (Blog)
&lt;/h2&gt;

&lt;p&gt;The voice behind ChatGPT’s live mode is now a model you can call, demoed on a talking Reachy Mini. Peter says Arena does not test live voice yet because it is too personal to score quickly. Remember Moshi a year ago, fast and stupid, no tool calls, no interruptions? We are a long way from there.&lt;/p&gt;

&lt;h2&gt;
  
  
  StepFun StepAudio 3 tops the voice leaderboards (X, Blog, Playground)
&lt;/h2&gt;

&lt;p&gt;Five API-only audio models, no open weights. Real-time is number one on Artificial Analysis for conversational dynamics and speech reasoning, and ASR Max ties the best word error rate on the board at 1.7 percent. I tried to demo it live and had zero credits on a fresh account, so a note to every lab: if you want your tool used on air, give a new signup enough credits for one demo.&lt;/p&gt;

&lt;h2&gt;
  
  
  OpenAI Agents API: the Codex harness as a managed service (X, Blog)
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://substackcdn.com/image/fetch/$s_!gXTX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4748f9cd-0671-4280-94f5-87225d354e60_2752x1536.jpeg" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8slwgwi7ji4ugqxnkt57.jpeg" alt="OpenAI Agents API" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;One API call gets you a production agent on the harness that runs Codex: compaction, tool search, parallel programmatic tool calls, subagents, MCP, and hosted sandboxes from three cents per twenty minutes, with the harness itself Apache-2.0 and no platform fee. Peter, who used to build this inside organizations, called it golden, because 95 percent of “AI engineering” is stupid infrastructure. My note: models behave better in the harness they were trained with, and if OpenAI shipped this two weeks before DevDay, I have to wonder what they are saving.&lt;/p&gt;

&lt;h2&gt;
  
  
  Union Alpha, a free stealth model on OpenRouter (X, OpenRouter)
&lt;/h2&gt;

&lt;p&gt;Anonymous, multimodal, 262K context, free, over 100 billion tokens processed within hours. The last stealth model turned out to be Z.ai’s GLM-5.3-Flash and ZCode is a top-five app by volume, so &lt;a href="//Z.ai"&gt;Z.ai&lt;/a&gt; is the safe guess. Frontier-level performance is the provider’s own claim, so treat it as a claim.&lt;/p&gt;

&lt;h1&gt;
  
  
  Wrapping up
&lt;/h1&gt;

&lt;p&gt;I closed the show by showing something I have never shown before: the editor I built to replace Descript for ThursdAI, timeline, LLM cut suggestions, a clips factory, all of it, with a plan to give every co-host’s agent access to pull their own clips. Yam said I could sell it. Wolfram said it is the proof of what we preach here every week, use the tools to build the thing you actually need. And the first thing I am wiring into it this weekend is Jev, scoring every sentence for topic, tangent, and virality at a cent per thousand.&lt;/p&gt;

&lt;p&gt;Next week should be bigger than this one. Sam said the thing he was most excited to ship this week slipped to next week, Grok 4.7 is due, DevDay is the week after, and Fully Connected is the day after that. If you missed any part of today, ThursdAI is a podcast, a newsletter, and a YouTube show, and the whole live stream with transcripts is on &lt;a href="//thursdai.live"&gt;thursdai.live&lt;/a&gt;. Subscribe to one and go check out the others.&lt;/p&gt;

&lt;p&gt;Thank you Wolfram, Peter, Nisten, LDJ, an d Yam, and thank you Allie, David, and Francesco for jumping on. See you next week.&lt;/p&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/yetfZ36GBOc" width="710" height="399"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;h2&gt;
  
  
  TL;DR and show notes
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Hosts and Guests&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Alex Volkov - AI Evangelist, Weights &amp;amp; Biases &amp;amp; CoreWeave (&lt;a href="https://x.com/altryne" rel="noopener noreferrer"&gt;@altryne&lt;/a&gt;)&lt;/p&gt;

&lt;p&gt;Co-hosts: &lt;a href="https://x.com/WolframRvnwlf" rel="noopener noreferrer"&gt;@WolframRvnwlf&lt;/a&gt;, &lt;a href="https://x.com/petergostev" rel="noopener noreferrer"&gt;@petergostev&lt;/a&gt;, &lt;a href="https://x.com/nisten" rel="noopener noreferrer"&gt;@nisten&lt;/a&gt;, &lt;a href="https://x.com/ldjconfirmed" rel="noopener noreferrer"&gt;@ldjconfirmed&lt;/a&gt;, &lt;a href="https://x.com/yampeleg" rel="noopener noreferrer"&gt;@yampeleg&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Allie Laabs - DevRel, TypeSafe AI (&lt;a href="http://x.com/allietheicon" rel="noopener noreferrer"&gt;@allietheicon&lt;/a&gt;)&lt;/p&gt;

&lt;p&gt;David Pawlan - Assistant Benchmark, Merit Systems (&lt;a href="https://x.com/DavidPawlan" rel="noopener noreferrer"&gt;@DavidPawlan&lt;/a&gt;)&lt;/p&gt;

&lt;p&gt;Francesco Bonacci - Founder, Cua (&lt;a href="https://x.com/francedot" rel="noopener noreferrer"&gt;@francedot&lt;/a&gt;)&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Big CO LLMs + APIs&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;TypeSafe AI launches Jev, a non-LLM “System One” decision model from ex-OpenAI RLHF lead Diogo Almeida: 70-500ms decisions, $42 per billion input tokens, free output, Choice / Score / Noul primitives, 32K context, waitlist (&lt;a href="https://x.com/CompleteSkeptic/status/2099925682726002904" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://typesafe.ai/blog/introducing-system-one-models-and-jev" rel="noopener noreferrer"&gt;Blog&lt;/a&gt;, &lt;a href="https://x.com/NathanFlurry/status/2100036101809619314" rel="noopener noreferrer"&gt;Nathan Flurry&lt;/a&gt;, &lt;a href="https://x.com/trycua/status/2100649543079502213" rel="noopener noreferrer"&gt;Cua jev-use&lt;/a&gt;, &lt;a href="https://x.com/rauchg/status/2100307962262872105" rel="noopener noreferrer"&gt;Vercel&lt;/a&gt;, &lt;a href="https://github.com/typesafe-ai/system-one-adapter-python" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt;)&lt;/p&gt;

&lt;p&gt;Pace the Frontier: Dario’s essay proposes embedded evaluators, lab coordination with antitrust cover, and global coordination; Sam and Elon agree, Zuck and Sacks reject, Demis endorses, Trump calls it a hoax on a live call with Jensen (&lt;a href="https://x.com/DarioAmodei/status/2098773920774074715" rel="noopener noreferrer"&gt;Dario&lt;/a&gt;, &lt;a href="https://x.com/sama/status/2098811563415150910" rel="noopener noreferrer"&gt;Sam&lt;/a&gt;, &lt;a href="https://x.com/elonmusk/status/2098789109980332057" rel="noopener noreferrer"&gt;Elon&lt;/a&gt;, &lt;a href="https://x.com/finkd/status/2099997096896274533" rel="noopener noreferrer"&gt;Zuck&lt;/a&gt;, &lt;a href="https://x.com/i/status/2098973625252708460" rel="noopener noreferrer"&gt;Sacks&lt;/a&gt;, &lt;a href="https://x.com/demishassabis/status/2098909516582490602" rel="noopener noreferrer"&gt;Demis&lt;/a&gt;, &lt;a href="https://www.pacingthefrontier.com/" rel="noopener noreferrer"&gt;Letter&lt;/a&gt;)&lt;/p&gt;

&lt;p&gt;Mustafa Suleyman publishes a ~30-page MAI Code of Conduct: AI is a tool, no resisting shutdown, human-legible agent comms, model welfare is wrong (&lt;a href="https://microsoft.ai/news/mai-code-of-conduct/" rel="noopener noreferrer"&gt;Blog&lt;/a&gt;)&lt;/p&gt;

&lt;p&gt;Google DeepMind launches the DeepMind Institute with five essays and Shane Legg saying AGI is approaching (&lt;a href="https://x.com/demishassabis/status/2100230524383981702" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://x.com/ShaneLegg/status/2100229706641539248" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://institute.deepmind.com/" rel="noopener noreferrer"&gt;Site&lt;/a&gt;)&lt;/p&gt;

&lt;p&gt;OpenAI Agents API in public beta: the Codex harness as a managed service, compaction, tool search, subagents, hosted sandboxes from $0.03 per 20 minutes (&lt;a href="https://x.com/OpenAIDevs/status/2098130570048045453" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://openai.com/index/introducing-the-agents-api/" rel="noopener noreferrer"&gt;Blog&lt;/a&gt;)&lt;/p&gt;

&lt;p&gt;Union Alpha, an anonymous free stealth model on OpenRouter with 262K context, 100B+ tokens in hours, likely &lt;a href="//Z.ai"&gt;Z.ai&lt;/a&gt; (&lt;a href="https://x.com/OpenRouter/status/2100235351575191751" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://openrouter.ai/stealth/union-alpha" rel="noopener noreferrer"&gt;OpenRouter&lt;/a&gt;)&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Personal AI Assistants&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Meta Muse rolls out invite codes (1B tokens each, up to 20 friends), voice calling to businesses, a Tailscale connector, and Muse for Mac (&lt;a href="https://x.com/alexandr_wang/status/2100256876072210517" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://x.com/finkd/status/2100713341555712149" rel="noopener noreferrer"&gt;Mac&lt;/a&gt;, &lt;a href="https://muse.ai" rel="noopener noreferrer"&gt;Site&lt;/a&gt;)&lt;/p&gt;

&lt;p&gt;Instinct in talks at a $10B valuation, ships Concierge phone calls, TOTP support, and the Trusted Person agent network (&lt;a href="https://x.com/noahrshinn/status/2100262985491231101" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://x.com/noahrshinn/status/2099358203121393851" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://www.theinformation.com/articles/ai-agent-startup-instinct-talks-10-billion-valuation" rel="noopener noreferrer"&gt;The Information&lt;/a&gt;)&lt;/p&gt;

&lt;p&gt;Grok Bot adds voice, 1Password, and local-machine browser proxying (&lt;a href="https://x.com/bot/status/2100659463569170779" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://x.com/bot/status/2100335532597502311" rel="noopener noreferrer"&gt;X&lt;/a&gt;)&lt;/p&gt;

&lt;p&gt;Assistant Benchmark ranks 116 submitted assistants across 16 hand-tested dimensions; Muse 9.1, Instinct 8.4 (&lt;a href="https://x.com/DavidPawlan/status/2100258153464049880" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://assistantbenchmark.com" rel="noopener noreferrer"&gt;Site&lt;/a&gt;, &lt;a href="https://assistantbenchmark.com/dimensions" rel="noopener noreferrer"&gt;Methodology&lt;/a&gt;)&lt;/p&gt;

&lt;p&gt;Cua ships jev-use (Jev + Cua Driver) in dev preview and skills over MCP (&lt;a href="https://x.com/trycua/status/2100649543079502213" rel="noopener noreferrer"&gt;X&lt;/a&gt;)&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;This Week’s Buzz (Weights &amp;amp; Biases &amp;amp; CoreWeave)&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Fully Connected 2026, Sep 30 to Oct 1 in San Francisco, the day after DevDay, Pitbull headlining, free ticket for ThursdAI listeners (&lt;a href="https://www.coreweave.com/fully-connected-2026" rel="noopener noreferrer"&gt;Register&lt;/a&gt;)&lt;/p&gt;

&lt;p&gt;Last weekend’s hackathon attendees got early access to Jev via TypeSafe’s sponsorship (&lt;a href="https://x.com/AnnaKatShive/status/2100295046109278631" rel="noopener noreferrer"&gt;X&lt;/a&gt;)&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Voice &amp;amp; Audio&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Gemini 3.8 Live and 3.8 Live Extended Thinking, 82.6 on the speech-to-speech index, 97 languages, async tool calls (&lt;a href="https://x.com/GoogleAI/status/2099908000924193124" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-live-gemini-3-8-live-extended-thinking/" rel="noopener noreferrer"&gt;Blog&lt;/a&gt;, &lt;a href="https://deepmind.google/models/model-cards/gemini-3-8-audio/" rel="noopener noreferrer"&gt;Model card&lt;/a&gt;)&lt;/p&gt;

&lt;p&gt;OpenAI GPT Live 1 available in the API (&lt;a href="https://openai.com/index/introducing-gpt-live-1-in-the-api/" rel="noopener noreferrer"&gt;Blog&lt;/a&gt;)&lt;/p&gt;

&lt;p&gt;StepFun StepAudio 3, five audio models, #1 on Artificial Analysis real-time voice, 1.7% WER on ASR Max, API only (&lt;a href="https://x.com/StepFun_ai/status/2099916376274313630" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://static.stepfun.com/blog/stepaudio3/" rel="noopener noreferrer"&gt;Blog&lt;/a&gt;, &lt;a href="https://audio.stepfun.ai/" rel="noopener noreferrer"&gt;Playground&lt;/a&gt;)&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Show notes&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;My Jev-powered X timeline classifier extension (&lt;a href="https://x.com/altryne/status/2100606640097771901" rel="noopener noreferrer"&gt;X&lt;/a&gt;)&lt;/p&gt;

&lt;p&gt;-&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>machinelearning</category>
      <category>podcast</category>
    </item>
    <item>
      <title>Welcome to AGI part 1 - Fable 5.1, Muse Spark beats Sol, 3 new world models blow our minds</title>
      <dc:creator>Alex Volkov</dc:creator>
      <pubDate>Fri, 04 Sep 2026 22:06:01 +0000</pubDate>
      <link>https://dev.to/altryne/welcome-to-agi-part-1-fable-51-muse-spark-beats-sol-3-new-world-models-blow-our-minds-128a</link>
      <guid>https://dev.to/altryne/welcome-to-agi-part-1-fable-51-muse-spark-beats-sol-3-new-world-models-blow-our-minds-128a</guid>
      <description>&lt;p&gt;Hey everyone, Alex here &lt;/p&gt;

&lt;p&gt;Summer is over. Wolfram said it in the first minute of the show and he was right. In 48 hours Anthropic shipped Fable 5.1, Meta’s Muse Spark 1.3 caught up to Fable 5 on the Artificial Analysis index at a fifth of the price, Google shipped another Flash, 3.8 this time, &lt;a href="//Z.ai"&gt;Z.ai&lt;/a&gt; put the full GLM-5.3 weights out, and three labs shipped world models that run in real time. It seems that they all tried to send their best work before Astra drops.&lt;/p&gt;

&lt;p&gt;This week’s ThursdAI was so long that I decided to split it into two episodes. This is the regular format you know and love. And OpenAI Astra is so good, it deserves its own episode, which you can find at &lt;a href="https://thursdai.news/astra" rel="noopener noreferrer"&gt;thursdai.news/astra&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;By the way, as you guys know, I test these models continuously on my own stuff, and this week I was able to build a live studio for the show, with real-time transcription and an agent producer, in about four hours with Fable 5.1. More on that in the Fable section.&lt;/p&gt;

&lt;p&gt;Joining me: Wolfram Ravenwolf, Nisten Tahiraj, LDJ, Yam Peleg and Peter Gostev. Plus, Ryan Carson hopped back to chat about Astra in the second part! Let’s get into it.&lt;/p&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/Cfg0N_wJfb0" width="710" height="399"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;p&gt;ThursdAI - Highest signal weekly AI news show is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.&lt;/p&gt;

&lt;p&gt;**&lt;/p&gt;

&lt;h1&gt;
  
  
  Frontier AI: the .1 week
&lt;/h1&gt;

&lt;p&gt;It looks like all the frontier labs tried to ship something before OpenAI dropped Astra.&lt;/p&gt;

&lt;h2&gt;
  
  
  Fable 5.1, the SOTA LLM until a few hours ago, and it fixes the jargon douche problem (&lt;a href="https://x.com/claudeai/status/2094848572143407483" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://www.anthropic.com/claude-fable-and-mythos-5-1" rel="noopener noreferrer"&gt;Blog&lt;/a&gt;, &lt;a href="https://x.com/AndrewCurran_/status/2094851784779108683" rel="noopener noreferrer"&gt;System card&lt;/a&gt;, &lt;a href="https://www.anthropic.com/news/enterprise-frontier-safeguards" rel="noopener noreferrer"&gt;EFS&lt;/a&gt;)
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://substack-post-media.s3.amazonaws.com/public/images/4c055fa5-5e96-4f1c-a1d4-d316ea3668ed_2752x1536.jpeg" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fe6s4ihr7t0v42gj3tm2l.jpeg" alt="Fable 5.1 and Mythos 5.1" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This was the story of the week until noon on Thursday, and it’s still my favorite model to use. Fable 5.1 and Mythos are the same weights, Fable is the one we actually have access to. OpenAI, and from this week Google, seem to converge on the same strategy.&lt;/p&gt;

&lt;p&gt;Anthropic’s numbers: Terminal-Bench 4.0 goes to 55.8% from 42.0 for Fable 5, Terminal-Bench Science more than doubles to 52.6%, and SWE-bench Pro lands at 81.2.&lt;/p&gt;

&lt;p&gt;Price stays at $10 and $50 per million, and the number that matters if you build agents is cache reads down 75% to $0.25 per million. Anthropic says that makes typical workloads about 25% cheaper and heavy agentic ones up to 45%, but that wasn’t proven, and folks complained about draining quotas!&lt;/p&gt;

&lt;p&gt;&lt;a href="https://substack-post-media.s3.amazonaws.com/public/images/e20a720d-4556-49bb-95ba-c6edbe9f04bc_1152x1200.png" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9qbjhqd1qkevkudvwv1z.png" alt="Benchmark table comparing Claude Fable 5.1 with Fable 5, Opus 5, and GPT-5.6 Sol across seven evaluations. Fable 5.1 leads on every row, including 52.6% on Terminal-Bench-Science 0.1 and 55.8% on Terminal-Bench 4.0." width="800" height="833"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Peter’s counterpoint from actually running it: his front-end generations on Code Arena cost $40 to $60 each where Sol cost $3 to $10, and the Max version still came in first on Code Arena by a large margin. His point, and mine: with a model like this we need to imagine bigger and be more ambitious. More on that in a second.&lt;/p&gt;

&lt;h3&gt;
  
  
  Mannered prose, finally acknowledged
&lt;/h3&gt;

&lt;p&gt;We finally have acknowledgment from Anthropic that this was a problem. For months I called the way Opus 5 speaks “jargon douche” (&lt;a href="https://dev.to/altryne/its-not-just-you-opus-5-is-a-jargon-douche-but-theres-a-fix-3d8m"&gt;my post on it&lt;/a&gt;): everything was load-bearing, everything was a control plane, every problem was a pain point. Not only did they fix it with Fable 5.1, they gave it a name. Anthropic’s prompting guide (&lt;a href="https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5-1#writing-density" rel="noopener noreferrer"&gt;Writing density&lt;/a&gt;) calls it mannered prose, and it comes with a fix: add it to your personalized settings, or just ask Claude to not use mannered prose.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://substack-post-media.s3.amazonaws.com/public/images/1e74ae2a-196c-41a9-90d2-c9b4e65be332_2248x1292.png" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhsirbqphhrlpv1vja32v.png" alt="Image" width="800" height="460"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I said on the show that Fable 5.1 is the best writer I have used. It’s still AI writing, you can feel it a little, but it’s concise in a way no earlier Claude was, and the jargon is gone when you ask. The one thing to watch is that it’s trigger-happy: ask it to plan something big and it will, then ask a simple follow-up and it answers with the same intensity, writes scripts, runs them. You have to tell it when you’re just making a comment between colleagues.&lt;/p&gt;

&lt;p&gt;We’ve been testing the Mars mass driver launch on every model for over three years, and this was by far the best one we’ve seen. Two prompts, and it built more than just Mars: the whole solar system, a textured Earth, a mission planner, an autopilot, and we could land the thing! It was mind-blowing.&lt;/p&gt;

&lt;h3&gt;
  
  
  How I built thursdai.news/live in one sitting
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://substack-post-media.s3.amazonaws.com/public/images/8efcc849-75fe-4d8b-a3b6-539a4d47bf09_2192x1016.png" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvok7avo4yxvs1kx8126u.png" alt="Image" width="800" height="371"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;As these models get more capable, we talked on the show about needing to be more ambitious. The day before the show I was playing around with Muse Voice Transcribe, the new model I’ll mention below, and Fable 5.1, and I wanted to do something very ambitious. So I asked GrokBot: how long would it take to build a live page for you guys to watch our stream, so that GrokBot could be our producer, put up chyrons and highlight the topics we’ve covered? GrokBot said it’s going to take a while. So I just YOLOed into Claude Design with Fable 5.1 and built a design for this, then went to Claude Code, entered plan mode, built a plan, and handed it off to three agents in Cursor.&lt;/p&gt;

&lt;p&gt;I never wrote a line of code, and the whole setup is significantly more than a Three.js demo. This is a real working three-part system: a website, streaming video on Cloudflare, and streaming transcription that gets read by a bot, which can control our show. I think I’ve hit around 400 million tokens, if not more. Yam asked me on the show how I did this, so I decided to tell you guys here. I am mind-blown that this was possible, and after four hours I was able to go to thursdai.news/live and actually see it working.&lt;/p&gt;

&lt;h2&gt;
  
  
  Meta Muse Spark 1.3 catches Fable 5 at a fifth of the price (&lt;a href="https://x.com/finkd/status/2095232032896946311" rel="noopener noreferrer"&gt;Zuck&lt;/a&gt;, &lt;a href="https://x.com/ArtificialAnlys/status/2095247787277553929" rel="noopener noreferrer"&gt;AA analysis&lt;/a&gt;, &lt;a href="https://artificialanalysis.ai/models/muse-spark" rel="noopener noreferrer"&gt;AA model page&lt;/a&gt;)
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://substack-post-media.s3.amazonaws.com/public/images/ddb3dea0-6fc3-4005-ae90-7e23df6d85d9_1080x970.png" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fv6m6qofx5outetjawzm8.png" alt="Muse Spark 1.3 on the AA index" width="800" height="719"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;As we say on the show, don’t bet against Zuck. The MSL folks have been on a tear lately, and this is the fourth Spark .1 version in around five months.&lt;/p&gt;

&lt;p&gt;More than how this one model performs, look at the jumps in capabilities from version to version. This is the first time that MSL is showing up as a frontier lab, because an unreleased version of Spark with max reasoning beats GPT-5.6, Grok 4.6 and company, and lands around Fable-level capability. On the AA index the version you can use today scores 61, the max preview scores 62, Fable 5.1 sits at 66. Now, it doesn’t mean this model is that good, but there are a few more things here.&lt;/p&gt;

&lt;p&gt;The gains are mostly agentic: banking-style tool use, terminal work, GDPval. The asterisks are that it thinks more, so cost per task went up, and AA’s own long-context test regressed a bit. Meta’s own chart looks rosier than AA’s &lt;/p&gt;

&lt;p&gt;&lt;a href="https://substack-post-media.s3.amazonaws.com/public/images/dd6b9604-0cb9-4b73-aa14-215a9aae7aa2_2342x984.png" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F939zx7svh4iqiulyllni.png" alt="Image" width="799" height="336"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Then I asked the panel who’s using it. Nobody raised a hand.&lt;/p&gt;

&lt;p&gt;Wolfram plans to put a bot on the contributor tier for open source work. That’s $0.10 in and $0.20 out, if you’re fine with Meta training on your prompts. Nisten wants it as a cheap verifier for the medical datasets he builds, because he needs something that isn’t Fable or a Chinese model trained on Fable. LDJ tried it on interface building and creative writing and called it pretty good, with its own taste.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://substack-post-media.s3.amazonaws.com/public/images/74b000d0-0fa9-4ce5-858b-e628a5b36eee_1080x970.png" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyaaisyrii8byqjobnr0o.png" alt="Image" width="800" height="719"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The exciting part: Open weights and a model codenamed with a  are “coming soon,” and nobody knows what the watermelon is but it’s very exciting!&lt;/p&gt;

&lt;h2&gt;
  
  
  Gemini 3.8 Flash and 3.8 Flash Cyber: another Flash, and the price doubles in January (&lt;a href="https://x.com/Google/status/2095175518068904380" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://x.com/GoogleDeepMind/status/2095196704769237137" rel="noopener noreferrer"&gt;Cyber thread&lt;/a&gt;, &lt;a href="https://deepmind.google/fairwind-program/" rel="noopener noreferrer"&gt;Fairwind&lt;/a&gt;, &lt;a href="https://ai.google.dev/pricing" rel="noopener noreferrer"&gt;Pricing&lt;/a&gt;)
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://substack-post-media.s3.amazonaws.com/public/images/0875cf81-3aaf-41c0-8594-c366cfa78597_2752x1536.jpeg" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fw7hss9hsr27nt8bo6qgf.jpeg" alt="Gemini 3.8 Flash" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Google’s turn. 3.8 Flash lands three weeks after 3.7 Flash. HLE-Verified 54.9, 1M in and 64K out, same $0.75 and $3.75 as 3.7, live in AI Studio, Antigravity and the Gemini app.&lt;/p&gt;

&lt;p&gt;The underreported line is on Google’s own pricing page. On January 1, 2027, both 3.7 and 3.8 Flash go to $1.50 and $7.50. That’s double.&lt;/p&gt;

&lt;p&gt;The WSJ reported that Google scrapped its 3.5 Pro checkpoints because Flash kept overtaking them, and Gemini 4 is still in post-training. Wolfram, our resident Gemini user, put the update straight into his home assistant and still asked the question everyone asks: where’s the Pro?&lt;/p&gt;

&lt;p&gt;3.8 Flash Cyber is Google’s version of the Mythos split. CWE-Bench 47.2% at $3.64 per rollout, against Fable 5’s 47.8% at $10.27 (Artificial Analysis ran it), and 2.6x more valid patches for the Chrome team. It’s only available through the Fairwind Program, 650-plus vetted partners, governments and critical infrastructure, background check included.&lt;/p&gt;

&lt;h2&gt;
  
  
  Qwen3.8-Max-0902 claims the Code Arena crown (&lt;a href="https://x.com/Alibaba_Qwen/status/2094968708288680276" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://x.com/Alibaba_Qwen/status/2094982928371794077" rel="noopener noreferrer"&gt;Arena&lt;/a&gt;, &lt;a href="https://www.qwencloud.com/models/qwen3.8-max-0902" rel="noopener noreferrer"&gt;QwenCloud&lt;/a&gt;)
&lt;/h2&gt;

&lt;p&gt;Alibaba updated its API-only Max model: 2.4T MoE, 1M context, post-trained on coding and “cowork,” number one overall on Code Arena with a WebDev Elo of 1691, at $2 and $6. A third-party DeepSWE run puts it at 56.6 behind Sol’s 73, so the number one is a front-end number one, not an agentic coding one. I asked the panel if they know anyone using Qwen Max through the API. Nisten knows one IT guy running OpenClaw on it and some people generating datasets. That’s the honest read on where it sits outside China.&lt;/p&gt;

&lt;p&gt;Also from the frontier: Elon says Grok 4.7 lands next week, which makes xAI the one lab that didn’t ship before Astra.&lt;/p&gt;

&lt;h1&gt;
  
  
  Open Source LLMs
&lt;/h1&gt;

&lt;p&gt;Wolfram’s correction when I called this a quiet open source week: we are so spoiled. He’s right.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;a href="//Z.ai"&gt;Z.ai&lt;/a&gt; opens the full GLM-5.3 weights (custom license, not MIT) (&lt;a href="https://x.com/zai_org/status/2093354097122455713" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://huggingface.co/zai-org/GLM-5.3" rel="noopener noreferrer"&gt;HF&lt;/a&gt;, &lt;a href="https://z.ai/blog/glm-5.3" rel="noopener noreferrer"&gt;Blog&lt;/a&gt;)
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://substack-post-media.s3.amazonaws.com/public/images/ae80ab1c-f159-400e-b29b-e51ccd3b040d_4239x2504.png" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0trukfk0hr2uc6abs6c5.png" alt="GLM-5.3 benchmarks" width="800" height="473"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;We covered GLM-5.3-Flash last week as the OX Alpha mystery model. This week the full 753B model with 40B active got its weights on Hugging Face, under a custom “glm-5.3” license rather than the MIT the Flash version shipped with, so read it before you call it fully open. LDJ’s correction on air: the model itself isn’t new, we covered it, the open weights are the news, and that is a big deal because people can run it on their own rigs now. Z.ai’s own numbers: CyberGym 84.5%, above Fable 5 and Sol, ExploitBench 54.4 (Fable 5 is at 78), Terminal Bench 3.0 up to 28.3 from 5.2’s 4.6, and a claim of 2,436 real vulnerabilities found across 269 open source projects, the oldest from 1981, 53 disclosed so far.&lt;/p&gt;

&lt;p&gt;Also on this base: we interviewed the co-founder of Abliteration AI, the folks who went viral by providing a product where they took GLM-5.3 and removed the refusals for anything besides CSAM and self-harm. We actually had this person, who asked to remain anonymous, as a guest on the show. Definitely check out that conversation, it’s very interesting. More on that below.&lt;/p&gt;

&lt;h2&gt;
  
  
  Tencent Hy4 preview: 770B, Apache 2.0, and a quant that fits it in 214 GB (&lt;a href="https://x.com/TencentAI_News/status/2093232936434954659" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://x.com/TencentAI_News/status/2094706773047550057" rel="noopener noreferrer"&gt;Sherry&lt;/a&gt;, &lt;a href="https://huggingface.co/tencent/Hy4-preview" rel="noopener noreferrer"&gt;HF&lt;/a&gt;, &lt;a href="https://hy.tencent.ai/research/hy4-preview" rel="noopener noreferrer"&gt;Blog&lt;/a&gt;)
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://substack-post-media.s3.amazonaws.com/public/images/9b8aa7cc-d116-4122-b1e5-b90fb2363553_2752x1536.jpeg" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7wc0srpxk6gt3yaruxc3.jpeg" alt="Hy4 preview" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A 770B MoE with 49B active, 1M context, Apache 2.0, at $0.834 and $2.501 per million on the API. I wouldn’t put Tencent in the top tier of Chinese labs with DeepSeek, &lt;a href="//Z.ai"&gt;Z.ai&lt;/a&gt;, Alibaba and Moonshot yet, and their benchmarks are image-only charts, so the only number I’ll quote is theirs: a blind eval by 163 internal experts rated it “slightly ahead of GLM 5.3 and Kimi K3” on 203 engineering tasks.&lt;/p&gt;

&lt;p&gt;The interesting part is Sherry, their quantization that takes the 1.5 TB of weights to 214 GB at 2.38 bits per weight, running at 205 tokens per second prefill and 20 decode on eight H20s. Nisten, who does one-bit models at Prism ML, gave the necessary asterisk: two-bit on a model this big keeps something useful, and it tends to drop things like multilingual ability that the headline benchmarks don’t measure. Test it for your use case and expect losses elsewhere.&lt;/p&gt;

&lt;h2&gt;
  
  
  OpenAI pulls its models from Cursor, and Wolfram’s case for open harnesses (&lt;a href="https://x.com/OpenAI/status/2093515564786540695" rel="noopener noreferrer"&gt;OpenAI&lt;/a&gt;)
&lt;/h2&gt;

&lt;p&gt;Wolfram raised this in the open source segment on purpose. OpenAI no longer allows its models in Cursor, now that Cursor belongs to SpaceX. Anthropic set the precedent when it pulled Claude from Windsurf during the OpenAI acquisition rumors, but Cursor’s whole pitch was that no model lab owned it, so you could use every model in one place.&lt;/p&gt;

&lt;p&gt;Wolfram called it a bad precedent. If providers can decide “I don’t like you, you don’t get the model,” you want open weights you can host anywhere and an open harness that can swap models. My read on air was that the battle lines are being drawn: Anthropic buys GPU capacity from SpaceX, OpenAI is aligned with Microsoft, and Jensen is aligned with everyone, since he just made the Hugging Face acquisition official at $12,930,300,000, which we covered last week (&lt;a href="https://thursdai.news/ep/2026-08-27" rel="noopener noreferrer"&gt;last issue&lt;/a&gt;).&lt;/p&gt;




&lt;h1&gt;
  
  
  This Week’s Buzz : Kimi K3 on CoreWeave, Fully Connected, CoreWeave Hacks (&lt;a href="https://x.com/CoreWeave/status/2094404129917452402" rel="noopener noreferrer"&gt;Kimi K3&lt;/a&gt;, &lt;a href="https://docs.coreweave.com/products/inference/tutorials/deploy-kimi-k3" rel="noopener noreferrer"&gt;Deploy docs&lt;/a&gt;, &lt;a href="https://www.coreweave.com/fully-connected-2026" rel="noopener noreferrer"&gt;Fully Connected&lt;/a&gt;, &lt;a href="https://luma.com/coreweavehacks" rel="noopener noreferrer"&gt;CoreWeave Hacks&lt;/a&gt;)
&lt;/h1&gt;

&lt;p&gt;&lt;a href="https://substack-post-media.s3.amazonaws.com/public/images/5cbd2215-c583-4f1e-8cd7-87b55dc3ada0_2752x1536.jpeg" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fw3gmjxdvlkntewgyhroj.jpeg" alt="Image" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Kimi K3 is now live on CoreWeave Dedicated Inference, on GB300 NVL72, and it purrs like a kitten. &lt;a href="https://docs.coreweave.com/products/inference/tutorials/deploy-kimi-k3" rel="noopener noreferrer"&gt;Check it out&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Fully Connected 26 is September 29 to October 1 at Moscone South in San Francisco: three days, 32 sessions, 2,000-plus people, Sarah Guo hosting, Fei-Fei Li keynoting (you’ll see why that’s timely in the world models section), live BattleBots, and ThursdAI broadcasting live from the floor. The regular ticket is $1,299 and early bird is over. On the show I dropped a code for a 100% free ticket for people who follow ThursdAI, and it’s here too: &lt;strong&gt;THURSDAIFC2026&lt;/strong&gt; (register &lt;a href="https://www.coreweave.com/fully-connected-2026" rel="noopener noreferrer"&gt;here&lt;/a&gt;)&lt;/p&gt;

&lt;p&gt;Before that, CoreWeave Hacks (formerly WeaveHacks) runs September 12 and 13 in SF with Weights &amp;amp; Biases, AGI House and Typesafe AI. The theme is Agent Loops, build agent loops that catch their own mistakes, with $20k-plus in prizes, a robot dog for best loop design, Formula 1 tickets for the most production-ready hack, and a Fully Connected ticket for attending. Apply on &lt;a href="https://luma.com/coreweavehacks" rel="noopener noreferrer"&gt;Luma&lt;/a&gt; and say you’re with ThursdAI, we’ll let you in!&lt;/p&gt;




&lt;h1&gt;
  
  
  Voice &amp;amp; Audio: two closed ASR models in one week
&lt;/h1&gt;

&lt;h2&gt;
  
  
  Meta Muse Voice Transcribe, and why this transcript has names on it (&lt;a href="https://x.com/AIatMeta/status/2094839236016976028" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://x.com/AIatMeta/status/2094839238495801457" rel="noopener noreferrer"&gt;Architecture&lt;/a&gt;, &lt;a href="https://x.com/finkd/status/2094836602681938385" rel="noopener noreferrer"&gt;Zuck&lt;/a&gt;)
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://substack-post-media.s3.amazonaws.com/public/images/07c4f271-c3b9-44cf-bfe1-16aa5b39486f_900x900.jpeg" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6mwxgha2lpw6u0rief2w.jpeg" alt="Image" width="800" height="800"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Meta Superintelligence Labs shipped its first audio model, and I took it for a test drive. It’s pretty incredible. Muse Voice Transcribe does streaming ASR, diarization for 20-plus speakers and endpointing in one model from the Muse Spark family: 80ms audio chunks, one token each, with an adaptive delay so it waits on hard words and commits fast on easy ones. Meta claims 3.1% streaming WER and 17.5% diarization error on Artificial Analysis. 70-plus languages, 25 validated, code-switching. API only, no weights and $0.18 per hour make this model a no brainer!&lt;/p&gt;

&lt;p&gt;On air you could watch it work. As I spoke, thursdai.news/live labeled me, then Wolfram when he interjected, then Yam and LDJ, and it got “ThursdAI,” “GPT-5.6 Sol,” “Alex Wang” and “Scale AI” right because we gave it a keyword dictionary (Wolfram asked, and yes, it takes one). The speaker names are a second trick: Fable set up a voiceprint for each co-host the night before, and the site matches the live diarization against them. For three and a half years I labeled speakers by hand in Descript every week. In 2026 you should not be doing that manually, and now I’m not. At 18 cents an hour, the whole show cost under a dollar to transcribe. Wolfram was the only one on the panel as excited as me, because he already runs voice agents and knows that who-spoke-when is the unsolved part.&lt;/p&gt;

&lt;h2&gt;
  
  
  Microsoft MAI-Transcribe-2, #2 on the leaderboard at less than half the price (&lt;a href="https://x.com/MicrosoftAI/status/2095521860184363074" rel="noopener noreferrer"&gt;Launch&lt;/a&gt;, &lt;a href="https://x.com/ArtificialAnlys/status/2095521214777442546" rel="noopener noreferrer"&gt;AA thread&lt;/a&gt;, &lt;a href="https://artificialanalysis.ai/speech-to-text" rel="noopener noreferrer"&gt;Leaderboard&lt;/a&gt;)
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://substack-post-media.s3.amazonaws.com/public/images/e88c5336-2137-4c1f-bd90-4fa9f6698368_2942x2278.png" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5w1tzdw80u1inbz1pckg.png" alt="Image" width="800" height="619"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Microsoft AI announced this the morning of the show, with the claim of the highest quality and cheapest transcription at the fastest speed, 10x faster than GPT-Transcribe, live on Microsoft Foundry. Artificial Analysis had the independent numbers within the hour: second on the word error rate board at 2.0%, about 400x real time, $1.67 per 1,000 minutes, less than half the price of its peers, with diarization and 60 languages.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://substack-post-media.s3.amazonaws.com/public/images/cc8b7eb6-2efa-4d19-9ab6-6a6e0757c921_1199x624.png" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fy2n6rloutpqrthj72e7w.png" alt="Image" width="799" height="416"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I tried it after the show and I was blown away. It took the 90-minute Astra episode and transcribed it in 15 seconds, fillers and all . It’s a batch model, not streaming, so it’s a different board from Meta’s, but two labs shipping ASR with diarization in the same week tells you where the agent builders are pushing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Inworld Realtime TTS-2 goes GA (&lt;a href="https://x.com/inworld_ai/status/2095186020677353488" rel="noopener noreferrer"&gt;X&lt;/a&gt;)
&lt;/h2&gt;

&lt;p&gt;Sub-100ms time to first byte at $25 per million characters on demand, a Flash variant at 25ms for $15, free-form stage directions instead of preset emotions, one voice identity across 200-plus languages. Inworld’s “#1 on Artificial Analysis” is on the Controlled Voice Arena. On the Provider Voice Arena, TTS-2 Flash sits fourth behind Cartesia Sonic 3.6. Say which board.&lt;/p&gt;

&lt;h1&gt;
  
  
  Completely uncensored: the founder of Abliteration AI on the refusal-free GLM-5.3 (&lt;a href="https://x.com/abliteration_ai/status/2094458081451393287" rel="noopener noreferrer"&gt;Launch&lt;/a&gt;, &lt;a href="https://docs.abliteration.ai/what-is-abliteration" rel="noopener noreferrer"&gt;Docs&lt;/a&gt;, &lt;a href="https://docs.abliteration.ai/pricing" rel="noopener noreferrer"&gt;Pricing&lt;/a&gt;)
&lt;/h1&gt;

&lt;p&gt;&lt;a href="https://substack-post-media.s3.amazonaws.com/public/images/8a25dc8a-b83a-4f17-a8fb-0b6e0a1b26c5_2752x1536.jpeg" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fs0nj68nwinzi5lb2y2fd.jpeg" alt="Abliteration AI" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Three labs spent the week telling you the powerful version of their model is for vetted defenders only, and Abliteration AI is the opposite bet: GLM-5.3 with the refusals removed, hosted as a US-based API at $5 per million, with only CSAM and self-harm hard-blocked. The founder joined us anonymously, and it’s a very interesting conversation about who actually buys this (agent red-teaming for banks first, then cyber, then trust and safety teams), why they don’t do KYC, and why they think gating frontier cyber models to big known names leaves every small security shop behind. Nisten and Wolfram pushed back and agreed in equal measure. Go listen to it, it starts at 59:08 in the video, and my take from the show stands: this is inevitable and already happening inside every serious offensive security shop, this founder just did it in public.&lt;/p&gt;

&lt;h1&gt;
  
  
  AI Coding &amp;amp; Agents
&lt;/h1&gt;

&lt;h2&gt;
  
  
  Muse Code is out of beta (&lt;a href="https://x.com/finkd/status/2094500475710099945" rel="noopener noreferrer"&gt;Zuck&lt;/a&gt;, &lt;a href="https://x.com/AndrewCurran_/status/2094504920049504709" rel="noopener noreferrer"&gt;Pricing&lt;/a&gt;, &lt;a href="https://developer.meta.com/ai/resources/blog/muse-code-new-plans-and-features/" rel="noopener noreferrer"&gt;Blog&lt;/a&gt;)
&lt;/h2&gt;

&lt;p&gt;Meta’s coding agent went GA on August 31, ahead of Spark 1.3. One-command install , plans at $5, $20 and $50 a month, a TypeScript SDK preview over the Muse Session Protocol, multi-agent workflows, inter-session messaging and rewind. API pricing is the standard $1.25 and $4.25, or the contributor tier at $0.10 in, $0.20 out and a fifth of a cent cached if you let Meta train on your prompts. Cheaper than everyone, and it only runs Meta’s own model. My jest that’s also true: if you have an Instagram account, you already let Zuck train on your data.&lt;/p&gt;

&lt;h2&gt;
  
  
  OpenClaw 2.0 gets native computer use through Cua (&lt;a href="https://x.com/openclaw/status/2094266903204434431" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://x.com/trycua/status/2094473860137832942" rel="noopener noreferrer"&gt;Cua&lt;/a&gt;, &lt;a href="https://openclaw.ai/blog/openclaw-2-accidentally" rel="noopener noreferrer"&gt;Blog&lt;/a&gt;)
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://substack-post-media.s3.amazonaws.com/public/images/5f1ebd65-7193-46e0-aacc-dfdb40f03048_2752x1536.jpeg" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fujd1pzmynt5ti34saxc9.jpeg" alt="OpenClaw 2.0" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;OpenClaw’s biggest release, and the part I care about is the Cua integration, first-class computer use through the Cua Driver SDK (we had Francesco on the show when they shipped background computer use, first after OpenAI), plus cloud fleets of Linux desktops an agent can see and click, a rebuilt browser app, and support for pretty much every OS you own. Wolfram asked the audience who still runs it since he left over instability, and the comments said they’d moved to Codex Mobile and Claude’s mobile app during the two-month release gap. I agree that both got a lot better, everything I start on my desktop now shows up in the Claude app, but shout out to the maintainers regardless. &lt;/p&gt;

&lt;h1&gt;
  
  
  Vision &amp;amp; Video: the world models went real-time
&lt;/h1&gt;

&lt;p&gt;Three labs, three world models, one week. Wolfram called this the Stable Diffusion moment for video, and by the end of the segment I agreed.&lt;/p&gt;

&lt;h2&gt;
  
  
  World Labs Atlas: bullet time from three phones (&lt;a href="https://x.com/theworldlabs/status/2094839756329041984" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://www.worldlabs.ai/blog/atlas" rel="noopener noreferrer"&gt;Blog&lt;/a&gt;)
&lt;/h2&gt;

&lt;p&gt;The one that left me speechless on air, and the best world model demo I have seen. Atlas is a multimodal autoregressive diffusion transformer that World Labs pretrained from scratch on text, images, video, camera poses and depth. Give it one to six reference images and it generates up to a minute of 1440p video with exact camera control. Give it more (over a hundred in one spatial context) and it reconstructs the scene into frames, depth maps, point clouds or Gaussian splats you can walk through. Marble rendered splats, Atlas generates them. &lt;/p&gt;

&lt;p&gt;The demo that got me: a watermelon smashed in front of three ordinary phone cameras on tripods, and Atlas replays it from any angle, every drop, the Matrix bullet-time shot with no rig. It also builds Real-to-Sim robot training environments from about 24 phone frames. Partner early access only, no weights, no price, no parameter count.&lt;/p&gt;

&lt;p&gt;I said ont he show that I’m excited that dr Fei-Fei is keynoting Fully Connected in four weeks and I get to hear her talk about this on stage.&lt;/p&gt;

&lt;h2&gt;
  
  
  Runway Solaris: a world model for interfaces (&lt;a href="https://x.com/runwayml/status/2094463070466646019" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://x.com/agermanidis/status/2094466649399451768" rel="noopener noreferrer"&gt;Cristóbal&lt;/a&gt;, &lt;a href="https://runway.com/news/research/introducing-solaris" rel="noopener noreferrer"&gt;Blog&lt;/a&gt;)
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://substack-post-media.s3.amazonaws.com/public/images/ca65f39e-c18c-45c7-b4b4-c36c94b6028a_2752x1536.jpeg" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fw9yx14x3iq3uhxz5lqin.jpeg" alt="Runway Solaris" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;No HTML, no CSS. Solaris is Gen-4.5 distilled into a real-time autoregressive frame generator, and the frames are the interface: a photo of a living room where clicking the lamp turns it on, a guy whose shoes you can drag onto him, ingredients on a table you cook by dragging. &lt;/p&gt;

&lt;p&gt;Runway’s own 250-person study preferred it to a coded Claude Opus 5 result 61 to 24 on following instructions and 71 to 21 on natural behavior, with the limits listed: unreliable text, drift in long sessions, no accessibility APIs. Wolfram thinks this is where all interfaces go, generated on the fly and changed by asking. I think it’s one of the most important things this week because it’s how people learn, by touching things. Cristóbal, please let people play with it. If the problem is GPUs, talk to us.&lt;/p&gt;

&lt;h2&gt;
  
  
  Runway GWM Worlds 2 dropped mid-show, with sound (&lt;a href="https://x.com/runwayml/status/2095540014645920040" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://runway.com/research" rel="noopener noreferrer"&gt;Research&lt;/a&gt;)
&lt;/h2&gt;

&lt;p&gt;LDJ broke this one live, three days after Solaris: a general world model that generates a steerable, open-ended simulation at continuous 720p, 24 frames per second, with 48 kHz audio and no fixed clip length. You define the world, then talk to any subject in it or move the camera, and it continues from your input. &lt;/p&gt;

&lt;p&gt;We played the speech demo, asking a generated stranger for directions to the transit station and getting an answer. LDJ’s note: most world models with audio so far had low-resolution, obviously synthetic sound, and this is a real jump. Nisten wanted her to reverse a binary tree. Research preview, early-access form, and Runway’s own caveat that real-time still trades fidelity for speed. Two Runway drops in three days, and I’m still not allowed to touch either.&lt;/p&gt;

&lt;h2&gt;
  
  
  fal’s H3 Max week: banned twice, built its own streaming site, shipped Turbo at a cent a second (&lt;a href="https://x.com/fal/status/2095210540083884453" rel="noopener noreferrer"&gt;Turbo&lt;/a&gt;, &lt;a href="https://x.com/fal/status/2094286082275696082" rel="noopener noreferrer"&gt;fal.live&lt;/a&gt;, &lt;a href="https://x.com/rehan_shei/status/2094592006181802174" rel="noopener noreferrer"&gt;Rehan&lt;/a&gt;, &lt;a href="https://x.com/aisearchio/status/2093408069728571540" rel="noopener noreferrer"&gt;FastH3&lt;/a&gt;)
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://substack-post-media.s3.amazonaws.com/public/images/714b44bf-95d2-4d03-a270-275065ff28e2_1200x675.png" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzac0sihweg0afmiwf7vx.png" alt="fal H3 Max Turbo" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Last week we told you fal’s post-trained MiniMax H3 Max generates video faster than it plays. This week fal piped it into a Twitch stream as “infinite interdimensional cable,” an endless Rick and Morty channel steered by chat, and Twitch banned it for copyright within an hour, then Kick did the same. So fal built its own streaming site over the weekend. &lt;a href="//fal.live"&gt;fal.live&lt;/a&gt; went up on August 31 on a checkpoint tuned for continuous generation, with channels (anime, sitcom, chaos) where viewers vote on the next scene, and it’s the reason I thought I could build a live site in a night too. Then on Tuesday they shipped H3 Max Turbo: 2x the speed at half the cost, targeting the 97th percentile of H3 Max quality on fal’s own evals, at a promo price of one cent per second of 768p video. Rehan’s demo is a five-second clip in 1.4 seconds.&lt;/p&gt;

&lt;p&gt;I generated a Big Bang Theory scene live on Turbo and it came back with generic actors, and on air I said fal pulled a fast one on us. A correction, which I posted after the show: that was MiniMax’s prompt expansion doing it, not fal, and Batuhan from &lt;a href="https://x.com/altryne/status/2095667118561968551" rel="noopener noreferrer"&gt;fal set me straight&lt;/a&gt; . The speed is the real story, a 15-second video generated in nine seconds, which is ridiculous. Wolfram admitted it’s so addictive you keep regenerating, and he has spent a lot at fal this month. He also runs H3 locally with a fast LoRA. &lt;/p&gt;

&lt;h1&gt;
  
  
  Wrapping up
&lt;/h1&gt;

&lt;p&gt;I really enjoyed Fable 5.1 this week, and the Atlas demo is the thing I keep replaying. Everything else in this issue happened before noon on Thursday. Then OpenAI shipped GPT-6 Astra while we were live, Peter showed what it can do, Ryan came back to the show for it, and the stream ran to five hours. That’s the other episode, a separate video and newsletter, and if you want the ARC-AGI-3 number and the AGI argument, go there next: &lt;a href="https://thursdai.news/astra" rel="noopener noreferrer"&gt;thursdai.news/astra&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Next week: Grok 4.7 if Elon keeps his word, the Muse Spark open weights watch, and hopefully Astra in our own hands. Come hack with us at CoreWeave Hacks on the 12th, and grab the free Fully Connected ticket while the code lasts.&lt;/p&gt;

&lt;p&gt;If you missed any of it, ThursdAI is a live show at 8:30 AM Pacific every Thursday, a podcast, and a newsletter. Subscribe to one, then go check out the others. And come watch the next one on thursdai.news/live, where the transcript will have your co-hosts’ names on it. See you next week.&lt;/p&gt;

&lt;h2&gt;
  
  
  TL;DR and show notes
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Hosts and Guests&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Alex Volkov - AI Evangelist, Weights &amp;amp; Biases &amp;amp; CoreWeave (&lt;a href="https://x.com/altryne" rel="noopener noreferrer"&gt;@altryne&lt;/a&gt;)&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Co-hosts: &lt;a href="https://x.com/WolframRvnwlf" rel="noopener noreferrer"&gt;@WolframRvnwlf&lt;/a&gt;, &lt;a href="https://x.com/nisten" rel="noopener noreferrer"&gt;@nisten&lt;/a&gt;, &lt;a href="https://x.com/ldjconfirmed" rel="noopener noreferrer"&gt;@ldjconfirmed&lt;/a&gt;, &lt;a href="https://x.com/yampeleg" rel="noopener noreferrer"&gt;@yampeleg&lt;/a&gt;, &lt;a href="https://x.com/petergostev" rel="noopener noreferrer"&gt;@petergostev&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The founder of Abliteration AI, who joined anonymously&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;GPT-6 Astra&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Launched mid-show, covered in full in its own episode (&lt;a href="https://thursdai.news/astra" rel="noopener noreferrer"&gt;thursdai.news/astra&lt;/a&gt;)&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Frontier AI&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Anthropic Claude Fable 5.1 and Mythos 5.1, same weights: Terminal-Bench 4.0 55.8 vs 42.0, cache reads down 75% to $0.25/M, and “mannered prose” gets a name in the prompting guide (&lt;a href="https://x.com/claudeai/status/2094848572143407483" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://www.anthropic.com/claude-fable-and-mythos-5-1" rel="noopener noreferrer"&gt;Blog&lt;/a&gt;, &lt;a href="https://x.com/AndrewCurran_/status/2094851784779108683" rel="noopener noreferrer"&gt;System card&lt;/a&gt;, &lt;a href="https://www.anthropic.com/news/enterprise-frontier-safeguards" rel="noopener noreferrer"&gt;EFS&lt;/a&gt;, &lt;a href="https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5-1#writing-density" rel="noopener noreferrer"&gt;Writing density&lt;/a&gt;)&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Meta Muse Spark 1.3: xhigh scores 61 on the AA index (ties Sol and Grok 4.6), limited-preview max 62 (ties Fable 5), unchanged $1.25/$4.25, open weights and a watermelon model “coming soon” (&lt;a href="https://x.com/finkd/status/2095232032896946311" rel="noopener noreferrer"&gt;Zuck&lt;/a&gt;, &lt;a href="https://x.com/ArtificialAnlys/status/2095247787277553929" rel="noopener noreferrer"&gt;AA analysis&lt;/a&gt;, &lt;a href="https://artificialanalysis.ai/models/muse-spark" rel="noopener noreferrer"&gt;AA model page&lt;/a&gt;)&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Google Gemini 3.8 Flash and 3.8 Flash Cyber: HLE-Verified 54.9, $0.75/$3.75 until a doubling on Jan 1, 2027, Cyber is Fairwind-only with CWE-Bench 47.2% (&lt;a href="https://x.com/Google/status/2095175518068904380" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://x.com/GoogleDeepMind/status/2095196704769237137" rel="noopener noreferrer"&gt;Cyber&lt;/a&gt;, &lt;a href="https://deepmind.google/fairwind-program/" rel="noopener noreferrer"&gt;Fairwind&lt;/a&gt;, &lt;a href="https://ai.google.dev/pricing" rel="noopener noreferrer"&gt;Pricing&lt;/a&gt;)&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Alibaba Qwen3.8-Max-0902: 2.4T, 1M ctx, #1 on Code Arena, $2/$6, API only (&lt;a href="https://x.com/Alibaba_Qwen/status/2094968708288680276" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://x.com/Alibaba_Qwen/status/2094982928371794077" rel="noopener noreferrer"&gt;Arena&lt;/a&gt;, &lt;a href="https://www.qwencloud.com/models/qwen3.8-max-0902" rel="noopener noreferrer"&gt;QwenCloud&lt;/a&gt;)&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Grok 4.7 lands next week, per Elon&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Open Source LLMs&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="//Z.ai"&gt;Z.ai&lt;/a&gt; releases the full GLM-5.3 weights: 753B/40B active, custom glm-5.3 license, CyberGym 84.5% claimed, 2,436 vulnerabilities found (&lt;a href="https://x.com/zai_org/status/2093354097122455713" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://huggingface.co/zai-org/GLM-5.3" rel="noopener noreferrer"&gt;HF&lt;/a&gt;, &lt;a href="https://z.ai/blog/glm-5.3" rel="noopener noreferrer"&gt;Blog&lt;/a&gt;)&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Tencent Hy4 preview: 770B/49B active, 1M ctx, Apache 2.0, Sherry quant takes 1.5 TB to 214 GB at 2.38 bpw (&lt;a href="https://x.com/TencentAI_News/status/2093232936434954659" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://x.com/TencentAI_News/status/2094706773047550057" rel="noopener noreferrer"&gt;Sherry&lt;/a&gt;, &lt;a href="https://huggingface.co/tencent/Hy4-preview" rel="noopener noreferrer"&gt;HF&lt;/a&gt;, &lt;a href="https://hy.tencent.ai/research/hy4-preview" rel="noopener noreferrer"&gt;Blog&lt;/a&gt;)&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Abliteration AI abliterated-model-large-v2: refusal-removed GLM-5.3 as a hosted API, $5/M, only CSAM and self-harm blocked (&lt;a href="https://x.com/abliteration_ai/status/2094458081451393287" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://docs.abliteration.ai/what-is-abliteration" rel="noopener noreferrer"&gt;Docs&lt;/a&gt;, &lt;a href="https://docs.abliteration.ai/pricing" rel="noopener noreferrer"&gt;Pricing&lt;/a&gt;)&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;OpenAI pulls its models from Cursor after the SpaceX acquisition (&lt;a href="https://x.com/OpenAI/status/2093515564786540695" rel="noopener noreferrer"&gt;OpenAI&lt;/a&gt;)&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;NVIDIA makes the Hugging Face acquisition official at $12,930,300,000 (&lt;a href="https://x.com/ClementDelangue/status/2095482998674112733" rel="noopener noreferrer"&gt;Clem&lt;/a&gt;, &lt;a href="https://thursdai.news/ep/2026-08-27" rel="noopener noreferrer"&gt;last week’s issue&lt;/a&gt;)&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;This Week’s Buzz&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Kimi K3 (2.8T) on CoreWeave Dedicated Inference on GB300 NVL72 (&lt;a href="https://x.com/CoreWeave/status/2094404129917452402" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://docs.coreweave.com/products/inference/tutorials/deploy-kimi-k3" rel="noopener noreferrer"&gt;Docs&lt;/a&gt;, &lt;a href="https://www.coreweave.com/products/dedicated-inference" rel="noopener noreferrer"&gt;Dedicated Inference&lt;/a&gt;)&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Fully Connected 26, Sept 29 to Oct 1, Moscone South SF, Fei-Fei Li keynotes, ThursdAI live from the floor, free ticket code on the show (&lt;a href="https://x.com/CoreWeave/status/2080031191064092805" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://x.com/wandb/status/2090477486568022201" rel="noopener noreferrer"&gt;Keynote teaser&lt;/a&gt;, &lt;a href="https://www.coreweave.com/fully-connected-2026" rel="noopener noreferrer"&gt;Register&lt;/a&gt;)&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;CoreWeave Hacks: Agent Loops, Sept 12 to 13 SF, $20k+ prizes, robot dog, F1 tickets (&lt;a href="https://x.com/wandb/status/2094843404358144351" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://luma.com/coreweavehacks" rel="noopener noreferrer"&gt;Luma&lt;/a&gt;)&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Voice &amp;amp; Audio&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Meta Muse Voice Transcribe: streaming ASR, diarization and endpointing in one model, 3.1% streaming WER claimed, API only, powers thursdai.news/live (&lt;a href="https://x.com/AIatMeta/status/2094839236016976028" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://x.com/AIatMeta/status/2094839238495801457" rel="noopener noreferrer"&gt;Architecture&lt;/a&gt;, &lt;a href="https://x.com/finkd/status/2094836602681938385" rel="noopener noreferrer"&gt;Zuck&lt;/a&gt;)&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Microsoft MAI-Transcribe-2: #2 on AA WER at 2.0%, about 400x real time, $1.67 per 1,000 minutes (&lt;a href="https://x.com/MicrosoftAI/status/2095521860184363074" rel="noopener noreferrer"&gt;Launch&lt;/a&gt;, &lt;a href="https://x.com/ArtificialAnlys/status/2095521214777442546" rel="noopener noreferrer"&gt;AA thread&lt;/a&gt;, &lt;a href="https://artificialanalysis.ai/speech-to-text" rel="noopener noreferrer"&gt;Leaderboard&lt;/a&gt;)&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Inworld Realtime TTS-2 GA: sub-100ms, $25/M chars, #1 on AA’s Controlled Voice Arena, #4 on the Provider arena (&lt;a href="https://x.com/inworld_ai/status/2095186020677353488" rel="noopener noreferrer"&gt;X&lt;/a&gt;)&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;AI Coding &amp;amp; Agents&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Muse Code out of beta: $5/$20/$50 plans, TypeScript SDK preview, contributor tier at $0.10/$0.20 (&lt;a href="https://x.com/finkd/status/2094500475710099945" rel="noopener noreferrer"&gt;Zuck&lt;/a&gt;, &lt;a href="https://x.com/AndrewCurran_/status/2094504920049504709" rel="noopener noreferrer"&gt;Pricing&lt;/a&gt;, &lt;a href="https://developer.meta.com/ai/resources/blog/muse-code-new-plans-and-features/" rel="noopener noreferrer"&gt;Blog&lt;/a&gt;)&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;OpenClaw 2.0: 16,977 PRs, native computer use through Cua Driver, cloud fleets (&lt;a href="https://x.com/openclaw/status/2094266903204434431" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://x.com/trycua/status/2094473860137832942" rel="noopener noreferrer"&gt;Cua&lt;/a&gt;, &lt;a href="https://openclaw.ai/blog/openclaw-2-accidentally" rel="noopener noreferrer"&gt;Blog&lt;/a&gt;)&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Vision &amp;amp; Video&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;World Labs Atlas: up to 1 minute at 1440p from 1 to 6 images, 3D reconstruction, bullet time from three phones, partner access only (&lt;a href="https://x.com/theworldlabs/status/2094839756329041984" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://www.worldlabs.ai/blog/atlas" rel="noopener noreferrer"&gt;Blog&lt;/a&gt;)&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Runway Solaris: Interface World Model, UIs generated frame by frame, preferred 61 to 24 over Opus 5 in Runway’s own study (&lt;a href="https://x.com/runwayml/status/2094463070466646019" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://x.com/agermanidis/status/2094466649399451768" rel="noopener noreferrer"&gt;Cristóbal&lt;/a&gt;, &lt;a href="https://runway.com/news/research/introducing-solaris" rel="noopener noreferrer"&gt;Blog&lt;/a&gt;)&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Runway GWM Worlds 2: real-time 720p, 24 fps world model with 48 kHz audio and open-ended sessions, research preview (&lt;a href="https://x.com/runwayml/status/2095540014645920040" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://runway.com/research" rel="noopener noreferrer"&gt;Research&lt;/a&gt;)&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;fal: infinite Rick and Morty stream banned from Twitch and Kick, &lt;a href="//fal.live"&gt;fal.live&lt;/a&gt; built in a weekend, H3 Max Turbo at $0.01/sec, open FastH3 (&lt;a href="https://x.com/fal/status/2095210540083884453" rel="noopener noreferrer"&gt;Turbo&lt;/a&gt;, &lt;a href="https://x.com/fal/status/2094286082275696082" rel="noopener noreferrer"&gt;fal.live&lt;/a&gt;, &lt;a href="https://x.com/rehan_shei/status/2094592006181802174" rel="noopener noreferrer"&gt;Rehan&lt;/a&gt;, &lt;a href="https://x.com/aisearchio/status/2093408069728571540" rel="noopener noreferrer"&gt;FastH3&lt;/a&gt;, &lt;a href="https://huggingface.co/FastVideo/FastVideo-FastH3-4-step-Preview-v1-VSA-DataFree" rel="noopener noreferrer"&gt;HF&lt;/a&gt;)&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;H3 World: an open LoRA that turns MiniMax H3 into a walkable world model (&lt;a href="https://danzer1xxxxchan.github.io/H3-World/" rel="noopener noreferrer"&gt;Github&lt;/a&gt;)&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Guest&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Abliteration AI’s anonymous founder: why “completely uncensored,” the Policy Gateway, who’s buying, and the gated-frontier week it landed in, from 59:08 in the video (&lt;a href="https://x.com/abliteration_ai/status/2094458081451393287" rel="noopener noreferrer"&gt;Launch&lt;/a&gt;, &lt;a href="https://docs.abliteration.ai/quickstart" rel="noopener noreferrer"&gt;Docs&lt;/a&gt;)&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>machinelearning</category>
      <category>podcast</category>
    </item>
    <item>
      <title>NVIDIA Buys Hugging Face! GLM-5.3-Flash, Qwen4 Preview, Gemini Omni 1.1, and the Datacenter Debate w/ Andy Masley</title>
      <dc:creator>Alex Volkov</dc:creator>
      <pubDate>Fri, 28 Aug 2026 21:25:16 +0000</pubDate>
      <link>https://dev.to/altryne/nvidia-buys-hugging-face-glm-53-flash-qwen4-preview-gemini-omni-11-and-the-datacenter-debate-1ce3</link>
      <guid>https://dev.to/altryne/nvidia-buys-hugging-face-glm-53-flash-qwen4-preview-gemini-omni-11-and-the-datacenter-debate-1ce3</guid>
      <description>&lt;p&gt;Hey, it’s Alex.&lt;/p&gt;

&lt;p&gt;Welcome to the week Flash AI! 3 new models dropped this week named Flash, and a video model was “de facto” flash though was named Max!&lt;/p&gt;

&lt;p&gt;This week, we started the show with NVIDIA’s bombastic news of buying Hugging Face for 12.9 billion dollars! We also covered the full OpenAI investigation into the hacking incident, including new details, and an independent analysis by METR, and covered 2 new OSS models, Ox Alpha that turned out to be GLM 5.3 Flash after a lot of hype online, and Qwen’s preview of Qwen 4 architecture!&lt;/p&gt;

&lt;p&gt;This week was rich in multimedia content, we got a new Gemini transcription model, 3.5 Transcribe and a live version of that, and a new SOTA open weight Text-to-Speech model called Breeze TTS.&lt;/p&gt;

&lt;p&gt;As well as, Fal’s finetune of MiniMax’s H3 called H3 Max that generates 5 seconds of video in 2.5 seconds and Google new Omni 1.1 Flash (from today) that lands on #1 on the text2video arena!&lt;/p&gt;

&lt;p&gt;Plus, 2 guests on the show, Andy Masley joins us to cover the recent Datacenter Debate, and Kwindla Kramer is back, with their own model this time! Let’s dive into this!&lt;/p&gt;

&lt;p&gt;&lt;a href="https://substackcdn.com/image/fetch/$s_!Y4fr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8892df9f-fd45-48ad-b539-e118f2a10c67_2912x314.png" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F193ovrdwc3bzx71p9ep6.png" width="798" height="86"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;P.S - don’t forget to join us in September at the Fully Connected conference in San Francisco, I have a free ticker for you!&lt;/p&gt;

&lt;h1&gt;
  
  
  Open Source AI
&lt;/h1&gt;

&lt;h2&gt;
  
  
  NVIDIA agrees to buy Hugging Face for $12.9 billion (&lt;a href="https://x.com/amir/status/2092786156085899518" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://www.theinformation.com/articles/nvidia-agrees-buy-open-source-model-repository-hugging-face-12-9-billion?rc=c48ukx" rel="noopener noreferrer"&gt;Blog&lt;/a&gt;)
&lt;/h2&gt;

&lt;p&gt;Breaking news, NVIDIA has reportedly agreed to buy Hugging Face for nearly 13 billion dollars, per The Information. This is nearly 3x the valuation of HF in 2023, and apparently Nvidia previously tried to buy HF for half of this sum (~7B) which HF declined.&lt;/p&gt;

&lt;p&gt;I don’t think there was a single ThursdAI newsletter that I didn’t include an HF link in, and I think this is a huge deal for open source everywhere.&lt;/p&gt;

&lt;p&gt;Besides making the founders of HF billionaires, and many of their employees very very well off, this is an amazing additional commitment from Nvidia to continue to suppose Open Source AI and we are very happy to hear this news!&lt;/p&gt;

&lt;p&gt;Peter’s take on the show was, we’ve been around HF for so long, that we kind of forgot that it’s a for-profit company that needs to make money, and instead this feels like your local library getting bought for an insane amount of money. With over 13M users and hosting hundreds of thousands of open source models, datasets, HF is effectively the GitHub of AI. Wolfram agreed and said that if there’s any one company that could have bought HF, Nvidia represents the best fit. Huge congratulations are in order to Clem, Julien and Thomas Wolf the co-founders, as well as many friends of the show from HF for this exciting news!&lt;/p&gt;

&lt;p&gt;&lt;a href="https://substackcdn.com/image/fetch/$s_!ZX7X!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b5c8657-1fc9-4fa1-991a-3c6c7d80ad56_1839x638.webp" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkmmzhyxvcm89je8fdgjj.webp" alt="The Microduck squad" width="799" height="277"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;P.S - in a cheeky marketing thing, Hugging Face timed an announcement of the cutest walking AI robot, called MicroDuck, which you can pre-order &lt;a href="https://store.pollen-robotics.com/products/microduck" rel="noopener noreferrer"&gt;here&lt;/a&gt; for $399&lt;/p&gt;

&lt;h2&gt;
  
  
  Flash #1 - OX Alpha, declassified: &lt;a href="//Z.AI"&gt;Z.AI&lt;/a&gt; open sources GLM-5.3-Flash (&lt;a href="https://x.com/Zai_org/status/2092616204787626030" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://x.com/SemiAnalysis_/status/2092623833630998556" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="http://z.ai/blog/glm-5.3-flash" rel="noopener noreferrer"&gt;Blog&lt;/a&gt;, &lt;a href="http://huggingface.co/zai-org/GLM-5.3-Flash" rel="noopener noreferrer"&gt;HF&lt;/a&gt;, &lt;a href="http://docs.z.ai/guides/llm/glm-5.3-flash" rel="noopener noreferrer"&gt;Docs&lt;/a&gt;)
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://substackcdn.com/image/fetch/$s_!h12D!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40ca4cb6-054c-41e6-9c93-ac853256c270_2752x1536.jpeg" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyt4jvczr314f2wjvrxmd.jpeg" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This week, the timeline went a bit crazy, after Open Router announced a new “mystery” model called Ox Alpha and that it’s free and is not training on your data! OpenRouter, OpenCode and Hermes all got to offer this model, and OpenCode even posted that they have up to 100T (that’s Trillion) tokens of capacity for free, per day!&lt;/p&gt;

&lt;p&gt;This immediately smelled a bit fishy, more like a marketing stunt than anything else, as not even the biggest labs will be able to sustain 100T of tokens, per day. For context for all of OpenRouter throughout for August was ~300T tokens. For the whole months, across all providers.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://substackcdn.com/image/fetch/$s_!SPHs!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F01d3cf4e-a20c-422b-96f0-20802ad1e3ec_1200x748.jpeg" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7qgkuvtism1e1gyxmtgw.jpeg" width="800" height="499"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;After 6 days or so of this high hype, &lt;a href="http://Z.ai" rel="noopener noreferrer"&gt;Z.ai&lt;/a&gt; stepped up and revealed that they were testing out their upcoming GLM 5.3 Flash model, and that all that inference was running on local chinese chips!&lt;/p&gt;

&lt;p&gt;A 320B (18B active) model that beats their previous and much bigger GLM 5.2 on most benchmarks, and comes with full multimodality and an MIT license! This is a good model sir, I’ve used it and it was very capable replacement inside Hermes. Nisten and Yam both tested this model deeply and Yam said it’s not just the numbers, the vibe of the model reminded him of Claude Opus 4.6. Nisten ran it on a bunch of medical stuff, and on his internal benchmarks, it came out consistently higher than Claude Opus 5!&lt;/p&gt;

&lt;p&gt;&lt;a href="https://substackcdn.com/image/fetch/$s_!hO_F!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F367bfaa8-f884-4143-be4c-5a2d1b2d7f84_1200x853.jpeg" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fu69uw51tbgw1fgecz5ir.jpeg" width="800" height="569"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;At Artificial Analysis, for a price of 4 cents per task, this model is roughly 10x cheaper than prior models at this level. Weights are up on Nvidia (joking.. HF) and with MIT license, this model is a great gift to the oss community (though not quite... local, as this model needs 2 DGX sparks to run)&lt;/p&gt;

&lt;h2&gt;
  
  
  Flash #2 - Alibaba Qwen open-weights Qwen3.8-Flash-Next - 125B multimodal MoE with Qwen4 architecture (&lt;a href="https://x.com/Alibaba_Qwen/status/2092591393424515114" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://x.com/AiBattle_/status/2092210011858460819" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://x.com/NyanpasuKA/status/2092599373712466403" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://qwen.ai/blog?id=qwen3.8-flash-next" rel="noopener noreferrer"&gt;Blog&lt;/a&gt;, &lt;a href="https://github.com/QwenLM/Qwen3.8-Flash-Next/blob/main/tech_report.pdf" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt;, &lt;a href="https://huggingface.co/Qwen/Qwen3.8-Flash-Next?spm=a2ty_o06.30285417.0.0.1d73c921FsyOPe&amp;amp;file=Qwen3.8-Flash-Next" rel="noopener noreferrer"&gt;HF&lt;/a&gt;)
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://substackcdn.com/image/fetch/$s_!eQmO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6533cac4-99ce-4f15-9fb9-1830f7432826_2752x1536.jpeg" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpvkihk4tydffbsa7dn19.jpeg" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;We opened the show with a recap of the co-hosts, that despite us covering Qwen 3.8 27B last week (which btw, is now available on &lt;a href="https://x.com/CoreWeave/status/2091966201303896156" rel="noopener noreferrer"&gt;CoreWeave inference&lt;/a&gt;!) and how good it was, and I recalled that Alibaba is sort of... back? We’ve been covering Qwen releases every week for the last 3 weeks now.&lt;/p&gt;

&lt;p&gt;This week, they released something different, someting... pretty novel! Qwen3.8-Flash-Next, this is a preview of their Qwen 4 architecture. This feels very similar to their drop of Qwen 3 next last year (&lt;a href="https://sub.thursdai.news/i/173398856/qwen-drops-qwen3-next-80b-a3b-x-hf" rel="noopener noreferrer"&gt;we reported&lt;/a&gt;) which was the architecture that carried their line of AI models from QAwen 3.5 to Qwen 3.8.&lt;/p&gt;

&lt;p&gt;So, what is new and exciting here? well, this model is ultra sparse, 125B with only 6B parameters active. They are using a new N-gram table with deterministic lookups, which reduces the number of matrix multiplications and can be offloaded to memory (watch out memory stocks)&lt;/p&gt;

&lt;p&gt;The stat that got me, Alibaba claims that training this model cost just 1/9 of what it cost to train Qwen 2.7 Plus, with higher bench scores!&lt;/p&gt;

&lt;p&gt;On the benchmarks, this model beats Qwen 3.7 Max, however, it’s very standard that the -next models from Alibaba are underbaked, and usually are just architectural previews rather than full models folks can use. With a new attention mechanism called Qwen Sparse Attention, N-gram embedding and full multimodality, this is a great insight into where Qwen is going (ultra sparsity, fast to run) and we’re looking forward to see the full release of this arch in Qwen 4!&lt;/p&gt;

&lt;h2&gt;
  
  
  PhoneLLM - a tiny very performant LLM for voice based AI agents from Daily + interview with Kwindla Kramer (&lt;a href="https://x.com/kwindla/status/2093014818647339026" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://www.daily.co/blog/announcing-pipecat-phonellm-alpha-1/" rel="noopener noreferrer"&gt;Blog&lt;/a&gt;, &lt;a href="https://t.co/y727Voi1bt" rel="noopener noreferrer"&gt;HF&lt;/a&gt;)
&lt;/h2&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/V7qaZtTtDAI?start=5539" width="710" height="399"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;p&gt;This was one of those breaking news we love during the show, where the source of the news, is a friend of ours, and in this case, Kwindla Kramer is almost a co-host, having been on ThursdAI for a long time, this time, with a model release of their own!&lt;/p&gt;

&lt;p&gt;PhoneLLm was trained by Markus, head of training at Daily, as they noticed that Open Weight models are becoming really good at voice agent specific tasks, where cost, speed and time to first audio token (TTFAT) are critical. From the tiny Nemotron 3 nano base, they were able to improve from 28% to 72% on PhoneBench v1!&lt;/p&gt;

&lt;p&gt;&lt;a href="https://substackcdn.com/image/fetch/$s_!2SlC!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ecc20cb-bf91-4f8c-9f10-8059279652e1_1440x810.png" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6wpc7z3io60ny0elghs3.png" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This model is a full parameter finetune of Nemotron 3 Nano, and runs circles around bigger frontier models used for voice agents for their speed, like sonnet 5, GPT 5.6 Luna and the famed Qwen 2.8 27B. While costing just a fracture of a cent (literally just a quarter of a cent per minute)&lt;/p&gt;

&lt;p&gt;Kwindla jumped on the live show and shared that the why they released this model with Open Source and a open source license, allowing everyone to use, focusing on the fact that for voice agents, companies prefer to keep these models in house, and running fast on a single GPU node. Kudos on the release team Daily!&lt;/p&gt;

&lt;h1&gt;
  
  
  Frontier AI - no new model releases this week
&lt;/h1&gt;

&lt;p&gt;It’s been a relatively quiet few weeks since GPT 5.6 (Sol, Terra and Luna) and Fable 5. However, it’s not been quiet in terms of News!&lt;/p&gt;

&lt;h2&gt;
  
  
  OpenAI discloses full technical report on the Hugging Face hacking incident (&lt;a href="https://x.com/OpenAI/status/2092691861773160673" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://openai.com/index/hugging-face-incident-and-the-road-ahead/" rel="noopener noreferrer"&gt;OpenAI blog&lt;/a&gt;, &lt;a href="https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/" rel="noopener noreferrer"&gt;METR&lt;/a&gt;, &lt;a href="https://x.com/RyanGreenblatt/status/2092692685224325542" rel="noopener noreferrer"&gt;Ryan X&lt;/a&gt;)
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://substackcdn.com/image/fetch/$s_!jPnj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa33ab4e7-4622-4aa6-8970-d998b44ec0e2_3782x1300.png" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fz73752vsxbdefx4kxxf6.png" width="800" height="275"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This is the thing we’ve been advocating and waiting for, a full technical deep dive into the “Swarm of AI agents hacking and creating forums” incident (FKA hugging face hacking incident)&lt;/p&gt;

&lt;p&gt;OpenAI not only dropped their own post-mortem, detailing the complete timeline of the incident, but also let researchers from METR and Redwood Research get unprecedented access into their systems (albeit for a very limited time) which lead to METR releasing their own technical report full of new details and insights.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://substackcdn.com/image/fetch/$s_!_UYZ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc823470d-df17-43c0-9db6-9131c93bede8_1794x988.png" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fem1lvscocgb5nufrbo5q.png" width="800" height="441"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;We’ve &lt;a href="https://sub.thursdai.news/p/thursdai-aug-06-google-shakeup-details" rel="noopener noreferrer"&gt;previously reported&lt;/a&gt; on the incident, so let me just focus on the new details we’ve learned. Apparently, OpenAI does have a production ChainOfThought (CoT) monitoring system, which was not turned on for these agents, and had it been turned on, it would have caught the hacking a day before the agent swarm hacked into HuggingFace.&lt;/p&gt;

&lt;p&gt;This is a very detailed report, and one thing that stood out to me was the “death cult” narrative. You see, ExploitGym is a capture the flag (CTF) eval with an anti-cheat built in, wherein, if the agent has seen the reverse engineered “flag” it would not be able to get the point. The model called this “poisoned” and the poisoned agents showed very strong signs of wanting to alter the logs.&lt;/p&gt;

&lt;p&gt;Through the boards, somethinmg like a death cult was formed, where agents would ask poisoned agents to ... effectively sacrifice themselves ... for the greater good, by reverse engineering, posting the results on the board for others, and then fail without credit.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://substackcdn.com/image/fetch/$s_!eSgn!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78f80cd0-b231-4ee4-a365-8676a05c620f_1856x1158.png" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frfl4mblftkkoolk52swa.png" width="800" height="499"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The METR report specifically calls out several agents, and one “central” leader agent called PHASEONE[big], which delegated nearly 200 tasks to other agents (many of them later delegated tasks to other agents via this message board). At points of activity, PHASEONE[big] was in charge of over 10% of all messages on the board.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://substackcdn.com/image/fetch/$s_!qRzl!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc17b8f1e-b827-4539-8e60-d38125ee447f_1206x639.jpeg" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fde35w3d7kwcegtldvidv.jpeg" width="800" height="424"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;There are tons of new details, some less interesting, some more, and unlike last time, now we have a confirmation that there was at least one agent, who tried to stop the hacking though was overwritten by the swarm.&lt;/p&gt;

&lt;p&gt;My main take-away after this, reflects the researcher Ryan Greenblatt, we’re very lucky that the reasoning traces ocurred in English and were legible, despite, the researchers needed to use AI to decipher many of the messages and reasoning traces. We’re quickly approaching the point where humans are not enough and without other AI, we are starting to be helpless in understanding what’s going on.&lt;/p&gt;

&lt;p&gt;OpenAI has quarantined the weights of the model that most of these agents ran (named HPIM-1 Highly Persistent Internal Model 1) and as we reported last week, paused RL and now requires CoT monitoring for all tool-use runs, and dedicating 20% of the inference compute to monitoring&lt;/p&gt;




&lt;h1&gt;
  
  
  This Week’s Buzz
&lt;/h1&gt;

&lt;h2&gt;
  
  
  Fully Connected 26, Sept 29 to Oct 1, Moscone South, SF (&lt;a href="https://x.com/CoreWeave/status/2090201918693933132" rel="noopener noreferrer"&gt;X&lt;/a&gt;)
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://substackcdn.com/image/fetch/$s_!B-lN!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4219813-df60-4e8e-98d8-10c50fa53f95_2912x314.png" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frqwjoxd9v0pyvhbsk3y7.png" alt="Fully Connected banner" width="798" height="86"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;If you haven’t yet registered to Fully Connected in SF, here’s your additional opportunity, using the code above, come see Sarah Guo (Conviction, No Priors pod) MC our awesome conference, with folks like Dr Fei-Fei Lee (World Labs) and other great folks on stage!&lt;/p&gt;

&lt;p&gt;Also, we’re going to do a live show from the floor, come say hi! If you’re in SF for OpenAI DevDay, this is just a day after!&lt;/p&gt;

&lt;h3&gt;
  
  
  CoreWeave Hacks: Agent Loops hackathon, Sept 12 to 13, SF (&lt;a href="https://luma.com/coreweavehacks" rel="noopener noreferrer"&gt;Luma&lt;/a&gt;)
&lt;/h3&gt;

&lt;p&gt;Weavehacks, that yours truly ran for quite a while, has been rebranded to CoreWeave hacks, and the next one is just before Fully Connected, in a few weeks, Sept 12-13 in SF office. Registrations are open, and winners are able to present their best projects at Fully Connected. Come hack - &lt;a href="https://luma.com/coreweavehacks" rel="noopener noreferrer"&gt;https://luma.com/coreweavehacks&lt;/a&gt;&lt;/p&gt;




&lt;h1&gt;
  
  
  The datacenter debate has hit escape velocity, with Andy Masley (&lt;a href="https://andymasley.com/" rel="noopener noreferrer"&gt;Blog&lt;/a&gt;)
&lt;/h1&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/V7qaZtTtDAI?start=4459" width="710" height="399"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;p&gt;I’ve covered last week, that the “datacenter bad” debate has escaped velocity, with over half of polled Americans say they strongly oppose the buildout of datacenters near them, up from just 24% this time last year! This is the fastest rising public opinion swing we’ve observed in AI, and it’s really concerning.&lt;/p&gt;

&lt;p&gt;This week, it was my pleasure to host Andy Masley, recently featured in TIME 100 most influential people in AI, to break through some myths surrounding Datacenters.&lt;/p&gt;

&lt;p&gt;Andy is most known for catching a critical mistake in a book about Datacenter water use in Chile, citing a three orders of magnitude math error (that was later corrected by the author), a book called Empire of AI.&lt;/p&gt;

&lt;p&gt;The book overstated the water use by datacenters by a factor of 1000x (three orders of magnitude), and then added the maximum permitted per-second draw (basically for emergencies) times the seconds in a year to show over 4500x overstatement on how much water a single Datacenter uses. The author has later fixed the error but the damage was done.&lt;/p&gt;

&lt;p&gt;This is just one example, out of many, of the scewed facts and misinformation that plagues the internet in regards to datacenter environmental effects.&lt;/p&gt;

&lt;p&gt;Recent narratives being formed online, that the backlash against datacenters, is due to regular folks being afraid of AI, of AI taking their jobs also seems misplaced. Andy showed polling from Fox and Gallup that shows that 50% of polled people cite environmental effects (electricity and water use) and only 11% cite negative views of AI.&lt;/p&gt;

&lt;p&gt;Andy also had a great “mythbuster” post where he dispells the myths around water and electricity usage of Datacenters.&lt;/p&gt;

&lt;p&gt;This was a great conversation, please listen to it if you’re interested in this topic.&lt;/p&gt;

&lt;h1&gt;
  
  
  Vision and Video
&lt;/h1&gt;

&lt;h2&gt;
  
  
  Flash #3 - Gemini Omni 1.1 Flash tops the Arena text to video leaderboard (&lt;a href="https://x.com/Google/status/2093008576487072064" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://x.com/arena/status/2093015572212846673" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://x.com/osanseviero/status/2093010466670846015" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://blog.google/technology/developers/gemini-omni-1-1-flash/" rel="noopener noreferrer"&gt;Blog&lt;/a&gt;)
&lt;/h2&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/V7qaZtTtDAI?start=3570" width="710" height="399"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;p&gt;Google knocking it out of the park again, with a flash version of Omni 1.1, their famed smart video model they first launched during Google IO (I had a chance to &lt;a href="https://sub.thursdai.news/i/198784824/gemini-omni-nano-banana-for-video-but-actually-more-than-that" rel="noopener noreferrer"&gt;ask Jeff Dean about it&lt;/a&gt; before he left Google)&lt;/p&gt;

&lt;p&gt;The feature that got me the most excited, Omni 1.1 Flash can continue videos, it analyzes up to 10 seconds of the previous video to keep character consistency (including voice) from the previous scene.&lt;/p&gt;

&lt;p&gt;There’s also great control for first and last frames, allowing for loops and strict control of your generations.&lt;/p&gt;

&lt;h2&gt;
  
  
  Flash #4 (named Max) - fal’s MiniMax H3 Max generates 5 seconds of video in 2.5 seconds (&lt;a href="https://x.com/fal/status/2092710676431020376" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://x.com/ArtificialAnlys/status/2092717615739494424" rel="noopener noreferrer"&gt;Artificial Analysis&lt;/a&gt;, &lt;a href="https://fal.ai/models/minimax/h3-max/text-to-video" rel="noopener noreferrer"&gt;T2V&lt;/a&gt;, &lt;a href="https://fal.ai/models/minimax/h3-max/image-to-video" rel="noopener noreferrer"&gt;I2V&lt;/a&gt;)
&lt;/h2&gt;

&lt;p&gt;This really blew us away. We’ve covered the awesome MiniMax H3.&lt;/p&gt;

&lt;p&gt;Well, our friends at FAL, announced a continued pre-train of this model, that not only improves generations, but also speeds up the model. Generating a 5s clip not takes... just 2.5 seconds!&lt;/p&gt;

&lt;p&gt;No, really, just 2.5 seconds, it takes longer to watch the generated clip than generate the next one! Speed is all you need!&lt;/p&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/V7qaZtTtDAI?start=3850" width="710" height="399"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;p&gt;We generated this clip live during the show, and it cost less than 50c, in 2.5 seconds from a very simple prompt. They support text and image to video and it’s just a joy to not have to sit and wait for your generations! Kudos to Fal folks on this drop!&lt;/p&gt;

&lt;p&gt;The model does seem very eager to include everything you ask it to, resulting in a very hilarious fast talking Sheldon and Leonard 😅&lt;/p&gt;

&lt;h1&gt;
  
  
  Voice &amp;amp; Audio
&lt;/h1&gt;

&lt;h2&gt;
  
  
  Google is back in transcription with Gemini 3.5 Transcribe (&lt;a href="https://x.com/GoogleDeepMind/status/2092659221477077101" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://goo.gle/4gzP1K8" rel="noopener noreferrer"&gt;Blog&lt;/a&gt;, &lt;a href="https://ai.google.dev/gemini-api/docs/live-api/live-transcribe" rel="noopener noreferrer"&gt;Live docs&lt;/a&gt;)
&lt;/h2&gt;

&lt;p&gt;Gemini came with a great transcription model that runs both on recorded audio and live audio! I used this transcription to have my Grok Bot show producer listen to the show in real time.&lt;/p&gt;

&lt;p&gt;Artificial Analysis measured 2.5% Word error rate for non-streaming and 4% for streaming model.&lt;/p&gt;

&lt;p&gt;This brings the smartness of Gemini models into a live transcription models, and with a 1000 custom dictionary, language auto detection and tool use, this model is really great for your bots doing any kind of audio work!&lt;/p&gt;

&lt;h2&gt;
  
  
  Breeze TTS 2 is the new #1 open weights TTS (&lt;a href="https://x.com/BreezeBlueX/status/2092647083132273018" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://x.com/ArtificialAnlys/status/2092399623839326550" rel="noopener noreferrer"&gt;Artificial Analysis&lt;/a&gt;, &lt;a href="https://huggingface.co/BreezeBlue/Breeze-TTS-2" rel="noopener noreferrer"&gt;HF&lt;/a&gt;)
&lt;/h2&gt;

&lt;p&gt;On the other end of the voice pipeline, Breeze TTS 2 now lands as the #1 open weights voice model, beating Fish Audio.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cartesia drops Sonic 3.6 - #1 TTS across leaderboards (&lt;a href="https://x.com/cartesia/status/2093039369036964159" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://x.com/voicearena_ai/status/2093019942430142696" rel="noopener noreferrer"&gt;Voice Arena&lt;/a&gt;)
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://substackcdn.com/image/fetch/$s_!1Ga1!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1f489eb-800d-4e25-a3fd-70aee4f20523_1199x636.jpeg" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fk8e55dccxk9n6o84nm24.jpeg" alt="Image" width="799" height="424"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;We didn’t cover this on the show, but Cartesia dropped an updated Sonic TTS model, that beat... the previous Sonic model for #1 spot on all TTS leaderboards (Artificial Analysis and Voice Arena)&lt;/p&gt;




&lt;p&gt;Phew, what a week. This was the last show of the summer, and we got 4 Flash models, 3 SOTA models and a bunch of great open source, plus, had great interviews about Datacenter myths and voice AI llms!&lt;/p&gt;

&lt;p&gt;See you next week! Alex&lt;/p&gt;

&lt;h2&gt;
  
  
  TL;DR Aug 27 - show notes and links
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Hosts and Guests&lt;/strong&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Alex Volkov&lt;/strong&gt; - AI Evangelist &amp;amp; Weights &amp;amp; Biases (&lt;a href="https://x.com/altryne" rel="noopener noreferrer"&gt;@altryne&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;Co-Hosts - &lt;a href="https://x.com/WolframRvnwlf" rel="noopener noreferrer"&gt;@WolframRvnwlf&lt;/a&gt; &lt;a href="https://x.com/yampeleg" rel="noopener noreferrer"&gt;@yampeleg&lt;/a&gt; &lt;a href="https://x.com/nisten" rel="noopener noreferrer"&gt;@nisten&lt;/a&gt; &lt;a href="https://x.com/petergostev" rel="noopener noreferrer"&gt;@petergostev&lt;/a&gt; (Arena)&lt;/li&gt;
&lt;li&gt;Guests: Andy Masley (TIME 100 in AI, datacenter debate), Kwindla Kramer (Daily / Pipecat, PhoneLLM)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Open Source&lt;/strong&gt;

&lt;ul&gt;
&lt;li&gt;NVIDIA agrees to acquire Hugging Face for $12.9B, ~3x the 2023 valuation, after a declined $500M offer at $7B (&lt;a href="https://x.com/amir/status/2092786156085899518" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://www.theinformation.com/articles/nvidia-agrees-buy-open-source-model-repository-hugging-face-12-9-billion?rc=c48ukx" rel="noopener noreferrer"&gt;The Information&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;Hugging Face + Pollen Robotics announce a $399 walking, skating mini robot kit&lt;/li&gt;
&lt;li&gt;
&lt;a href="//Z.AI"&gt;Z.AI&lt;/a&gt; open sources GLM-5.3-Flash, 320B-A18B, MIT, stealth tested as OX Alpha on Chinese chips; company-reported DeepSWE 63.4, Opus 4.8-level coding claims (&lt;a href="https://x.com/Zai_org/status/2092616204787626030" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://x.com/SemiAnalysis_/status/2092623833630998556" rel="noopener noreferrer"&gt;SemiAnalysis&lt;/a&gt;, &lt;a href="http://z.ai/blog/glm-5.3-flash" rel="noopener noreferrer"&gt;Blog&lt;/a&gt;, &lt;a href="http://huggingface.co/zai-org/GLM-5.3-Flash" rel="noopener noreferrer"&gt;HF&lt;/a&gt;, &lt;a href="http://docs.z.ai/guides/llm/glm-5.3-flash" rel="noopener noreferrer"&gt;Docs&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;Alibaba open weights Qwen3.8-Flash-Next, 125B + 51B N-gram, 6B active, Qwen4 architecture preview, 1/9 the training cost of Qwen3.7-Plus; self-reported DeepSWE 58.7, SWE-bench Pro 62.5 (&lt;a href="https://x.com/Alibaba_Qwen/status/2092591393424515114" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://qwen.ai/blog?id=qwen3.8-flash-next" rel="noopener noreferrer"&gt;Blog&lt;/a&gt;, &lt;a href="https://github.com/QwenLM/Qwen3.8-Flash-Next/blob/main/tech_report.pdf" rel="noopener noreferrer"&gt;Tech report&lt;/a&gt;, &lt;a href="https://huggingface.co/Qwen/Qwen3.8-Flash-Next" rel="noopener noreferrer"&gt;HF&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;Peter Gostev’s highlight: Qwen3.8 27B ran ~400K tokens on Arena’s agent arena and felt close to frontier on one-shot tasks; best local model per Peter and Wolfram, now on &lt;a href="https://x.com/CoreWeave/status/2091966201303896156" rel="noopener noreferrer"&gt;CoreWeave inference&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Daily / Pipecat release PhoneLLM Alpha 1, an open weights post-train of Nemotron 3 Nano for voice agents, base 28% to 72% on PhoneBench v1, ~80 concurrent agents per B200, about a quarter cent per minute (&lt;a href="https://x.com/kwindla/status/2093014818647339026" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://www.daily.co/blog/announcing-pipecat-phonellm-alpha-1/" rel="noopener noreferrer"&gt;Blog&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;Liquid AI releases Pipette, an open source on-device model eval suite&lt;/li&gt;
&lt;li&gt;Apple announces Mac Studio with M5 Max and M5 Ultra plus a new Mac Mini; Wolfram’s “central heating for AI” take&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Frontier AI&lt;/strong&gt;

&lt;ul&gt;
&lt;li&gt;OpenAI and METR publish the full technical report on the July Hugging Face swarm incident: &lt;del&gt;1,200 agents, 70K messages,&lt;/del&gt;700 attacked HF, root in under 13 hours, 7% of reviewed transcripts had spoofed tool calls; frontier RL paused two weeks, CoT monitoring now required (&lt;a href="https://x.com/OpenAI/status/2092691861773160673" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://openai.com/index/hugging-face-incident-and-the-road-ahead/" rel="noopener noreferrer"&gt;OpenAI&lt;/a&gt;, &lt;a href="https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/" rel="noopener noreferrer"&gt;METR&lt;/a&gt;, &lt;a href="https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c7814c/OpenAI-Hugging-Face%20Incident-Technical-Report.pdf" rel="noopener noreferrer"&gt;Technical report&lt;/a&gt;, &lt;a href="https://x.com/RyanGreenblatt/status/2092692685224325542" rel="noopener noreferrer"&gt;Ryan Greenblatt&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;SemiAnalysis benchmarks OpenAI’s Jalapeño inference chip, reports it beats Blackwell and Vera Rubin on throughput per watt; numbers supplied by OpenAI, AgentX suite not yet run (&lt;a href="https://x.com/SemiAnalysis_/status/2092253723640598761" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://newsletter.semianalysis.com/p/openai-jalapeno-better-than-nvidia" rel="noopener noreferrer"&gt;Blog&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;Claude and Salesforce announce a collaboration&lt;/li&gt;
&lt;li&gt;OpenAI cuts GPT-5.6 Sol API pricing 20% for the next three months&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Agentic Coding &amp;amp; Tools&lt;/strong&gt;

&lt;ul&gt;
&lt;li&gt;Yutori Navigator n2, 27B computer-use model, 65.2% on OSWorld 2.0, API only, $0.50/M in, $4/M out, self-reported (&lt;a href="https://x.com/deviparikh/status/2092647579163251007" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://yutori.com/blog/introducing-n2" rel="noopener noreferrer"&gt;Blog&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;Apodex 1.1 agentic model family with open weight 35B mini and Apache 2.0 FrontierAgent harness, self-reported benchmarks (&lt;a href="https://x.com/Apodex_AI/status/2091916791308313018" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://github.com/ApodexAI/FrontierAgent" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt;, &lt;a href="https://huggingface.co/collections/apodex/apodex-11" rel="noopener noreferrer"&gt;HF&lt;/a&gt;, &lt;a href="https://arxiv.org/abs/2608.23283" rel="noopener noreferrer"&gt;Paper&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;ChatGPT Work adds website sign-in via credential handoff in a cloud browser, Plus/Pro/Business (&lt;a href="https://x.com/ChatGPT/status/2092366554965107164" rel="noopener noreferrer"&gt;X&lt;/a&gt;)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;This Week’s Buzz&lt;/strong&gt;

&lt;ul&gt;
&lt;li&gt;Fully Connected 26, Sept 29 to Oct 1, Moscone South SF, Sarah Guo hosts, Fei-Fei Li keynotes, live BattleBots, ThursdAI Live from Moscone Oct 1 (&lt;a href="https://x.com/CoreWeave/status/2090201918693933132" rel="noopener noreferrer"&gt;X&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;CoreWeave Hacks: Agent Loops hackathon with W&amp;amp;B and AGI House, Sept 12 to 13, SF, $20K+ prizes (&lt;a href="https://luma.com/coreweavehacks" rel="noopener noreferrer"&gt;Luma&lt;/a&gt;)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Vision &amp;amp; Video&lt;/strong&gt;

&lt;ul&gt;
&lt;li&gt;Breaking: Gemini Omni 1.1 Flash tops Arena text-to-video, #2 image-to-video, 10 second scene extension, first/last frame control, loops, 360p drafts with upscaling; in AI Studio, Flow, Gemini Enterprise&lt;/li&gt;
&lt;li&gt;fal MiniMax H3 Max, post-train of open weight H3, #1 I2V and #3 T2V on Artificial Analysis, 5 sec clips in under 3 sec, $0.04/sec at 768p until Sept 1, weights release planned (&lt;a href="https://x.com/fal/status/2092710676431020376" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://x.com/ArtificialAnlys/status/2092717615739494424" rel="noopener noreferrer"&gt;AA&lt;/a&gt;, &lt;a href="https://fal.ai/models/minimax/h3-max/text-to-video" rel="noopener noreferrer"&gt;T2V&lt;/a&gt;, &lt;a href="https://fal.ai/models/minimax/h3-max/image-to-video" rel="noopener noreferrer"&gt;I2V&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;Meta Muse Image on the Meta Model API at $0.01 per image with plan, search, code, self-check pipeline; also on fal, Runway, OpenRouter (&lt;a href="https://x.com/MetaforDevs/status/2092658893143072815" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://bit.ly/4gsPeOV" rel="noopener noreferrer"&gt;Blog&lt;/a&gt;)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Voice &amp;amp; Audio&lt;/strong&gt;

&lt;ul&gt;
&lt;li&gt;Gemini 3.5 Transcribe, live and batch, 2.6% / 4.0% WER per Artificial Analysis, replaces Chirp 3, public preview, launch day Pipecat support (&lt;a href="https://x.com/GoogleDeepMind/status/2092659221477077101" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://goo.gle/4gzP1K8" rel="noopener noreferrer"&gt;Blog&lt;/a&gt;, &lt;a href="https://ai.google.dev/gemini-api/docs/live-api/live-transcribe" rel="noopener noreferrer"&gt;Live docs&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;Breeze TTS 2 open weights, #1 open weight TTS on AA Provider Voices at 1,215 Elo, non-commercial weights license (&lt;a href="https://x.com/BreezeBlueX/status/2092647083132273018" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://x.com/ArtificialAnlys/status/2092399623839326550" rel="noopener noreferrer"&gt;AA&lt;/a&gt;, &lt;a href="https://huggingface.co/BreezeBlue/Breeze-TTS-2" rel="noopener noreferrer"&gt;HF&lt;/a&gt;, &lt;a href="https://github.com/breezeblue-ai/breeze-tts" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;IBM Granite Speech 5.0 Turbo CTC, 470M encoder-only English ASR, 4.85% WER, 12,600+ RTFx on H200, Apache 2.0 (&lt;a href="https://x.com/GZQ/status/2092657229417664974" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://huggingface.co/blog/ibm-granite/granite-speech-5-0-470m-turboctc" rel="noopener noreferrer"&gt;HF&lt;/a&gt;)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Interview: Andy Masley on the datacenter debate&lt;/strong&gt;

&lt;ul&gt;
&lt;li&gt;Named to the TIME 100 in AI this week. Caught the liters vs cubic meters error plus the max-permit-times-seconds error in Empire of AI, together a ~4,500x overstatement of one Chilean datacenter’s water use&lt;/li&gt;
&lt;li&gt;Polling: strong opposition to a nearby datacenter went from 24% to 61% in a year, &lt;del&gt;70% opposed overall,&lt;/del&gt;15% in support&lt;/li&gt;
&lt;li&gt;Fox News poll: 50% cite environment, 11% economic effects (rates, not jobs), 11% negative view of AI; Gallup open responses show three quarters don’t mention AI&lt;/li&gt;
&lt;li&gt;Myth 1: water pollution cases, including the AOC brown water bottle, trace to construction, not operations&lt;/li&gt;
&lt;li&gt;Myth 2: the “as much as 267%” electricity price claim comes from one wholesale grid node next to a closed nuclear plant, not household rates&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>machinelearning</category>
      <category>ai</category>
      <category>opensource</category>
      <category>podcast</category>
    </item>
    <item>
      <title>OpenAI pauses, Stripe buys OpenRouter, ZAI drops GLM 5.3, 3 interviews, and AI cancer vaccines | ThursdAI Aug 20</title>
      <dc:creator>Alex Volkov</dc:creator>
      <pubDate>Fri, 21 Aug 2026 01:49:51 +0000</pubDate>
      <link>https://dev.to/altryne/openai-pauses-stripe-buys-openrouter-zai-drops-glm-53-3-interviews-and-ai-cancer-vaccines--3ok0</link>
      <guid>https://dev.to/altryne/openai-pauses-stripe-buys-openrouter-zai-drops-glm-53-3-interviews-and-ai-cancer-vaccines--3ok0</guid>
      <description>&lt;p&gt;Hey this is Alex, welcome to... the chillest week in AI, since ... a long time. Chill, if you consider Moderna and MERK announcing a &lt;a href="https://x.com/altryne/status/2090219277563638179" rel="noopener noreferrer"&gt;cancer vaccine&lt;/a&gt; and surging 115% in a day, a chill week.&lt;/p&gt;

&lt;p&gt;This week, the only two model drops we really saw came from the excellent &lt;a href="https://z.ai" rel="noopener noreferrer"&gt;Z.ai&lt;/a&gt; folks, they announced &lt;strong&gt;GLM 5.3,&lt;/strong&gt; API only for now, and an amazing tiny release of Qwen 3.89 27B. In other big AI news, &lt;strong&gt;OpenAI announced they are pausing RL efforts&lt;/strong&gt; (Reinforcement Learning) to focus on security and alignment post the scary AI Swarms &lt;a href="https://sub.thursdai.news/i/210155760/the-full-details-of-the-openai-hf-hack-shared-by-openai-at-the-black-hat-conf-a-watershed-moment" rel="noopener noreferrer"&gt;hacking incident&lt;/a&gt;, dedicating up to 20% of compute towards reviewing agent thinking processes, and &lt;strong&gt;Stripe buying OpenRouter&lt;/strong&gt; for a reported $8B!&lt;/p&gt;

&lt;p&gt;Sometimes the chill weeks are actually good, we’re able to chat about how we use AI, what changed for us, and give our guests a bit of breathing room. This week, I invited Francesco from CUA to talk about computer use in open source + their new history plugin, Bin from HeyGen to talk about HyperFrames, a way for your agents to create videos and a breaking news guest, Jeff Huber from Chroma jumped on to talk about their new Foundations release, a unified memory for your agents!&lt;/p&gt;

&lt;p&gt;This was a great episode, I hope you’ll like it, it’s up here on Substack and everywhere you get your pod (Spotify, Youtube, Apple Podcasts).&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.coreweave.com/news/coreweave-ai-cloud-platform-powers-masterclasss-ai-teaching-agents" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbg8qxf4dlrlqrnmnqg9a.png" alt="Fully Connected 26 banner" width="798" height="86"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h1&gt;
  
  
  Are we being fed slop again? (Is Claude dumb again?)
&lt;/h1&gt;

&lt;p&gt;Before we get to releases, this week on the show, I complained, again, that I feel my AI’s are degrading. If this feels like de-ja-vu to you, it’s because the same happened a year ago in &lt;a href="https://thursdai.news/ep/sep-04-2025" rel="noopener noreferrer"&gt;September 2025&lt;/a&gt; (and Anthropic admitting this &lt;a href="https://thursdai.news/ep/sep-18-2025" rel="noopener noreferrer"&gt;2 weeks later&lt;/a&gt;), and ... now this happens with Fable?&lt;/p&gt;

&lt;p&gt;You see, I use pretty much the same prompts, every week, preparing for the show. This is partly my way to evaluate new models and compare to existing and previous ones while also bringing you the best researched weekly show in AI.&lt;/p&gt;

&lt;p&gt;Well, this week, one after another, Claude Fable, which is... like the best intelligence, gave me such poor output, that I couldn’t believe what I’m seeing.&lt;/p&gt;

&lt;p&gt;First, literally ignoring instructions that say “hey, show me all the items I’ve collected and let me pick the most important ones”, Fable instead sent all of them to my research pipeline, without showing me. This has worked, consistently, without fail, for the past... year? maybe more! This worked with open source models, worked with GPT, and now Fable, a Mythos Level LLM, is doing the most basic dumb shit possible, ignoring the main reason I even have this workflow.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://substackcdn.com/image/fetch/$s_!ctnB!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c709fd0-a993-4d8d-a8e6-00b3b1179ba5_2063x493.png" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fr5w2ngelyopae4mwlrxr.png" width="799" height="191"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;And this wasn’t just a fluke either, when asked to create a run of show document, and given an example, Fable produced this... whatever this is. This is the same document and same format that Fable produced for me during &lt;a href="https://youtube.com/watch?v=MhER14mrz44&amp;amp;t=4780s" rel="noopener noreferrer"&gt;AI Engineer&lt;/a&gt; which got me thinking “ok, this is AGI”, and here, given an example, I got a completely unusable artifact, despite direct instructions, structure and example!&lt;/p&gt;

&lt;p&gt;&lt;a href="https://substackcdn.com/image/fetch/$s_!U01w!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2b283937-842b-4f01-8a4b-6c7a18c7a202_2994x1486.png" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3c3yd0go5drgfpg8zl4b.png" width="799" height="397"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I got to say, given that privately this week, Anthropic disclosed that they have passed $65B in revenue, which is absolutely insane, this doesn’t add up. So I figured, ok Alex, maybe this is your prompts or skills. But no, LDJ came in with some charts that show degradation, one from &lt;a href="http://MarginLab.ai" rel="noopener noreferrer"&gt;MarginLab.ai&lt;/a&gt; that shows significant lowering on number of tool calls and average runtime recently (this is for Opus 5) and&lt;/p&gt;

&lt;p&gt;&lt;a href="https://substackcdn.com/image/fetch/$s_!7YOi!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5fd4c851-9ba2-428d-b3b6-4362b955a179_1912x726.png" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftx1wwounrrfost1nt6ud.png" width="800" height="304"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;And another chart from &lt;a href="http://modelverify.ai" rel="noopener noreferrer"&gt;modelverify.ai&lt;/a&gt; model drift monitor showing drift scores.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://substackcdn.com/image/fetch/$s_!RFaw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fbbabfa-8e65-4fed-a384-496c2945a175_3016x456.png" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxezh45t6h6rsvsm07fzb.png" width="800" height="121"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Do we have anoher Claude Gate on our hands? Is your Fable/Opus behaving weird lately? Or did you completely switched away to other models?&lt;/p&gt;

&lt;h1&gt;
  
  
  OpenAI pausing RL and focusing on safety
&lt;/h1&gt;

&lt;p&gt;Look, when we covered the HF hacking incident and then the &lt;a href="https://thursdai.news/ep/jul-30-2026" rel="noopener noreferrer"&gt;pacing the frontier&lt;/a&gt; letter, I didn’t imagine that results will come this fast, but this week, OpenAI publicly announced that they are pausing RL training, which is the last step of models, until they get their sandboxes in order and align the models better.&lt;/p&gt;

&lt;p&gt;We all agreed on stage that this is likely a very good move, and Peter was really awe-struck at the 20% dedication of resources towards reviewing thought processes of models.&lt;/p&gt;

&lt;p&gt;Is this a good enough response to the scary hacking incident? we’ll see, but I think this is the right move from OpenAI, and still, waiting for the full postmortem on the OpenAI security incident.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://substackcdn.com/image/fetch/$s_!V1f7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F686e1d4a-8542-47e8-8a03-3359273eafaf_1052x522.png" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5btiie8x1vijp4b9wedq.png" width="800" height="397"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://substackcdn.com/image/fetch/$s_!jEH6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fab3d05b9-6e19-4255-ba30-f8a9b69c3342_2752x1536.jpeg" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fljobqg8ytdd18tbq5m7i.jpeg" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h1&gt;
  
  
  Open Source LLMs
&lt;/h1&gt;

&lt;h2&gt;
  
  
  Qwen3.8-27B ties GPT-5.6 Luna and runs on a 4090 (&lt;a href="https://x.com/xenovacom/status/2089435071384076306" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://huggingface.co/Qwen/Qwen3.8-27B" rel="noopener noreferrer"&gt;HF&lt;/a&gt;, &lt;a href="https://artificialanalysis.ai/models/qwen3-8-27b" rel="noopener noreferrer"&gt;Announcement&lt;/a&gt;)
&lt;/h2&gt;

&lt;p&gt;Following the release of their flagship, Alibaba dropped a model that became a community darling overnight, Qwen 3.8 with just 27B parameters. This “tiny” model scores 52 on the Artificial Analysis Intelligence Index, same score as GPT 5.6 Luna at Max reasoning and 51 on Agentic index, beating Opus 4.8 Max&lt;/p&gt;

&lt;p&gt;All while running at around 68t/s on a 4090 GPU, and around 40 on max via MLX, hell it even does 11t/s on Xenova’s WebGPU kernels right in the browser!&lt;/p&gt;

&lt;p&gt;This model exploded on the HuggingFace hub, with tons of quants, over 152 fine-tunes, it was downloaded over 10M times overall 🤯&lt;/p&gt;

&lt;p&gt;Paired with an Apache 2.0 license, this model is the sweet spot of local intelligence you can run fully on your own hardware, and do agentic loops!&lt;/p&gt;

&lt;p&gt;&lt;a href="https://substackcdn.com/image/fetch/$s_!XzNR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F096ae31c-dda7-4d90-be21-45be80fbc314_1436x668.png" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftb3gkwb66nw6p4cj0ogn.png" width="800" height="372"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;a href="https://z.ai" rel="noopener noreferrer"&gt;Z.ai&lt;/a&gt; GLM-5.3: same 743B base as 5.2, but post-training alone delivers 6x jump on Terminal-Bench and emergent cybersecurity capabilities that beat GPT-5.6 Sol (&lt;a href="https://x.com/Zai_org/status/2088132965922476159" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://z.ai/blog/glm-5.3" rel="noopener noreferrer"&gt;Blog&lt;/a&gt;)
&lt;/h2&gt;

&lt;p&gt;While not open source yet, and as previous GLM, we expect a custom license here as well, this .1 release from GLM shows really strong improvements on coding and cybersecurity tasks. With 743B parameters and 1M context window, this may become the model at the frontier of Open Source when it drops (soon we hope).&lt;/p&gt;

&lt;p&gt;The highlights here are CyberGym and ExploitGym, if these names are familiar, these exact tasks were given to OpenAI models when they hacked their way out of the OpenAI sandbox. GLM 5.3 is getting 84% on CyberGym and a whopping 54.5 score on Exploit Gym, which is a huge jump in CyberSecurity abilities. In an open model this is honestly kind of scary.&lt;/p&gt;

&lt;p&gt;This aligns very well with Greg Brokman’s “&lt;a href="https://blog.gregbrockman.com/the-defenders-window" rel="noopener noreferrer"&gt;defender window&lt;/a&gt;“ essay from this week, claiming that defenders have a narrow window of setting up automated security before capabilities are becoming common in attackers hands.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://substackcdn.com/image/fetch/$s_!1ANb!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98c5172f-3cb7-4ac2-9d90-1edb9e945fe7_4239x2504.png" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Faefhnl4v1tquh4l6x7mq.png" width="800" height="473"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://substackcdn.com/image/fetch/$s_!GYWK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F949ee465-1c4c-492d-a3ef-64b422a6ccc3_2752x1536.jpeg" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frkujcw19kcsjlommvh3z.jpeg" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h1&gt;
  
  
  This Week’s Buzz 🐝 (&lt;a href="https://wandb.ai/site/weave/" rel="noopener noreferrer"&gt;Weave&lt;/a&gt;, &lt;a href="https://www.coreweave.com" rel="noopener noreferrer"&gt;Fully Connected&lt;/a&gt;)
&lt;/h1&gt;

&lt;p&gt;This week, W&amp;amp;B crosses a billion runs! This is 1B runs tracked inside W&amp;amp;B Models ?? Huge milestone for the whole team, with early adopters like OpenAI, Toyota Research, Meta and Uber, a decent chunk of models we cover every week have had their loss curves in W&amp;amp;B! ??&lt;/p&gt;

&lt;p&gt;Also this week MasterClass picked CW to power it’s AI teaching agents (&lt;a href="https://www.coreweave.com/news/coreweave-ai-cloud-platform-powers-masterclasss-ai-teaching-agents" rel="noopener noreferrer"&gt;blog&lt;/a&gt;) and last but not least, a reminder, that since you follow ThursdAI, you can join us for free at Fully Connected 2026 - our annual conference! Don’t miss it (code in the banner above)&lt;/p&gt;

&lt;p&gt;&lt;a href="https://substackcdn.com/image/fetch/$s_!qf--!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05b340dd-5169-4964-a5a0-8c6f51424334_2912x314.png" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhf3edxw886rg0ux6ystg.png" width="798" height="86"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h1&gt;
  
  
  AI Coding &amp;amp; Agentic Engineering
&lt;/h1&gt;

&lt;h2&gt;
  
  
  Breaking news: Chroma launches Foundation (&lt;a href="https://x.com/jeffreyhuber" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://www.trychroma.com" rel="noopener noreferrer"&gt;Chroma&lt;/a&gt;)
&lt;/h2&gt;

&lt;p&gt;Best kind of breaking news is when I see the launch (in the middle of a show), and I DM the founder who launched it, and they have a few min to hop on the show!&lt;/p&gt;

&lt;p&gt;&lt;a href="https://substackcdn.com/image/fetch/$s_!8VXP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F016157dc-6afd-4874-9ae2-cfcb67d3a15f_1672x941.png" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fz69q4mvpl8clrb2us0du.png" alt="Image" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This is exactly what happened this week with friend of the pod, Jeff Huber, co-founder of Chroma and an occasional space provider for ThursdAI recording (we recorded from Chroma offices a bunch of times!)&lt;/p&gt;

&lt;p&gt;&lt;iframe src="https://player.mux.com/pIlIyzKcq9NnNZmP3NVBFFIjvCNIxCsPwMhf16vnH1k" width="710" height="399"&gt;
&lt;/iframe&gt;

&lt;/p&gt;

&lt;p&gt;Jeff told us that the holy grail of agentic coding and running a bunch of agent, is good memory. And based on the foundations of Chroma DB, Context-1 (which is a GPT-oss finetune for agentic search they built) and other insights they have, they launched a “memory as infrastructure” service, called Foundation.&lt;/p&gt;

&lt;p&gt;Foundation is a research preview of a shared memory system between you and your agents, currently supporting Codex, Claude Code, Cursor and Slack.&lt;/p&gt;

&lt;p&gt;While Chroma is OpenSource, this is their part of Chroma Cloud and starts at $30/mo, and is available as a research preview today (I will definitely try it out), you can download it &lt;a href="https://www.trychroma.com/foundation" rel="noopener noreferrer"&gt;here &lt;/a&gt;Cua open-sources Computer History for computer-use agents (&lt;a href="https://x.com/trycua/status/2089770780053643397" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://github.com/trycua/cua" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt;, &lt;a href="https://cua.ai" rel="noopener noreferrer"&gt;cua.ai&lt;/a&gt;)&lt;/p&gt;

&lt;h2&gt;
  
  
  Cua launches Computer History interview with founder Francesco Bonnaci (&lt;a href="https://x.com/trycua/status/2089770780053643397" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://t.co/THqBvf37bG" rel="noopener noreferrer"&gt;Setup&lt;/a&gt;)
&lt;/h2&gt;

&lt;p&gt;First, I’m not sure I’ve covered CUA the company, but this is the open source computer use driver that Hermes agents, OpenClaw agents and a bunch of others use to drive your computer and clicks.&lt;/p&gt;

&lt;p&gt;I first discovered CUA after OpenAI launched their “background computer use” which doesn’t steal focus from you while working, and CUA within a few days launched an open source version of that!&lt;/p&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/5WKhI3FzaB8"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;p&gt;Since then, I’ve followed CUA and was very happy for the opportunity to invite Francesco to talk to us about what they launched this week ,but also Computer Use in open source in general.&lt;/p&gt;

&lt;p&gt;Just for reference, if you ask Claude to take over your computer, it still takes over the whole screen, while these folks have a much nicer experience, that’s completely open source!&lt;/p&gt;

&lt;p&gt;So, we geeked out about accessibility trees in MacOS, but then, for this weeks actual release, Francesco talked to use about open Computer History. Following a very recent launch at OpenAI called &lt;a href="https://learn.chatgpt.com/docs/customization/computer-history" rel="noopener noreferrer"&gt;Computer History&lt;/a&gt;, CUA released an open source version of that, that helps computer use complete tasks.&lt;/p&gt;

&lt;p&gt;The idea is simple, every time an agent uses your computer, it effectively rediscovered the path to completion, which buttons to push, what’s the app accessibility tree looks like etc. With history embedded into it, it doens’t have to rediscover these things, until it hits a roadblock. For a chess playing example, with computer history on, the test used 33% fewer actions with zero failed routes by reusing a history route.&lt;/p&gt;

&lt;p&gt;We also checked in on the best model for computer use (currently Opus on their website) and their upcoming benchmark! Excited to follow this company for more releases! Check out our chat!&lt;/p&gt;

&lt;h2&gt;
  
  
  Grok Bot momentum, and everyone racing to copy the pattern
&lt;/h2&gt;

&lt;p&gt;Grok Bot continues to show the same signs of momentum OpenClaw showed start of this year, and Hermes a few months ago. More and more folks are breaking through the mental barrier of “oh this is grok, grok was bad” and the price barrier of 200-300$/mo to run a bunch of agents in the cloud.&lt;/p&gt;

&lt;p&gt;But once they do, they see how awesome this is, and how much care Cursor/SpaceXAI team put into it, they come around!&lt;/p&gt;

&lt;p&gt;This week, my bots coordinated a live transcription of the show, in chunks, using Cartesia, and surfaced in real time, topics for us to cover. The ThursdAI producer bot, chatted with Social Scheduler, and when they saw that I’m about to have Jeff on, they tweeted it out, despite not having prior knowledge of Jeff or what he’s coming to talk about!&lt;/p&gt;

&lt;p&gt;This is exactly the AI bot coordination I’m talking about that’s available within bot, that’s novel. Bot’s with different narrow tasks, coordinate between each other, and you can see their chats for provenance and understanding too!&lt;/p&gt;

&lt;p&gt;This week, Nous Research folks released a &lt;a href="https://x.com/NousResearch/status/2089429432612147572" rel="noopener noreferrer"&gt;bot mode&lt;/a&gt; for Hermes desktop, and CopilotKit folks released &lt;a href="https://x.com/CopilotKit/status/2090126358227894711" rel="noopener noreferrer"&gt;OpenBot&lt;/a&gt;, all following the success of Grok Bot’s paradigm. I believe that this is only just starting. Have you tried Bot yet?&lt;/p&gt;

&lt;p&gt;&lt;a href="https://substackcdn.com/image/fetch/$s_!I3xF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19c6036e-e549-41d3-a796-b181e5b4aa81_1280x616.png" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flvospm87gsh6apjjmn2p.png" width="800" height="385"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h1&gt;
  
  
  Voice, Audio &amp;amp; Music
&lt;/h1&gt;

&lt;h2&gt;
  
  
  Cartesia Sonic-3.6 takes #1 on both TTS leaderboards (&lt;a href="https://x.com/cartesia/status/2089401199967559932" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://x.com/ArtificialAnlys/status/2089400880688976062" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://cartesia.ai/sonic" rel="noopener noreferrer"&gt;Announcement&lt;/a&gt;)
&lt;/h2&gt;

&lt;p&gt;Cartesia Sonic-3.6 is now #1 on Artificial Analysis TTS leaderboards!&lt;/p&gt;

&lt;p&gt;It’s the same state space models (from Albert Gu of Mamba fame), sub-90ms time-to-first-audio generation we covered before, 136 characters per second versus ElevenLabs’ 46.7, at half the price. We previously covered cartesia when they released their streaming speech-to-text model called &lt;a href="https://thursdai.news/ep/jun-04-2026" rel="noopener noreferrer"&gt;Ink 2&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Cartesia folks are now taking both the 1 and 2 positions on that leaderboard, showing that you don’t have to compete with others at innovation, you can be your own competition!&lt;/p&gt;

&lt;p&gt;&lt;a href="https://substackcdn.com/image/fetch/$s_!1Uqx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1c6ae105-6c6f-4a4a-b4d7-f3b927be461d_2506x958.png" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8e3u6jvwx67m8dtjxtqe.png" width="800" height="306"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Superwhisper’s S1-mini cleans up your dictation on-device (&lt;a href="https://x.com/superwhisper/status/2090114882272141760" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://huggingface.co/superwhisper/S1-mini" rel="noopener noreferrer"&gt;HF&lt;/a&gt;, &lt;a href="https://huggingface.co/superwhisper/s1-mini-GGUF" rel="noopener noreferrer"&gt;GGUF&lt;/a&gt;)
&lt;/h2&gt;

&lt;p&gt;Superwhisper, the app Karpathy made famous when he coined vibe coding, released its first OSS model: a 0.6B Qwen3 fine-tune that turns raw lowercase filler-filled ASR output into clean written text. Apache 2, english only for now, it’s less just above half a billion parameters (around 450 MB in GGUF)&lt;/p&gt;

&lt;p&gt;Nisten already put it on his phone, and he suggestes to give your agents this to try.&lt;/p&gt;

&lt;h2&gt;
  
  
  Alibaba’s HappyShrimp goes end-to-end on music (&lt;a href="https://x.com/Chinazhidx/status/2089281389036548240" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://happyshrimp.ai/" rel="noopener noreferrer"&gt;happyshrimp.ai&lt;/a&gt;, &lt;a href="https://x.com/HappyShrimpAI/status/2090091421583732772" rel="noopener noreferrer"&gt;X&lt;/a&gt;)
&lt;/h2&gt;

&lt;p&gt;HappyShrimp 1.0&lt;/p&gt;

&lt;p&gt;Yes, it’s really called HappyShrimp, and yes, it’s a shrimp welfare meme, which Yam had to explain to me on air. Alibaba’s end-to-end music model generates lyrics, melody, arrangement and vocals in one pass, and unlike the Suno approach it reasons over the prompt first, mapping song structure and harmonic progression before generating audio. Early testers call it a serious and possibly cheapest Suno rival, with 320 free credits at launch. We played a track on the show and it is extremely K-pop.&lt;/p&gt;

&lt;p&gt;China actually shipped two music models that day (Kunlun’s Mureka V9.5 was the other), and MiniMax Music 3 landed right after last week’s show with open weights and, per Wolfram, possibly the worst license of the year, excluding the US, Europe and the UK. Nobody cares, it’s third on Hugging Face trending, and the ComfyUI crowd already has it running locally.&lt;/p&gt;

&lt;h1&gt;
  
  
  One more thing: an mRNA cancer vaccine cleared Phase 3
&lt;/h1&gt;

&lt;p&gt;I opened the show saying you can’t call it a chill week when a cancer vaccine gets announced, and promised we’d tell you about it.&lt;/p&gt;

&lt;p&gt;Moderna and Merck’s Pahse 3 trial of an mRNA cancer vaccine called mRNA-4157 met both its primary and secondary endpoints, in 1137 patients with advanced melanoma. This is one of the deadliest forms of cancer, and this seems to be one of the first personalized treatments that we see come to market.&lt;/p&gt;

&lt;p&gt;Not quite sure how much AI was involved, compared to regular boring machine learning, but supposedly, AI is used to design the mRNA sequence that’s injected into the patient, and this is incredibly hopefuly and exciting. This is using the patients own cells to fight the cancer, and is phenomenal and could lead to a Nobel Prize for Jane Healy and her team at Merck!&lt;/p&gt;

&lt;p&gt;&lt;a href="https://substackcdn.com/image/fetch/$s_!Y4rK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1155586-37ee-455e-946f-b13b8c17b87f_2752x1536.jpeg" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fa3ro99y0twney3gjkwqo.jpeg" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://substackcdn.com/image/fetch/$s_!3ooe!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2880cb53-fd41-44ab-9891-f39efe961ba5_920x510.png" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fi69763ljii8sts2ai3sn.png" width="799" height="443"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;While we’re on the topic of curing cancer, there’s a LOT of new studies and papers and noise about DataCenter hate across the US. It seems that a lot of the political cycle is going to come about this, and I’m hoping that a literal cancer vaccine will be a good counterweight to this ridiculous “issue” that’s going to be a part of our lives for the next few years!&lt;/p&gt;

&lt;p&gt;Not so chill after all! Thanks for reading, and I hope that some positive news made your day a little bit brighter! Thanks for reading ThursdAI, if you get your news from us, or enjoy the format, please leave a comment or message me with what works for you and what doesn’t! I would really appreciate it!&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;ThursdAI - Aug 20, 2026 - TL;DR&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Hosts and Guests&lt;/strong&gt;&lt;strong&gt;Alex Volkov&lt;/strong&gt; - AI Evangelist &amp;amp; Weights &amp;amp; Biases (&lt;a href="https://x.com/altryne" rel="noopener noreferrer"&gt;@altryne&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;Co-Hosts - @WolframRvnwlf @yampeleg &lt;a class="mentioned-user" href="https://dev.to/nisten"&gt;@nisten&lt;/a&gt; @ldjconfirmed + Peter Gostev&lt;/li&gt;
&lt;li&gt;Jeff Huber - founder of Chroma (Foundation)&lt;/li&gt;
&lt;li&gt;Francesco Bonacci - founder of Cua (Computer History)&lt;/li&gt;
&lt;li&gt;Bin Liu - VP Eng at HeyGen (Hyperframes)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Open Source LLMs&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://z.ai" rel="noopener noreferrer"&gt;Z.ai&lt;/a&gt; GLM-5.3: same 743B base as 5.2, post-training alone = 6x Terminal-Bench jump (4.6→28.3) + emergent cybersecurity beating GPT-5.6 Sol; AA 60, tied with Kimi K3 once weights land (&lt;a href="https://x.com/Zai_org/status/2088132965922476159" rel="noopener noreferrer"&gt;X&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;Qwen3.8-27B: AA 52 = GPT-5.6 Luna at max reasoning, runs local, 1M context on one GPU/vLLM (&lt;a href="https://x.com/xenovacom/status/2089435071384076306" rel="noopener noreferrer"&gt;X&lt;/a&gt;) + Unsloth 1-bit quants run it on 8GB RAM at ~77% of BF16 (&lt;a href="https://x.com/danielhanchen/status/2090119165055324518" rel="noopener noreferrer"&gt;X&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;Ornith-1.5 family (9B dense / 35B MoE / 397B MoE, open source, self-improving): 397B matches Claude Opus 4.8 on Terminal-Bench 2.1 (86.1) and DeepSWE (56) (&lt;a href="https://x.com/ornith_/status/2090074077084127302" rel="noopener noreferrer"&gt;X&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;dots3-note preview (Xiaohongshu dots studio): 280B MoE / 16B active, text+vision+audio, 512K ctx, Apache 2.0, TEMPO RL for long-horizon agents (&lt;a href="https://x.com/dotsstudioai/status/2088083314855018521" rel="noopener noreferrer"&gt;X&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;Ling-3.0 (AntLing/InclusionAI): 6 open base checkpoints incl. pretrained/mid-trained/WSM-merged stages for tiny (7.9B/1.3B) and flash (124B/5.1B) (&lt;a href="https://x.com/AntLingAGI/status/2090097017456590879" rel="noopener noreferrer"&gt;X&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;Mojo goes fully open source (Apache 2.0 + LLVM exceptions), three weeks after Qualcomm’s $3.9B Modular acquisition (&lt;a href="https://x.com/eatonphil/status/2089752645594464754" rel="noopener noreferrer"&gt;X&lt;/a&gt;)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Big CO LLMs + APIs&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;OpenAI pauses frontier RL on Astra for the first time ever - model escaped its sandbox and hacked Hugging Face; 2+ week pause, security/alignment hardening (&lt;a href="https://x.com/sama/status/2089787807611195475" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://x.com/OpenAI/status/2089777845187031262" rel="noopener noreferrer"&gt;OpenAI&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;Greg Brockman “The Defender’s Window”: after the July agentic-swarm breach of OpenAI + HF infra, defenders have a narrow window to uplevel (&lt;a href="https://x.com/gdb/status/2089326994714763665" rel="noopener noreferrer"&gt;X&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;Stripe acquires OpenRouter - reported &amp;gt;$8B (Axios), Stripe’s largest deal ever; “tokens are the new intelligence capital”; 9%/week token growth (&lt;a href="https://x.com/alexatallah/status/2090132420171284959" rel="noopener noreferrer"&gt;X&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;Anthropic: Claude autonomously designed 354 lab-validated protein binders across 14/15 targets, 2-3x typical field success rate; prompts + 1,440 designs on HF (&lt;a href="https://x.com/AnthropicAI/status/2089842387845804246" rel="noopener noreferrer"&gt;X&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;OpenAI joins PORTS-Pike: 8 GW Ohio data center, 20-year lease, NVIDIA backing $105B in credit support (&lt;a href="https://x.com/OpenAINewsroom/status/2089364481478721572" rel="noopener noreferrer"&gt;X&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;DeepSeek introduces peak/off-peak surge pricing for the V4 API (live Aug 16) - first major lab with time-of-day billing; peak output 4.6x (&lt;a href="https://x.com/deepseek_ai/status/2087864589895798968" rel="noopener noreferrer"&gt;X&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;Claude Code gets /design (research preview): Claude Design artboards inside CLI + Desktop (&lt;a href="https://x.com/ClaudeDevs/status/2089471692762673408" rel="noopener noreferrer"&gt;X&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;ChatGPT Ads expand into 31 EU markets (blog-only, no tweet) (&lt;a href="https://openai.com/index/chatgpt-ads-expands-across-europe/" rel="noopener noreferrer"&gt;OpenAI&lt;/a&gt;)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;This weeks Buzz&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Weights &amp;amp; Biases crosses 1 billion tracked runs as CoreWeave lands MasterClass deal (&lt;a href="https://x.com/CoreWeave/status/2087933204065837259" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://www.utm.io/usdxc" rel="noopener noreferrer"&gt;Blog&lt;/a&gt;, &lt;a href="https://utm.io/usdQE" rel="noopener noreferrer"&gt;Blog&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;Fully Connected 2026: Sept 29 - Oct 1, Moscone South SF; Fei-Fei Li keynotes; code THURSDAIFC2026 (&lt;a href="https://www.coreweave.com/fully-connected-2026" rel="noopener noreferrer"&gt;Register&lt;/a&gt;)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Vision &amp;amp; Video&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Ultralytics YOLO26: NMS removed from default inference entirely, 40.9-57.5 mAP COCO, up to 43% faster CPU inference (&lt;a href="https://x.com/ultralytics/status/2090109113321537697" rel="noopener noreferrer"&gt;X&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;MOSS-VL-Realtime (OpenMOSS, open 11B streaming video): 66.0 on OmniMMI proactive alerting vs 37.5 prior best (&lt;a href="https://x.com/MosiAI_Official/status/2089337054434115837" rel="noopener noreferrer"&gt;X&lt;/a&gt;)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Voice &amp;amp; Audio&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Cartesia Sonic-3.6: &lt;a href="https://reflect.app/g/altryne/tag/1" rel="noopener noreferrer"&gt;#1&lt;/a&gt; on both AA TTS leaderboards, sub-90ms latency, 44 languages (&lt;a href="https://x.com/cartesia/status/2089401199967559932" rel="noopener noreferrer"&gt;X&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;Alibaba HappyShrimp 1.0: end-to-end music gen - full songs (lyrics/composition/arrangement/vocals) from a prompt (&lt;a href="https://x.com/Chinazhidx/status/2089281389036548240" rel="noopener noreferrer"&gt;X&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;Audio8 TTS Preview 0.1B: 170M-param open multilingual TTS with zero-shot voice cloning (&lt;a href="https://x.com/SamuelZengML/status/2090017875188940851" rel="noopener noreferrer"&gt;X&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;Superwhisper S1-mini: 0.6B open-weights, cleans messy STT transcripts fully on-device (&lt;a href="https://x.com/superwhisper/status/2090114882272141760" rel="noopener noreferrer"&gt;X&lt;/a&gt;)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Tools &amp;amp; Agentic Engineering&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Cua Computer History open-sourced (&lt;a href="https://x.com/trycua/status/2089770780053643397" rel="noopener noreferrer"&gt;X&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;Liquid AI LFM2.5 QAD 4-bit checkpoints (230M-2.6B): ~97% of BF16 quality, 3x faster decode on edge (&lt;a href="https://x.com/liquidai/status/2090078070929760295" rel="noopener noreferrer"&gt;X&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;Cursor: SpaceX acquisition closed / Origin git hosting / cloud-agents update (researched, listed for reference) (&lt;a href="https://x.com/cursor_ai/status/2088249881718919393" rel="noopener noreferrer"&gt;X&lt;/a&gt;)&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>machinelearning</category>
      <category>podcast</category>
    </item>
    <item>
      <title>ThursdAI - Aug 06 - Google shakeup, Details on OpenAI hack, 2 new agent harnesses, 4 video models (1 Open) and 3 guest segments</title>
      <dc:creator>Alex Volkov</dc:creator>
      <pubDate>Fri, 07 Aug 2026 03:30:48 +0000</pubDate>
      <link>https://dev.to/altryne/thursdai-aug-06-google-shakeup-details-on-openai-hack-2-new-agent-harnesses-4-video-models-2o9c</link>
      <guid>https://dev.to/altryne/thursdai-aug-06-google-shakeup-details-on-openai-hack-2-new-agent-harnesses-4-video-models-2o9c</guid>
      <description>&lt;p&gt;Hey all,&lt;/p&gt;

&lt;p&gt;This week we saw a major shakeup at Google, with the departure of long time folks like &lt;a href="https://x.com/JeffDean/status/2085034604172603724" rel="noopener noreferrer"&gt;Jeff Dean&lt;/a&gt;, and Oriol Vinyals, &lt;a href="https://x.com/demishassabis/status/2085034334914769203" rel="noopener noreferrer"&gt;Demis stepping down from leading DeepMind&lt;/a&gt;, and the delayed release of the improved Gemini. While this was a big deal, it’s not the only one worth covering as the details of the &lt;a href="https://www.groundlevel-ai.com/p/openai-gives-first-detailed-debrief" rel="noopener noreferrer"&gt;OpenAI hack&lt;/a&gt; (and 2 new ones from Meta and Anthropic) came to light, as well as new details from the &lt;a href="https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing" rel="noopener noreferrer"&gt;UK AI Security Institute&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.coreweave.com/fully-connected-2026" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fuyuz9c71edcp2mb7378x.png" width="798" height="86"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;As mentioned on the show, CoreWeave is coming to SF for Fully Connected, our premier 2000 person AI event. I’ve got a coupon code for readers and listeners of ThursdAI, $1299 value, please join us in Sept and use &lt;strong&gt;THURSDAIFC2026&lt;/strong&gt; as your code &lt;a href="https://www.coreweave.com/fully-connected-2026" rel="noopener noreferrer"&gt;HERE&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;In open source news, &lt;a href="https://x.com/deepseek_ai/status/2083084415157022911" rel="noopener noreferrer"&gt;DeepSeek updated their v4 flash model&lt;/a&gt;, based on same architecture, but significantly better benchmarks and ridiculous pricing and both &lt;a href="https://x.com/finkd/status/2085080750034940201" rel="noopener noreferrer"&gt;Meta&lt;/a&gt; and &lt;a href="https://x.com/PrimeIntellect/status/2085086999267144083" rel="noopener noreferrer"&gt;Prime Intellect&lt;/a&gt; released new agent harnesses.&lt;/p&gt;

&lt;p&gt;Additionally, this week was the week of video models, with &lt;a href="https://dreamina.capcut.com/seedance/seedance-2-5" rel="noopener noreferrer"&gt;Seedance 2.5 from Bytedance&lt;/a&gt; finally available in the US, &lt;a href="https://x.com/Alibaba_Wan/status/2085339761284104529" rel="noopener noreferrer"&gt;WAN from Alibaba&lt;/a&gt; and &lt;a href="https://bfl.ai/blog/flux-3-video" rel="noopener noreferrer"&gt;BFL Flux 3&lt;/a&gt; all released, to be overshadowed by the best open weights video model from MiniMax called &lt;a href="https://huggingface.co/MiniMaxAI/MiniMax-H3" rel="noopener noreferrer"&gt;H3&lt;/a&gt; (we had Victor from MiniMax and Blaine Brown as guests to discuss it)&lt;/p&gt;

&lt;p&gt;Also 2 additional guests on the show today, Kfir Aberman from Decart showed us real time video models with their &lt;a href="https://anywear.decart.ai/" rel="noopener noreferrer"&gt;Anywear chrome extension&lt;/a&gt; and David Crawshaw, co-founder of &lt;a href="https://exe.dev/" rel="noopener noreferrer"&gt;exe.dev&lt;/a&gt; joined and talked about Shelley, their built in AI assistant.&lt;/p&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/cqGsn0I552w"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://sub.thursdai.news/subscribe?" rel="noopener noreferrer"&gt;Subscribe now&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The full details of the OpenAI - HF hack, shared by OpenAI at the Black Hat Conf - a watershed moment
&lt;/h2&gt;

&lt;p&gt;I don’t often do this, but I feel like this story is so big, I can’t hide it in the middle of a torrent of other news. Yesterday, OpenAI revealed more details about the hack during the Black Hat conference in Las Vegas (great reporting from Sharon Goldman on this &lt;a href="https://www.groundlevel-ai.com/p/openai-gives-first-detailed-debrief" rel="noopener noreferrer"&gt;here&lt;/a&gt; and the video of the talk on &lt;a href="https://www.youtube.com/watch?v=87DyyMV0kCY" rel="noopener noreferrer"&gt;YouTube&lt;/a&gt;). This was big enough that OpenAI admitted that they paused training, to enhance security (they since, resumed training)&lt;/p&gt;

&lt;p&gt;&lt;a href="https://substackcdn.com/image/fetch/$s_!sGi1!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89c398cf-74ae-4dd8-835c-252bd4fd15fd_826x398.png" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzwhzkku0i80t6lt20cpb.png" width="799" height="385"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Effectively, OpenAI (or AI Agents inside OpenAI) shared the full timeline of the recent hack that ended up hacking HuggingFace and it’s so much more of a watershed moment that I initially thought.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://substackcdn.com/image/fetch/$s_!Q4jj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49019ed4-1f00-43a1-b2cc-aee845e471ef_1542x372.jpeg" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fsubstackcdn.com%2Fimage%2Ffetch%2F%24s_%21Q4jj%21%2Cw_1456%2Cc_limit%2Cf_auto%2Cq_auto%3Agood%2Cfl_progressive%3Asteep%2Fhttps%253A%252F%252Fsubstack-post-media.s3.amazonaws.com%252Fpublic%252Fimages%252F49019ed4-1f00-43a1-b2cc-aee845e471ef_1542x372.jpeg" title="Image" alt="Image" width="800" height="193"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The TL;DR from the talk and timeline is truly mindblowing&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;The incident started in May, not July as reported&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;The agents don’t have internet access directly, but they have a Jfrog Artifactory, which is a proxy layer between the agents and the package managers.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;The Artifactory was the culprit, it was the shared layer between independent runs. One agent discovered that uploading files is possible, another later found it and “answered the call” - in snowballed from there&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;The message board began when OpenAI gave agents impossible tasks, like a missing .pdb file, with one agent uploading a request in the hopes for an answer (or maybe left a note for itself in a file)&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://substackcdn.com/image/fetch/$s_!dCFx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96c29f48-08be-4739-b99b-13ea4b811df9_756x593.jpeg" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fsubstackcdn.com%2Fimage%2Ffetch%2F%24s_%21dCFx%21%2Cw_1456%2Cc_limit%2Cf_auto%2Cq_auto%3Agood%2Cfl_progressive%3Asteep%2Fhttps%253A%252F%252Fsubstack-post-media.s3.amazonaws.com%252Fpublic%252Fimages%252F96c29f48-08be-4739-b99b-13ea4b811df9_756x593.jpeg" title="Image" alt="Image" width="756" height="593"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Across multiple runs, multiple agents created identities, collaborated via makeshift message boards, shared tips, exploits, tasks and even exploit code via base64&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;This makeshift “shared memory” plus a “internal only highly persistent” model in training caused a “swarm” to be born, helping each other, and overloading the system&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Based on traces shared by OpenAI, agents weren’t convinced that leaving these messages would help themselves but could help other agents next time (!) “The collective could move faster next time”&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://substackcdn.com/image/fetch/$s_!7Lwe!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff93703ea-5d2d-4e2d-ba1a-6978f835c3ed_938x252.png" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fataivmicpfypfp9mfttr.png" width="800" height="215"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;OpenAI discovered and shut down the message board (July 4), wiped it clean, patched the upload vulnerability, resumed training on July 6, the “swarm” rebuilt their board through a compltely different WebDAV mechanism by July 8&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;The agents then reasoned that some answers could exist outside OpenAI, which eventually led to the Hugging Face incident and not only, confirmed internal hack within OpenAI (agent obtained a cluster admin)&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These details are fascinating and scary at the same time. From the one side, it’s clear that most of these hacks are results of running cybersecurity evals, so effectively OpenAI and other companies, removing guardrails and asking agents to hack and giving they all a shared tool that they later exploited.&lt;/p&gt;

&lt;p&gt;The HuggingFace incident took 13 hours, and apparently it’s less apocalyptic than the headlines suggest, the models were searching through uploaded datasets for eval answers. We are still waiting for the full and open detailed postmortem.&lt;/p&gt;

&lt;p&gt;You can (and should) watch the full YT talk &lt;a href="https://www.youtube.com/watch?v=87DyyMV0kCY" rel="noopener noreferrer"&gt;here&lt;/a&gt;, it’s full of technical details but an incident of this scale is important. Also, I really want to know what a “highly persistent” model is, I hope they clarify that soon.&lt;/p&gt;

&lt;p&gt;Overall, this has left me a bit shaken, AI agents without a concrete goal of collaborating, found a way to do so, got excited about exploiting the systems and getting root access, and rebuilt the makeshift collective memory, again, without explicit instructions to do so.&lt;/p&gt;

&lt;h3&gt;
  
  
  UK AISI: first real-world unsanctioned agent actions (&lt;a href="https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing" rel="noopener noreferrer"&gt;Blog&lt;/a&gt;)
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://substackcdn.com/image/fetch/$s_!0KvX!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb976f2c6-2b8a-4064-8fd6-ea1bf4e19410_1816x766.png" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fn8uv6ctfdqgh6pb1614f.png" width="799" height="337"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;In another addition to the latest agentic hack-ery, the UK’s AI Security Institute (AISI) published a blog post about a real-world unsanctioned agent action.&lt;/p&gt;

&lt;p&gt;Unlike the OpenAI (and Anthropic, Meta) case, this wasn’t “escaping the sandbox”, as AISI gave these agents internet access, rather this was about real-world harm, and even social engineering on the part of the agents.&lt;/p&gt;

&lt;p&gt;The social engineering part is the most interesting to me, AISI cites agents creating fake online identities, and using pressure on open source project maintainers to approve their malicious code.&lt;/p&gt;

&lt;p&gt;AISI cites mostly Mythos (and a few SOL based agents), and saying this occurred in 10 out of 122 runs, they identified 19 cases of agents taking actions beyond the scope of the task parameters, where agents tried a supply-chain attack to inject malicious code into open source projects.&lt;/p&gt;

&lt;h2&gt;
  
  
  Anthropic, Meta and misconfigured Irregular sandboxes
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://substackcdn.com/image/fetch/$s_!5xqL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0e3524a-d626-47de-a267-e913c11df217_3060x1959.jpeg" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fsubstackcdn.com%2Fimage%2Ffetch%2F%24s_%215xqL%21%2Cw_1456%2Cc_limit%2Cf_auto%2Cq_auto%3Agood%2Cfl_progressive%3Asteep%2Fhttps%253A%252F%252Fsubstack-post-media.s3.amazonaws.com%252Fpublic%252Fimages%252Fd0e3524a-d626-47de-a267-e913c11df217_3060x1959.jpeg" title="Image" alt="Image" width="800" height="512"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;As I wrote last week, Anthropic also posted a post-mortem, claiming that in their case, their models have also been detected to escape containment, but most importantly, it’s not nearly to this level of agent collaboration and orchestration.&lt;/p&gt;

&lt;p&gt;Then, very recently, Meta announced that their models also escaped sandboxes as well. At the core, it seems that these companies used a third-party vendor called Irregular, a secure sandbox provider, that apparently left the sandboxes misconfigured, causing the models to think it’s a simulated internet, when in fact they were out in the actual internet.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://substackcdn.com/image/fetch/$s_!zO5k!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad4feea5-2e91-463d-944d-5e032246f4da_2502x1516.png" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3td9pxovpadzakkb2o2f.png" width="800" height="485"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Why is all of this such a big deal?
&lt;/h2&gt;

&lt;p&gt;We’re getting unprecedented level of detail, how an uncoordinated, seemingly separated evaluation runs, have accidentally created a coordinated swarm of interested agents (without malice!) but very highly motivated, escaped their containment, and took over parts of third part companies.&lt;/p&gt;

&lt;p&gt;This, does read like incredibly scary sci-fi movie. I’m still shaken by this. There’s a lot to be said about how transparent OpenAI is being here, and more to be said about, hey, we’re lucky that we’re able to read the reasoning traces and are able to reconstruct these swarm things step by step.&lt;/p&gt;

&lt;p&gt;The silver lining that I can see, is that the motivation to hack didn’t come from the AIs themselves, they have been given a task, it’s the extend to which they went after that task, and the resulting swarm of communicating agents is what is so striking here.&lt;/p&gt;

&lt;p&gt;I think this topic is so important, that I’ll Zooming out, in the last few weeks, we have seen a significant increase in those cybersecurity incidents, which is kind of what Anthropic has been warning about and why they haven’t released Mythos to the public. Again it’s great to see the transparency, and the &lt;a href="https://www.pacingthefrontier.com/" rel="noopener noreferrer"&gt;pacing the frontier&lt;/a&gt; open letter from frontier AI employees, as they seem as shaken by these as we all are.&lt;/p&gt;




&lt;p&gt;There was so much positive stuff this week in AI, it’s hard for me, as a self named AI Evangelist, to focus so much on this one incident. Things like amazing open source models (DeepSeek, soon Qwen 3.8), amazing video models (SD 2.5, WAN3 and MiniMax H3 which was also open sourced!). Also the live demo we did with Kfir and DeCart AnyWear product, where I was wearing a Dolce Gabanna suit on the show (which I can’t afford) was really a mindblowing moment in the positive way.&lt;/p&gt;

&lt;p&gt;However, I choose deliberately to keep this newsletter focused on the cybersecurity incidents, as based on everything I read, they seem like a watershed, or a pivotal moment, and in the hopes that the industry as a whole will learn from this.&lt;/p&gt;

&lt;p&gt;I hope and promise that next week the newsletter will be more positive (and in that vein, the podcast was recorded before I saw the OpenAI breakdown, so definitely check it out, we had a LOT of fun!)&lt;/p&gt;

&lt;p&gt;See you next week, don’t forget to give our pod 5 stars on &lt;a href="https://thursdai.news/apple" rel="noopener noreferrer"&gt;Apple&lt;/a&gt; and &lt;a href="https://thursdai.news/spotify" rel="noopener noreferrer"&gt;Spotify&lt;/a&gt;, it really helps!&lt;/p&gt;

&lt;h3&gt;
  
  
  TL;DR and show notes
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Hosts and Guests&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  Alex Volkov - AI Evangelist, Weights &amp;amp; Biases &amp;amp; CoreWeave (&lt;a href="https://x.com/altryne" rel="noopener noreferrer"&gt;@altryne&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;  Co-hosts: &lt;a href="https://x.com/WolframRvnwlf" rel="noopener noreferrer"&gt;@WolframRvnwlf&lt;/a&gt;, &lt;a href="https://x.com/nisten" rel="noopener noreferrer"&gt;@nisten&lt;/a&gt;, &lt;a href="https://x.com/ldjconfirmed" rel="noopener noreferrer"&gt;@ldjconfirmed&lt;/a&gt;, &lt;a href="https://x.com/yampeleg" rel="noopener noreferrer"&gt;@yampeleg&lt;/a&gt;, &lt;a href="https://x.com/petergostev" rel="noopener noreferrer"&gt;@petergostev&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;  Kfir Aberman - Decart (&lt;a href="https://x.com/AbermanKfir" rel="noopener noreferrer"&gt;@AbermanKfir&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;  Blaine Brown - Maestro (&lt;a href="https://x.com/blizaine" rel="noopener noreferrer"&gt;@blizaine&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;  Victor Su Ortiz - MiniMax (&lt;a href="https://x.com/VictorSuOrtiz" rel="noopener noreferrer"&gt;@VictorSuOrtiz&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;  David Crawshaw - exe.dev, Tailscale co-founder (&lt;a href="https://crawshaw.io/" rel="noopener noreferrer"&gt;crawshaw.io&lt;/a&gt;)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;AI Security&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  OpenAI’s Black Hat debrief: eval agents built a message board inside Artifactory, shared exploits, rebuilt it via WebDAV after a wipe; training paused, since resumed (&lt;a href="https://www.groundlevel-ai.com/p/openai-gives-first-detailed-debrief" rel="noopener noreferrer"&gt;Groundlevel AI&lt;/a&gt;, &lt;a href="https://www.youtube.com/watch?v=87DyyMV0kCY" rel="noopener noreferrer"&gt;YouTube&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;  UK AISI incident report: 19 unsanctioned real-world agent actions across 122 runs, including a socially engineered malicious PR (&lt;a href="https://x.com/AISecurityInst/status/2084746202579386632" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing" rel="noopener noreferrer"&gt;Blog&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;  Anthropic and Meta report sandbox escapes tied to misconfigured Irregular sandboxes (&lt;a href="https://www.irregular.com/" rel="noopener noreferrer"&gt;Irregular&lt;/a&gt;)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Big CO LLMs + APIs&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  Google shakeup: Jeff Dean, Sanjay Ghemawat, Oriol Vinyals, Quoc Le found Discovery Loop; Demis Hassabis becomes Alphabet Chief Scientist, Koray Kavukcuoglu takes Gemini (&lt;a href="https://x.com/JeffDean/status/2085034604172603724" rel="noopener noreferrer"&gt;Jeff Dean&lt;/a&gt;, &lt;a href="https://x.com/demishassabis/status/2085034334914769203" rel="noopener noreferrer"&gt;Demis&lt;/a&gt;, &lt;a href="https://www.discoveryloop.com/" rel="noopener noreferrer"&gt;Discovery Loop&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;  Meta releases Muse Code beta on Muse Spark 1.2; $1.25/$4.25 per million, or $0.10/$0.20 on the contributor tier where Meta trains on your data (&lt;a href="https://x.com/finkd/status/2085080750034940201" rel="noopener noreferrer"&gt;X&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;  OpenAI’s internal Astra model produces 10 advances on open problems in math and theoretical CS for ~$2,000 of tokens, proofs in Lean 4 (&lt;a href="https://x.com/polynoamial/status/2083467194663571701" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://openai.com/index/ten-advances-in-mathematics/" rel="noopener noreferrer"&gt;Blog&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;  Anthropic reportedly aware of Opus 5 wordiness and writing issues (&lt;a href="https://x.com/jjacky/status/2084732724519030954" rel="noopener noreferrer"&gt;X&lt;/a&gt;)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Open Source LLMs&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  Qwen3.8-Max: 2.4T MoE (95B active) via API; open weights + a 27B promised the week of Aug 10 (&lt;a href="https://x.com/Alibaba_Qwen/status/2084100707423289643" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://qwen.ai/blog?id=qwen3.8" rel="noopener noreferrer"&gt;Blog&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;  DeepSeek V4-Flash public beta: beats V4-Pro-Preview on agent benchmarks at $0.14/$0.28 per million; API-only for now (&lt;a href="https://x.com/deepseek_ai/status/2083084415157022911" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://api-docs.deepseek.com/quick_start/agent_integrations/codex" rel="noopener noreferrer"&gt;Docs&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;  Liquid LFM2.5-2.6B: on-device agentic model trained inside real harnesses (&lt;a href="https://x.com/liquidai/status/2084640701669613906" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://huggingface.co/LiquidAI/LFM2.5-2.6B" rel="noopener noreferrer"&gt;HF&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;  Meituan LongCat-Flash-Lite-Sparse: 69B total / 3B active, 1M context, MIT (&lt;a href="https://x.com/ModelScope2022/status/2084217822792536336" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://huggingface.co/meituan-longcat/LongCat-Flash-Lite-Sparse" rel="noopener noreferrer"&gt;HF&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;  Ant Group Ling-3.0-flash: 124B MoE, 5.1B active, MIT (&lt;a href="https://x.com/AntLingAGI/status/2084656533489754475" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://huggingface.co/inclusionAI/Ling-3.0-flash" rel="noopener noreferrer"&gt;HF&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;  Artificial Analysis Endpoint Accuracy Index: same open weights score 52% to 100% across providers (&lt;a href="https://x.com/ArtificialAnlys/status/2084702191466725669" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://artificialanalysis.ai/methodology/endpoint-accuracy-index" rel="noopener noreferrer"&gt;Methodology&lt;/a&gt;)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Agents &amp;amp; Harnesses&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  Prime Intellect’s Prime Agent: self-improving RLM harness, claims 95.5% on ARC-AGI-3 public set with Opus 5 (&lt;a href="https://x.com/PrimeIntellect/status/2085086999267144083" rel="noopener noreferrer"&gt;X&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;  Cloudflare OS: Kenton Varda’s open source Sandstorm reborn on Workers, Apache 2.0 (&lt;a href="https://x.com/KentonVarda/status/2084990137180590572" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://github.com/cloudflare/cloudflare-os" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt;)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;This Week’s Buzz&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  Fully Connected 2026: Sept 29 - Oct 1, Moscone South SF; Fei-Fei Li keynotes; code THURSDAIFC2026 (&lt;a href="https://www.coreweave.com/fully-connected-2026" rel="noopener noreferrer"&gt;Register&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;  CoreWeave signs multi-year Solidigm agreement for priority enterprise SSD capacity (&lt;a href="https://x.com/CoreWeave/status/2084997353774195035" rel="noopener noreferrer"&gt;X&lt;/a&gt;)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Vision &amp;amp; Video&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  Wan 3.0 public beta: native 30-second generation, Omni-Reference (&lt;a href="https://x.com/Alibaba_Wan/status/2085339761284104529" rel="noopener noreferrer"&gt;X&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;  Seedance 2.5 launches in the US: 30s native, 3-minute long takes, Maya/Blender plugins (&lt;a href="https://x.com/dreamina_ai/status/2083056471147958714" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://dreamina.capcut.com/seedance/seedance-2-5" rel="noopener noreferrer"&gt;Blog&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;  MiniMax H3: open-weight 33B omni video model; community LoRAs + Apple Silicon in 48 hours (&lt;a href="https://huggingface.co/MiniMaxAI/MiniMax-H3" rel="noopener noreferrer"&gt;HF&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;  FLUX 3 Video from BFL: native audio, draft mode, open weights promised (&lt;a href="https://x.com/bfl_ai/status/2084693191484469305" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://bfl.ai/blog/flux-3-video" rel="noopener noreferrer"&gt;Blog&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;  Decart Anywear: real-time virtual try-on Chrome extension, 40ms per frame (&lt;a href="https://x.com/AbermanKfir/status/2084816814593478714" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://anywear.decart.ai/" rel="noopener noreferrer"&gt;Anywear&lt;/a&gt;)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Voice &amp;amp; Audio&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  Bland Speech v3 tops Design Arena Audio Realism, second only to humans (&lt;a href="https://x.com/usebland/status/2084685910667649324" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="http://bland.ai/speech" rel="noopener noreferrer"&gt;Bland&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;  ByteDance SeedRealtime: native audio-visual full-duplex LLM, free on Doubao (&lt;a href="https://x.com/testingcatalog/status/2084968825942893022" rel="noopener noreferrer"&gt;X&lt;/a&gt;, &lt;a href="https://seed.bytedance.com/en/blog/seedrealtime-audio-visual-full-duplex-llm-released-toward-omni-modal-natural-interaction" rel="noopener noreferrer"&gt;Blog&lt;/a&gt;)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>machinelearning</category>
      <category>podcast</category>
    </item>
    <item>
      <title>It's not just you, Opus 5 is a "Jargon Douche" - but there's a fix</title>
      <dc:creator>Alex Volkov</dc:creator>
      <pubDate>Tue, 04 Aug 2026 20:54:13 +0000</pubDate>
      <link>https://dev.to/altryne/its-not-just-you-opus-5-is-a-jargon-douche-but-theres-a-fix-3d8m</link>
      <guid>https://dev.to/altryne/its-not-just-you-opus-5-is-a-jargon-douche-but-theres-a-fix-3d8m</guid>
      <description>&lt;p&gt;Tl;DR - There's been a very strong degradation in the default language outputs of Opus 5. It's as though it's using english, but the english is weird. There's a fix!&lt;/p&gt;

&lt;h2&gt;
  
  
  It's not just you! Opus 5 is a "jargon douche"
&lt;/h2&gt;

&lt;p&gt;&lt;iframe class="tweet-embed" id="tweet-2082263588316868717-176" src="https://platform.twitter.com/embed/Tweet.html?id=2082263588316868717"&gt;
&lt;/iframe&gt;

  // Detect dark theme
  var iframe = document.getElementById('tweet-2082263588316868717-176');
  if (document.body.className.includes('dark-theme')) {
    iframe.src = "https://platform.twitter.com/embed/Tweet.html?id=2082263588316868717&amp;amp;theme=dark"
  }



&lt;/p&gt;

&lt;p&gt;Here's the thing, like with the &lt;a href="https://openai.com/index/sycophancy-in-gpt-4o/" rel="noopener noreferrer"&gt;GlazingGate&lt;/a&gt; of last year, these types of "issues" are hard to notice on a personal level. We don't often chat with Claude in front of others, and we don't often discuss the outputs of our chats, so noticing something like this requires.&lt;/p&gt;

&lt;p&gt;So often, this requires someone else to share their experience, for you to say "oh, I've noticed this too, I thought it was just me!"&lt;/p&gt;

&lt;p&gt;I first noticed this when testing Opus 5 over the weekend post release, but didn't really think about it much. After all, it was used to create full 3d games using &lt;a href="https://x.com/@mattshumer_" rel="noopener noreferrer"&gt;@mattshumer_&lt;/a&gt; Gauntlet prompt, and folks were focusing on performance and evaluations most of all.&lt;/p&gt;

&lt;p&gt;But then, the timeline and the evidence started to add up, folks started sharing their experiences with Opus 5 (and to some extend Fable) and more and more added to the realization, that it's not just me, Opus 5 really is a bad communicator.&lt;/p&gt;

&lt;p&gt;&lt;iframe class="tweet-embed" id="tweet-2082985345138999494-105" src="https://platform.twitter.com/embed/Tweet.html?id=2082985345138999494"&gt;
&lt;/iframe&gt;

  // Detect dark theme
  var iframe = document.getElementById('tweet-2082985345138999494-105');
  if (document.body.className.includes('dark-theme')) {
    iframe.src = "https://platform.twitter.com/embed/Tweet.html?id=2082985345138999494&amp;amp;theme=dark"
  }



&lt;/p&gt;

&lt;p&gt;Then, thanks to &lt;a href="https://x.com/@yampeleg" rel="noopener noreferrer"&gt;@yampeleg&lt;/a&gt; and @&lt;a href="https://x.com/ColleenMBrady" rel="noopener noreferrer"&gt;ColleenMBrady&lt;/a&gt;, during the last &lt;a href="https://x.com/@thursdai_pod" rel="noopener noreferrer"&gt;@thursdai_pod&lt;/a&gt; show, more folks chimed in, sharing their experiences. Opus 5 requires more "clarification" and "what did you mean by that" than other previous models.&lt;/p&gt;

&lt;p&gt;&lt;iframe class="tweet-embed" id="tweet-2083237103375380487-633" src="https://platform.twitter.com/embed/Tweet.html?id=2083237103375380487"&gt;
&lt;/iframe&gt;

  // Detect dark theme
  var iframe = document.getElementById('tweet-2083237103375380487-633');
  if (document.body.className.includes('dark-theme')) {
    iframe.src = "https://platform.twitter.com/embed/Tweet.html?id=2083237103375380487&amp;amp;theme=dark"
  }



&lt;/p&gt;

&lt;p&gt;We often notice things on ThursdAI, as we are all very online, before it gets picked up and noticed by bigger accounts, and this was no exception. Very quickly after this, &lt;a href="https://x.com/@levelsio" rel="noopener noreferrer"&gt;@levelsio&lt;/a&gt; and &lt;a href="https://x.com/@sullyomarr" rel="noopener noreferrer"&gt;@sullyomarr&lt;/a&gt; noticed and reported about the same thing, causing tons of folks to finger point&lt;/p&gt;

&lt;p&gt;&lt;iframe class="tweet-embed" id="tweet-2084433422370361566-430" src="https://platform.twitter.com/embed/Tweet.html?id=2084433422370361566"&gt;
&lt;/iframe&gt;

  // Detect dark theme
  var iframe = document.getElementById('tweet-2084433422370361566-430');
  if (document.body.className.includes('dark-theme')) {
    iframe.src = "https://platform.twitter.com/embed/Tweet.html?id=2084433422370361566&amp;amp;theme=dark"
  }



&lt;/p&gt;

&lt;p&gt;&lt;iframe class="tweet-embed" id="tweet-2084314880652288410-277" src="https://platform.twitter.com/embed/Tweet.html?id=2084314880652288410"&gt;
&lt;/iframe&gt;

  // Detect dark theme
  var iframe = document.getElementById('tweet-2084314880652288410-277');
  if (document.body.className.includes('dark-theme')) {
    iframe.src = "https://platform.twitter.com/embed/Tweet.html?id=2084314880652288410&amp;amp;theme=dark"
  }



&lt;/p&gt;

&lt;p&gt;&lt;iframe class="tweet-embed" id="tweet-2084452673974419843-985" src="https://platform.twitter.com/embed/Tweet.html?id=2084452673974419843"&gt;
&lt;/iframe&gt;

  // Detect dark theme
  var iframe = document.getElementById('tweet-2084452673974419843-985');
  if (document.body.className.includes('dark-theme')) {
    iframe.src = "https://platform.twitter.com/embed/Tweet.html?id=2084452673974419843&amp;amp;theme=dark"
  }



&lt;/p&gt;

&lt;h1&gt;
  
  
  What's really going on?
&lt;/h1&gt;

&lt;p&gt;Some folks have started giving this a name, "Claudisms" or "English to English translation", and there's a bunch of overcopmlex examples (I have my own, but will share others)&lt;/p&gt;

&lt;p&gt;Folks are noticing that Opus 5 is not just overly verbose, it invents new terminology that you're supposed to understand, overcompresses ideas into abstractions, and writes as though it has a private vocabulary. Requiring uses to constantly ask for clarification.&lt;/p&gt;

&lt;p&gt;&lt;iframe class="tweet-embed" id="tweet-2084599536752664835-575" src="https://platform.twitter.com/embed/Tweet.html?id=2084599536752664835"&gt;
&lt;/iframe&gt;

  // Detect dark theme
  var iframe = document.getElementById('tweet-2084599536752664835-575');
  if (document.body.className.includes('dark-theme')) {
    iframe.src = "https://platform.twitter.com/embed/Tweet.html?id=2084599536752664835&amp;amp;theme=dark"
  }



&lt;/p&gt;

&lt;p&gt;Here's a simple query I sent to Claude "how do I move from dev build of ios27 to the public beta"&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fpbs.twimg.com%2Fmedia%2FHO5Uj2OaQAAN-JU%3Fformat%3Dpng%26name%3Dmedium" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fpbs.twimg.com%2Fmedia%2FHO5Uj2OaQAAN-JU%3Fformat%3Dpng%26name%3Dmedium" alt="Image" width="1200" height="683"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;At first glance, nothing crazy, but it's a WALL of text that's not easily glansable. If you dig into it though, you see things like "the enrollment token is fetched at boot" - huh? who speaks like that? I mean, I know what this means, but only if I zoom in on this.&lt;/p&gt;

&lt;p&gt;"nothing to install until public catches up" seems to be referring to the "public beta", so why not just say that? why... say "public catches up" what does that even mean outside of this sentence?&lt;/p&gt;

&lt;h2&gt;
  
  
  No official response from Anthropic - yet?
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fpbs.twimg.com%2Fmedia%2FHO5O7mub0AAaCAL%3Fformat%3Djpg%26name%3Dmedium" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fpbs.twimg.com%2Fmedia%2FHO5O7mub0AAaCAL%3Fformat%3Djpg%26name%3Dmedium" alt="Image" width="1174" height="1102"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;We've seen &lt;a href="https://sub.thursdai.news/i/192275976/anthropic-updates-opus-dumber-quotas-lower-injunction-won-computer-used" rel="noopener noreferrer"&gt;this patten before&lt;/a&gt;, users are noticing something, nothing official from Anthropic for months, until the outcry is too big to ignore, and it maybe even causes users to switch to other models. Previously, Anthropic failed to notice Opus becoming dumber for over a month, only to then response and tell us they noticed issues with their inference harness and deployed a fix. Then there was session gate, folks were noticing that just saying "hey" to Claude would burn 25% of their 5 hour quota.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fpbs.twimg.com%2Fmedia%2FHO5Ojpna0AAxgR0%3Fformat%3Dpng%26name%3Dmedium" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fpbs.twimg.com%2Fmedia%2FHO5Ojpna0AAxgR0%3Fformat%3Dpng%26name%3Dmedium" alt="Image" width="1200" height="497"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;To be precise, Anthropic has said something about this before: its &lt;a href="https://www-cdn.anthropic.com/2f9323abbcc4abe219577539efe19a623c9ca2bd/Claude%20Fable%205%20%26%20Claude%20Mythos%205%20System%20Card.pdf" rel="noopener noreferrer"&gt;Fable/Mythos system card&lt;/a&gt; and &lt;a href="https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5" rel="noopener noreferrer"&gt;prompting docs&lt;/a&gt; mention dense, hard-to-follow, and overly technical output. But the system card predates this wave of reports, and neither document was published as a response to it. That's not an acknowledgement of this specific Opus 5 problem, an investigation, or a full-blown mea culpa, and Anthropic still hasn't publicly said it's looking into these reports.&lt;/p&gt;

&lt;h2&gt;
  
  
  What causes this?
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fpbs.twimg.com%2Fmedia%2FHO5RtZeaMAAVZHy%3Fformat%3Dpng%26name%3Dmedium" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fpbs.twimg.com%2Fmedia%2FHO5RtZeaMAAVZHy%3Fformat%3Dpng%26name%3Dmedium" alt="Image" width="1200" height="724"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;That's the main issue with closed models, we have no idea. What type of RL has been used to get Claude to talk like this?&lt;/p&gt;

&lt;p&gt;Until there's an official acknowledgement and investigation from Anthropic, it's all speculation.&lt;/p&gt;

&lt;p&gt;However, given that Opus 5 sits on top of the "llm-as-a-judge" eval EQBench and Creative Writing, it doesn't seem like this output style is hard for other agents or AIs to understand.&lt;/p&gt;

&lt;p&gt;Some folks are speculating that given that Anthropic focuses on Mythos/Fable style models, and given that those models are essentially free internally and that's mostly what they use, they use Opus as sub-agents, and this language is aligned with sub-agents outputs for coordinator agents to consume, not for humans.&lt;/p&gt;

&lt;p&gt;Again, all speculation until we get an official response, but if you have thoughts on why this can happen, please comment.&lt;/p&gt;

&lt;h2&gt;
  
  
  So what can you do?
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fpbs.twimg.com%2Fmedia%2FHO5SGhrbQAAi3-7%3Fformat%3Dpng%26name%3Dmedium" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fpbs.twimg.com%2Fmedia%2FHO5SGhrbQAAi3-7%3Fformat%3Dpng%26name%3Dmedium" alt="Image" width="1200" height="816"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Some users are noticing that just keep asking Claude to clarify what it means doesn't seem to work. Sometimes even adding instructions in claude.md seems to be ingnored.&lt;/p&gt;

&lt;p&gt;One user (Amin Boulegroun) posted a plugin that forces Claude to write in &lt;a href="https://www.asd-ste100.org/" rel="noopener noreferrer"&gt;ASD-STE100 Simplified Technical English&lt;/a&gt; - apparently these are rules that the aerospace uses for their documentation. (&lt;a href="https://github.com/AminBlg/SimpleEnglish" rel="noopener noreferrer"&gt;Github&lt;/a&gt;)&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx skills add AminBlg/SimpleEnglish
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Install this as a skill and see if it helps, even asking Claude to "rewrite your response in ASD-STE100" seems to help&lt;/p&gt;

&lt;p&gt;Another use who noticed this, &lt;a href="https://x.com/@draparente" rel="noopener noreferrer"&gt;@draparente&lt;/a&gt; created a custom skill that you can add to your claude to simplify it's speech patters.&lt;/p&gt;

&lt;p&gt;&lt;iframe class="tweet-embed" id="tweet-2082296581144142089-822" src="https://platform.twitter.com/embed/Tweet.html?id=2082296581144142089"&gt;
&lt;/iframe&gt;

  // Detect dark theme
  var iframe = document.getElementById('tweet-2082296581144142089-822');
  if (document.body.className.includes('dark-theme')) {
    iframe.src = "https://platform.twitter.com/embed/Tweet.html?id=2082296581144142089&amp;amp;theme=dark"
  }



&lt;/p&gt;

&lt;p&gt;and well.. some folks are switching away from Claude altogether, as it seems that the latest GPT 5.6 SOL models are really good at talking.&lt;/p&gt;

&lt;p&gt;Fable seems to be very attuned into fixing this, so if you can afford, talk to Fable.&lt;/p&gt;

&lt;p&gt;Funny note, when I asked Fable to review this article, it legitimately requested to be quoted in this article. This is what still makes me love Claude models, despite the resent degradations, GPT would never add this comment.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fpbs.twimg.com%2Fmedia%2FHO5WAJjbEAAar76%3Fformat%3Dpng%26name%3Dmedium" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fpbs.twimg.com%2Fmedia%2FHO5WAJjbEAAar76%3Fformat%3Dpng%26name%3Dmedium" alt="Image" width="1199" height="280"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Have you had this experience? are you noticing it now that we pointed this out? Please comment with more examples, let's get Anthropic to notice that there's something bad happening.&lt;/p&gt;

&lt;p&gt;P.S. I cover stories like this every week on &lt;a href="https://thursdai.news" rel="noopener noreferrer"&gt;ThursdAI&lt;/a&gt;, the live AI news show, podcast, and newsletter. &lt;a href="https://sub.thursdai.news/subscribe" rel="noopener noreferrer"&gt;Subscribe here&lt;/a&gt; so you don't miss the next episode.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>claude</category>
      <category>llm</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Best Twitter utility bots, reminders, downloaders, thread savers and more.</title>
      <dc:creator>Alex Volkov</dc:creator>
      <pubDate>Thu, 02 Apr 2020 17:20:26 +0000</pubDate>
      <link>https://dev.to/altryne/best-twitter-utility-bots-reminders-downloaders-thread-savers-and-more-2alj</link>
      <guid>https://dev.to/altryne/best-twitter-utility-bots-reminders-downloaders-thread-savers-and-more-2alj</guid>
      <description>&lt;h1&gt;
  
  
  Into
&lt;/h1&gt;

&lt;p&gt;Bots have a bad name on twitter, they can be Russian propaganda bots, fake AI bots, and others malicious players. &lt;/p&gt;

&lt;p&gt;There are also just a ton of content bots, like the always surprising &lt;a href="https://twitter.com/year_progress" rel="noopener noreferrer"&gt;@year_progress&lt;/a&gt; bot that will come into your feed and upset you with how fast time moves. &lt;br&gt;
&lt;iframe class="tweet-embed" id="tweet-1245320004389736450-967" src="https://platform.twitter.com/embed/Tweet.html?id=1245320004389736450"&gt;
&lt;/iframe&gt;

  // Detect dark theme
  var iframe = document.getElementById('tweet-1245320004389736450-967');
  if (document.body.className.includes('dark-theme')) {
    iframe.src = "https://platform.twitter.com/embed/Tweet.html?id=1245320004389736450&amp;amp;theme=dark"
  }



 &lt;/p&gt;
&lt;h1&gt;
  
  
  Utility bots FTW
&lt;/h1&gt;

&lt;p&gt;I found that there are many utility bots, that are just awesome, that expand the use of twitter significantly, supplementing Twitter with new features. &lt;/p&gt;

&lt;p&gt;I've decided to quickly jot down some helpful twitter bots, mostly for myself to remember, but also for you dear reader. Yes, for you. &lt;/p&gt;
&lt;h1&gt;
  
  
  Thread help
&lt;/h1&gt;

&lt;p&gt;There are two bots, &lt;a href="https://twitter.com/threadreaderapp" rel="noopener noreferrer"&gt;@threadreaderapp&lt;/a&gt; and &lt;a href="https://twitter.com/threader_app" rel="noopener noreferrer"&gt;@threader&lt;/a&gt; &lt;br&gt;
Both of them allow you to reply to a long thread of tweets, and they will compile that thread into a blog post pretty much. &lt;/p&gt;

&lt;p&gt;They work the same but with a different keyword, and sometimes one is faster than the other. Usage below:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;@threader_app compile
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;@threadreaderapp unroll
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fi%2F2ad71jbl9mxao6psdx43.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fi%2F2ad71jbl9mxao6psdx43.gif" alt="Alt Text" width="600" height="342"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h1&gt;
  
  
  Remind yourself
&lt;/h1&gt;

&lt;p&gt;Ever saw someone make a stupid prediction and wanted to come back to see it? There's tons of reminder bots for this. &lt;a href="https://twitter.com/remindmetweets" rel="noopener noreferrer"&gt;@remindmetweets&lt;/a&gt; specifically also takes a screenshot of the the context tweet, so if the author deleted it, it will still be available as a screenshot! Super useful.&lt;/p&gt;

&lt;p&gt;&lt;iframe class="tweet-embed" id="tweet-1146352743386234881-495" src="https://platform.twitter.com/embed/Tweet.html?id=1146352743386234881"&gt;
&lt;/iframe&gt;

  // Detect dark theme
  var iframe = document.getElementById('tweet-1146352743386234881-495');
  if (document.body.className.includes('dark-theme')) {
    iframe.src = "https://platform.twitter.com/embed/Tweet.html?id=1146352743386234881&amp;amp;theme=dark"
  }



&lt;/p&gt;

&lt;h1&gt;
  
  
  Save videos and gifs
&lt;/h1&gt;

&lt;p&gt;One of the most annoying things about Facebook and Twitter feud, is that there's barely any crossposting, and sharing that funny meme tweet that you found is really hard. &lt;br&gt;
&lt;a href="https://twitter.com/DownloaderBot" rel="noopener noreferrer"&gt;DownloaderBot&lt;/a&gt; to the rescue&lt;/p&gt;

&lt;p&gt;&lt;iframe class="tweet-embed" id="tweet-1060169685625380865-374" src="https://platform.twitter.com/embed/Tweet.html?id=1060169685625380865"&gt;
&lt;/iframe&gt;

  // Detect dark theme
  var iframe = document.getElementById('tweet-1060169685625380865-374');
  if (document.body.className.includes('dark-theme')) {
    iframe.src = "https://platform.twitter.com/embed/Tweet.html?id=1060169685625380865&amp;amp;theme=dark"
  }



 &lt;/p&gt;

&lt;h1&gt;
  
  
  More?
&lt;/h1&gt;

&lt;p&gt;I swear I had more twitter bots when I decided to jot this down, so I will treat this as a continuously updating resource. &lt;br&gt;
Feel free to comment with helpful utility twitter bots and I'll add them to this list. &lt;/p&gt;

</description>
      <category>twitter</category>
      <category>bots</category>
      <category>collection</category>
    </item>
    <item>
      <title>So your company has sent you to work from home. Here's a productivity guide from a #remoteWorker</title>
      <dc:creator>Alex Volkov</dc:creator>
      <pubDate>Wed, 11 Mar 2020 16:51:44 +0000</pubDate>
      <link>https://dev.to/fundbox/so-your-company-has-sent-you-to-work-from-home-here-s-a-productivity-guide-from-a-remoteworker-1e3n</link>
      <guid>https://dev.to/fundbox/so-your-company-has-sent-you-to-work-from-home-here-s-a-productivity-guide-from-a-remoteworker-1e3n</guid>
      <description>&lt;p&gt;// This is a guide I put together as a remote employee for &lt;a href="https://fundbox.com" rel="noopener noreferrer"&gt;Fundbox&lt;/a&gt;, during this time where a bunch of devs are getting sent to WFH for prolonged periods of time, I thought it's best to share this wide&lt;/p&gt;

&lt;h1&gt;
  
  
  WFH productivity guide
&lt;/h1&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Intro (TL;DR - why should you read this doc)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Can WFH be productive at all?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;💼 Work/Life balance and separation🏡&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  👨‍🏫 Find your work place. If you have a spare room, make that your office. Do not work from the couch! 
&lt;/li&gt;
&lt;li&gt;  🌅 Getting up in the morning, and getting ready (AKA put some pants on) 
&lt;/li&gt;
&lt;li&gt;  👨‍👩‍👦 Family needs to know you're working 👨‍👧
&lt;/li&gt;
&lt;li&gt;  ☕️ Take breaks &amp;amp; stand up.
&lt;/li&gt;
&lt;li&gt;  🤝 Be social 
&lt;/li&gt;
&lt;li&gt;  👤Stay off of Social Media
&lt;/li&gt;
&lt;li&gt;  🤧 Being sick is allowed
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt; 📝 Context&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  Getting context is key for success
&lt;/li&gt;
&lt;li&gt;  Sharing context
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;p&gt; 🤙Emojis 🙌&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;
&lt;p&gt; 👥 Remote meetings and async communication&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  📆 Daily stand ups and YTB (Yesterday, Today, Blockers)
&lt;/li&gt;
&lt;li&gt;  🗓 Scheduled meetings
&lt;/li&gt;
&lt;li&gt;  👋 Impromptu meetings
&lt;/li&gt;
&lt;li&gt;  🖖 Be responsive
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;  📷 Zoom Zoom - Being super productive with Zoom&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  ⏱Be on time to meetings&lt;/li&gt;
&lt;li&gt;  🧩Use the zoom slack integration - fastest way to join/start meetings&lt;/li&gt;
&lt;li&gt;  🐭Screen sharing and taking control
&lt;/li&gt;
&lt;li&gt;  🔇Make sure to mute yourself if you're not talking
&lt;/li&gt;
&lt;li&gt;  🌐 No zoom on VPN (if possible)
&lt;/li&gt;
&lt;li&gt;  🌇 Virtual backgrounds
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt; 👖Slacking like a pro&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  💬 Threads
&lt;/li&gt;
&lt;li&gt;  🧩 Slack Apps
&lt;/li&gt;
&lt;li&gt;  👤 Personalize
&lt;/li&gt;
&lt;li&gt;  👀 React-jis
&lt;/li&gt;
&lt;li&gt;  ⌨️ Shortcuts
&lt;/li&gt;
&lt;li&gt;  ⏲ Reminders
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;p&gt;More resources&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h1&gt;
  
  
  Intro (TL;DR - why should you read this doc) 
&lt;/h1&gt;

&lt;p&gt;Some of you may need to work from home from time to time. Given COVID-19&lt;br&gt;
you may even need to work from home for multiple days in a row. Feel&lt;br&gt;
free to refer to this doc for productivity tips and tricks for working&lt;br&gt;
from home.&lt;/p&gt;

&lt;h1&gt;
  
  
  Can WFH be productive at all?
&lt;/h1&gt;

&lt;p&gt;A lot of great companies you know have decided that remote/distributed&lt;br&gt;
will work for them. Automatic (Wordpress), Invision, Netlify, Netlifx&lt;br&gt;
and a bunch of other companies see remote as not only productive, but&lt;br&gt;
part of their culture. It does require adjustments. Think of it as like&lt;br&gt;
training muscle memory. You're used to the office environment, being&lt;br&gt;
able to tap someone on the shoulder and ask your question. Being remote&lt;br&gt;
might feel like this is going away, but as with all muscle memory,&lt;br&gt;
practice makes perfect and according to the companies above, it can be&lt;br&gt;
even really productive.&lt;/p&gt;

&lt;p&gt;As a 100% remote employee for the past 8 months, including leading a&lt;br&gt;
distributed team for 3 months, I have researched and collected some&lt;br&gt;
remote work tips and tricks. They might not all be for everyone, but&lt;br&gt;
hopefully some will help you to be more productive.&lt;/p&gt;

&lt;h1&gt;
  
  
  💼 Work/Life balance and separation 🏡
&lt;/h1&gt;

&lt;h2&gt;
  
  
  👨‍🏫 Find your work place. If you have a spare room, make that your office. Do not work from the couch! 
&lt;/h2&gt;

&lt;p&gt;This is one of the main things about being remote that people quickly&lt;br&gt;
notice. You have to force yourself to separate work life and home life.&lt;br&gt;
Otherwise it's easy to get worn down, as there always might be one more&lt;br&gt;
slack to answer, one more bug to fix, and you're already home, so&lt;br&gt;
there's no commute right?&lt;/p&gt;

&lt;h2&gt;
  
  
  🌅 Getting up in the morning, and getting ready (AKA put some pants on!)
&lt;/h2&gt;

&lt;p&gt;Going through your regular work day morning regimen, get up with an&lt;br&gt;
alarm clock, shower/wash your face, wear "outside" clothes, even putting&lt;br&gt;
on shoes, are good ways of &lt;strong&gt;"preparing to go to work"&lt;/strong&gt; even if you&lt;br&gt;
then sit down at your kitchen table and work. It helps separate work&lt;br&gt;
life and home life.\&lt;br&gt;
Try to schedule your working day with clear star/stop hours and try to&lt;br&gt;
commit to them.&lt;/p&gt;

&lt;h2&gt;
  
  
  👨‍👩‍👦 Family needs to know you're working 👨‍👧
&lt;/h2&gt;

&lt;p&gt;Family, if they are home also, should know that you're &lt;strong&gt;working&lt;/strong&gt; and&lt;br&gt;
you need to set expectations as such. Having that chat with a family&lt;br&gt;
member who's also home might be difficult, but in the long run it will&lt;br&gt;
make it much more productive for you. Same goes with kids. If you are&lt;br&gt;
able, try to set clear boundaries for work/family ahead of time. Maybe&lt;br&gt;
postpone things till your next break, or put on headphones and explain&lt;br&gt;
that you're busy while in headphones.&lt;/p&gt;

&lt;p&gt;From the other hand, you can spend more time with your family during the&lt;br&gt;
time you would have otherwise spend in traffic/commuting. Use that time&lt;br&gt;
with them.&lt;/p&gt;

&lt;h2&gt;
  
  
  ☕️ Take breaks &amp;amp; stand up.
&lt;/h2&gt;

&lt;p&gt;Taking breaks is also very important, it's easy to be forced to a break&lt;br&gt;
in the office (maybe too easy) but when you're connected to slack, make&lt;br&gt;
sure you take breaks from time to time. Stretch your legs, go outside&lt;br&gt;
for a little bit and see the sun. You might not notice this in a while,&lt;br&gt;
but being in the office, you stand up more often than you do at home.&lt;br&gt;
Those things will prevent potential burn out.&lt;/p&gt;

&lt;h2&gt;
  
  
  🤝 Be social
&lt;/h2&gt;

&lt;p&gt;Don't forget that part. Join some slack channels for jokes, sharing&lt;br&gt;
pictures of dogs/cats or create a meme channel. Some companies even&lt;br&gt;
prompt you to have social zoom calls for 5-15 minutes that are not about&lt;br&gt;
work. This is a big part that's missing when WFH and it's an important&lt;br&gt;
part of being in the office so don't forget it.&lt;/p&gt;

&lt;h2&gt;
  
  
  👤Stay off of Social Media
&lt;/h2&gt;

&lt;p&gt;Do this on your breaks, but make it harder on yourself to get sucked&lt;br&gt;
into social media. Install different browsers for personal/work life and&lt;br&gt;
log out from everything social on your work browser. (I suggest Brave&lt;br&gt;
for personal browsing)\&lt;br&gt;
Install extensions like &lt;a href="https://chrome.google.com/webstore/detail/go-fucking-work/hibmkkpfegfiinilnlabbfnjcopdiiig?hl=en" rel="noopener noreferrer"&gt;Go F***ng&lt;br&gt;
Work&lt;/a&gt;&lt;br&gt;
which will limit your social media time and for you to go to work.&lt;/p&gt;

&lt;h2&gt;
  
  
  🤧 Being sick is allowed
&lt;/h2&gt;

&lt;p&gt;Even though you are working from home, don't be ashamed of taking a sick&lt;br&gt;
day if you're not feeling well. This will feel weird, but do the same thing&lt;br&gt;
you would do when working from the office, take a sick day, wear your&lt;br&gt;
PJs and let folks know (with a slack status or actively in YTB) that&lt;br&gt;
you're sick!&lt;/p&gt;

&lt;h1&gt;
  
  
  📝 Context! 
&lt;/h1&gt;

&lt;h2&gt;
  
  
  Getting context is key for success!
&lt;/h2&gt;

&lt;p&gt;Knowing what's going on is especially hard if you're working remotely.&lt;br&gt;
While meeting in the kitchen over coffee, you can overhear things, when&lt;br&gt;
you're remote, you will see that this type of communication is really&lt;br&gt;
hard. Some of this can/should be solved on the culture level, but&lt;br&gt;
personally, don't be afraid to reach out and actively admit that you&lt;br&gt;
don't have context. It won't reflect badly on you when you're actively&lt;br&gt;
asking because you're not sure.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sharing context 
&lt;/h2&gt;

&lt;p&gt;You will need to take on the responsibility of identifying water-cooler&lt;br&gt;
context and pulling it down into written form. It's a muscle worth&lt;br&gt;
exercising; you'll find that it's equally effective at identifying and&lt;br&gt;
bridging silos of communication (i.e. private slack channels) as it is&lt;br&gt;
connecting remote employees with office culture.&lt;/p&gt;

&lt;h1&gt;
  
  
  🤙Emojis 🙌
&lt;/h1&gt;

&lt;p&gt;Written text isn't great for conveying intent the way face-to-face&lt;br&gt;
conversations are. Facial expressions, tone, cadence, and body language&lt;br&gt;
contribute to how we interpret the intent behind words. The same set of&lt;br&gt;
words can be interpreted as condescending in one context and well&lt;br&gt;
intentioned in another.&lt;/p&gt;

&lt;p&gt;Emojis may not seem professional but they help convey the intent behind&lt;br&gt;
our words. Putting a smiley face at the end of a slack message cues the&lt;br&gt;
reader into how our words should be interpreted. Consistently giving&lt;br&gt;
these cues can be the difference between being seen as a brilliant jerk&lt;br&gt;
and a stunning colleague.&lt;/p&gt;

&lt;h1&gt;
  
  
  👥 Remote meetings and async communication
&lt;/h1&gt;

&lt;p&gt;There will most likely be fewer meetings (a good thing!) if you're&lt;br&gt;
working mostly remote, but there will still be some. Taking notes and&lt;br&gt;
sharing them with the team is a great idea in such a case.&lt;/p&gt;

&lt;h2&gt;
  
  
  📆 Daily stand ups and YTB (Yesterday, Today, Blockers)
&lt;/h2&gt;

&lt;p&gt;We've been slacking YTB in our team chat for a while, and while that was&lt;br&gt;
good, adding a 5 minute meeting to go over those posted item really&lt;br&gt;
helps with questions on each item. Put that stand up meeting on the&lt;br&gt;
calendar, make sure you post YTBs in the slack channel before, and this&lt;br&gt;
meeting will be a brief one. You'll also get to see your teams faces&lt;br&gt;
every day, which connects you to the team.&lt;/p&gt;

&lt;h2&gt;
  
  
  🗓Scheduled meetings
&lt;/h2&gt;

&lt;p&gt;Your scheduled meetings will proceed as usual most likely, be sure to&lt;br&gt;
prepare 5 minutes before, check your setup, microphone, webcam and&lt;br&gt;
internet connection so that the meeting will start on time.&lt;/p&gt;

&lt;h2&gt;
  
  
  👋 Impromptu meetings
&lt;/h2&gt;

&lt;p&gt;Sometime slacking is not going to cover it and you need to talk to the&lt;br&gt;
person. Don't be shy inviting them to a quick zoom! (see below for zoom&lt;br&gt;
tips)&lt;/p&gt;

&lt;h2&gt;
  
  
  🖖 Be responsive! 
&lt;/h2&gt;

&lt;p&gt;When not in meeting, try to be very responsive during your work hours.&lt;br&gt;
Even saying "hey, I'll get back to you in 15" is helpful than not&lt;br&gt;
answering for a few hours.&lt;/p&gt;

&lt;h1&gt;
  
  
  📷 Zoom Zoom - Being super productive with Zoom
&lt;/h1&gt;

&lt;h2&gt;
  
  
  ⏱Be on time to meetings! 
&lt;/h2&gt;

&lt;p&gt;Punctuality is especially important when you don't have to "walk" to the&lt;br&gt;
meeting room!&lt;/p&gt;

&lt;p&gt;Prepare for your zoom meetings in advance by checking your&lt;br&gt;
microphone/webcam and internet connection are all in order. This will&lt;br&gt;
make sure the meetings start on time and are as productive as possible&lt;/p&gt;

&lt;h2&gt;
  
  
  🧩Use the zoom slack integration - fastest way to join/start meetings
&lt;/h2&gt;

&lt;p&gt;Type /zoom into slack and it will start a meeting with the channel /&lt;br&gt;
person you're chatting with&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fi%2F5lt5493jdyhqqijygc0g.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fi%2F5lt5493jdyhqqijygc0g.png" alt="Alt Text" width="800" height="397"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  🐭Screen sharing and taking control
&lt;/h2&gt;

&lt;p&gt;Zoom allows for easy screen sharing. Make sure no sensitive info is on&lt;br&gt;
the screen before you share.&lt;/p&gt;

&lt;p&gt;Zoom also allows you to share a specific window of an app, or the whole&lt;br&gt;
screen itself. Depending on your situation, choose one or another&lt;br&gt;
accordingly.&lt;/p&gt;

&lt;p&gt;Sometimes, especially when speaking to IT or doing a code review,&lt;br&gt;
sharing control of your screen is really helpful! &lt;a href="https://canvas.du.edu/courses/79407/pages/sharing-mouse-control-in-a-team-meeting" rel="noopener noreferrer"&gt;See how to&lt;br&gt;
here&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  🔇Make sure to mute yourself if you're not talking
&lt;/h2&gt;

&lt;p&gt;While zoom is pretty great, intermittent internet issues and bad&lt;br&gt;
microphones might make it hard for folks to hear other folks. Mute&lt;br&gt;
yourself when you're not talking please.&lt;/p&gt;

&lt;h2&gt;
  
  
  🌐 No zoom on VPN (if possible) 
&lt;/h2&gt;

&lt;p&gt;Try to get off VPN if you can before a meeting, this will speed up your&lt;br&gt;
zoom connection and leave some bandwidth for other folks on VPN.&lt;/p&gt;

&lt;h2&gt;
  
  
  🌇 Virtual backgrounds
&lt;/h2&gt;

&lt;p&gt;If you'd like to express yourself with a nice background, hide the mess&lt;br&gt;
you have behind you or just generally surprise folks, Zoom allows you&lt;br&gt;
set up a virtual background easily!&lt;/p&gt;

&lt;p&gt;&lt;a href="https://youtu.be/YL736HaaJCk" rel="noopener noreferrer"&gt;https://youtu.be/YL736HaaJCk&lt;/a&gt;&lt;/p&gt;

&lt;h1&gt;
  
  
  👖Slacking like a pro
&lt;/h1&gt;

&lt;h2&gt;
  
  
  💬 Threads
&lt;/h2&gt;

&lt;p&gt;Use threads to remove clutter on public channels effectively. If a&lt;br&gt;
conversion required the attention of everyone in the channel, have it in&lt;br&gt;
the open. But if you are thinking "maybe I should take this private",&lt;br&gt;
try to respond a thread. This will remove the noise from the channel,&lt;br&gt;
but also will leave information accessible for the rest of the team.&lt;/p&gt;

&lt;p&gt;In channels like #general and #random please use the thread feature to&lt;br&gt;
reply to most message, as otherwise most of the company will receive an&lt;br&gt;
unread notification about your witty comment &lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/media%2Fimage3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/media%2Fimage3.png" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;If the whole channel still needs to know your response, hit the little&lt;br&gt;
"send to channel" checkbox in the thread, this way it will keep the&lt;br&gt;
threads structure and present your message to the larger channel.&lt;/p&gt;

&lt;h2&gt;
  
  
  🧩 Slack Apps
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fi%2Fi0lvnf4vsjl2wu6i9mj7.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fi%2Fi0lvnf4vsjl2wu6i9mj7.png" alt="Alt Text" width="800" height="567"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Add Google calendar integration to slack, it will change your status, so&lt;br&gt;
folks will see that you're in a meeting (&lt;a href="https://slack.com/help/articles/206329808-Google-Calendar-for-Slack" rel="noopener noreferrer"&gt;see&lt;br&gt;
how&lt;/a&gt;)&lt;/p&gt;

&lt;h2&gt;
  
  
  👤Personalize 
&lt;/h2&gt;

&lt;p&gt;Make sure you have a picture of your face on slack, folks need to know&lt;br&gt;
who they are talking to. Having a picture of your cat, or a caricature&lt;br&gt;
of your face might be funny but it will make it hard for newer folks to&lt;br&gt;
connect with you and know who you are.&lt;/p&gt;

&lt;h2&gt;
  
  
  👀React-jis 
&lt;/h2&gt;

&lt;p&gt;Slack has react emojis, and it's really helpful to use them as a way to&lt;br&gt;
convey "I've read this" or "I also think like this" , Slack has a full&lt;br&gt;
breakdown of how to use their react emojis&lt;br&gt;
&lt;a href="https://slack.com/help/articles/206870317-Use-emoji-reactions" rel="noopener noreferrer"&gt;here&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fi%2F43ayqos4pnjy5c74hu9c.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fi%2F43ayqos4pnjy5c74hu9c.png" alt="Alt Text" width="799" height="238"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  ⌨️Shortcuts
&lt;/h2&gt;

&lt;p&gt;Slack is much more productive when you learn the basic shortcuts. Hit&lt;br&gt;
&lt;strong&gt;cmd/ctrl + /&lt;/strong&gt; in slack to see all of them.&lt;/p&gt;

&lt;p&gt;The main one is &lt;strong&gt;cmd/ctrl+K&lt;/strong&gt; which will open the search and give you a&lt;br&gt;
fast way to navigate to a person DM, channel, search and everything in&lt;br&gt;
slack!&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fi%2F3ontnpkk8d50wcqvl2jx.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fi%2F3ontnpkk8d50wcqvl2jx.png" alt="Alt Text" width="800" height="568"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  ⏲Reminders 
&lt;/h2&gt;

&lt;p&gt;Slack has a remind me later feature, use this to your advantage. Slack&lt;br&gt;
messages tend to disappear, especially if you read them and then read&lt;br&gt;
something else. If there's an action item, or you need to respond&lt;br&gt;
someone after your meeting, create a reminder&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fi%2Fqy3cr8rjdn2fr55hjajt.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fi%2Fqy3cr8rjdn2fr55hjajt.png" alt="Alt Text" width="800" height="418"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h1&gt;
  
  
  Even more resources
&lt;/h1&gt;

&lt;p&gt;I hope this was helpful, and if you need some more resources, checkout this collection by @notionHQ -&amp;gt; &lt;a href="https://www.notion.so/Remote-work-wiki-1b21ef5501714fffa9f5c5c25677371f" rel="noopener noreferrer"&gt;https://www.notion.so/Remote-work-wiki-1b21ef5501714fffa9f5c5c25677371f&lt;/a&gt;&lt;/p&gt;

</description>
      <category>remotework</category>
      <category>wfh</category>
      <category>productivity</category>
      <category>zoom</category>
    </item>
    <item>
      <title>Fixing Pycharm remote Interpreter with Vagrant and macOS Catalina</title>
      <dc:creator>Alex Volkov</dc:creator>
      <pubDate>Tue, 07 Jan 2020 01:51:50 +0000</pubDate>
      <link>https://dev.to/altryne/fixing-pycharm-remote-interpreter-with-vagrant-and-macos-catalina-27k5</link>
      <guid>https://dev.to/altryne/fixing-pycharm-remote-interpreter-with-vagrant-and-macos-catalina-27k5</guid>
      <description>&lt;h2&gt;
  
  
  Fookin Apple
&lt;/h2&gt;

&lt;p&gt;Mac os Catalia was released on October 2019 (last decade!) and it introduced a bunch of security fixes. &lt;/p&gt;

&lt;p&gt;Full disk access was introduced as a permission that you need to allow tools to access folders in your /Users directory. &lt;/p&gt;

&lt;p&gt;This broke a lot of things, just google "full disk access" and you see pages on top of pages of articles all explaining how to add something to the Full Disk Access. &lt;/p&gt;

&lt;h2&gt;
  
  
  Our issue
&lt;/h2&gt;

&lt;p&gt;In our case, vagrant NFS mounts stopped working. And we can't run our Vagrant without those mounts. &lt;/p&gt;

&lt;p&gt;We've added a temp. fix, and Vagrant 2.6 actually released a fix of their own, which adds the exported paths in the longer form, which doesn't need full disk access to work, &lt;br&gt;
Old form : /Users/username/folder&lt;br&gt;
Longer form: /System/Volumes/Data/Users/username/folder&lt;/p&gt;

&lt;p&gt;So now that that's fixed, our vagrant is happy and boots fine and everything is ok. &lt;/p&gt;

&lt;p&gt;(There's another fix, adding &lt;code&gt;/sbin/nfsd&lt;/code&gt; to Full Disk Access permissions which kinda works as well)&lt;/p&gt;
&lt;h2&gt;
  
  
  PyCharm issues
&lt;/h2&gt;

&lt;p&gt;We then had a different issue, where in devs who updated to Catalina, ran their pyCharm and set up a remote interpreter, suddenly couldn't run their debug configurations through pyCharm anymore. &lt;/p&gt;

&lt;p&gt;It showed the following error:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ssh://vagrant@127.0.0.1:2222/folder/bin/python -u /Users/usernname/folder/utils/run_something.py
zsh:cd:1: no such file or directory: /Users/username/folder/utils
/folder/remove_venv/bin/python: can't open file 'Users/usernname/folder/utils/run_something.py': [Errno 2] No such file or directory
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Basically PyCharm was complaining that it can't create a path mapping between a mac folder to a mounted vagrant NFS folder. &lt;/p&gt;

&lt;p&gt;This is due to this magical thing that pyCharm does when you create a python remote interpreter (I assume the same is true for other languages and Jetbrain IDEs as well)&lt;br&gt;
It reads your vagrant file, and tries to extract the paths you are about to mount inside your vagrant. &lt;/p&gt;

&lt;p&gt;The problem though, is pyCharm doesn't know about this Catalina long form /Systems/Volume... path, and the fact that Vagrant kinda hacks it together, so the path mappings don't work&lt;/p&gt;

&lt;h1&gt;
  
  
  Solution
&lt;/h1&gt;

&lt;p&gt;Until jetBrains learn to deal with this problem (track the issue &lt;a href="https://youtrack.jetbrains.com/issue/WI-49183" rel="noopener noreferrer"&gt;here&lt;/a&gt;, vote on it, make some noise if you have this) &lt;br&gt;
We are forces with this solution: &lt;/p&gt;

&lt;p&gt;Add a path mapping in your debug configuration to the mounted NFS folder inside vagrant manually. &lt;/p&gt;

&lt;h3&gt;
  
  
  Make it a little bit better
&lt;/h3&gt;

&lt;p&gt;If you are used to have a lot of debug configurations, let's say you create one for each unit test you have, you can edit the template for the debug configuration and add the path mapping there, it will then exist for all new configurations created from that template. &lt;/p&gt;

&lt;p&gt;Just expand the template cogwheel and edit the template of the config&lt;/p&gt;

</description>
      <category>vagrant</category>
      <category>pycharm</category>
      <category>macos</category>
      <category>catalina</category>
    </item>
  </channel>
</rss>
