DEV Community

HIROKI II
HIROKI II

Posted on

AI Daily Digest β€” August 4, 2026: Qwen 3.8-Max Ships, Mythos Cracks HAWK, Fireworks Hits $1B ARR

πŸ€–πŸ’» AI Daily Digest β€” August 4, 2026


Alibaba's Qwen 3.8-Max Codes Autonomously for 16 Days β€” and Open-Sources the Result

Alibaba released Qwen3.8-Max on August 3, its largest flagship model to date: 2.4 trillion total parameters with only 95 billion activated per token, using a sparse mixture-of-experts architecture paired with hybrid attention. It carries a 1M-token context window, native vision, and a ranking profile that is getting hard to dismiss β€” fifth on Text Arena, second on Vision Arena, fourth on CodeArena, and 93.0 on PaperBench against Fable 5's 88.8. The API is live on Alibaba Cloud Model Studio, and the weights are scheduled to open next week alongside Qwen3.8-27B.

The flagship demo is the one people keep quoting. Tasked with building a self-evolving agent framework from an empty folder, the model ran an engineering loop on its own for about 16 days β€” generating code, testing, previewing, reading logs, folding in user feedback and community practices β€” and delivered "oh-my-cli," a working self-evolving agent harness, open-sourced on GitHub with its full operation history public. The team also reports a 125-hour autonomous research run that reproduced published papers and pushed AIME24 up another 2.7 points, and a quant workflow where the model coordinated roughly 330 sub-agents through about 6,000 factor backtests.

The pricing makes the release read like a coordinated squeeze on closed flagships. Domestically Qwen3.8 runs at Β₯12 per million input tokens and Β₯36 output, with cache hits at Β₯1.5; internationally the prices land at roughly 40% and 24% of Opus 5. Alibaba also shipped QwenWork the same day, an all-in-one workplace agent product that surfaces the model through a PC client and DingTalk. The open-weight race keeps compressing what a closed lab can charge for comparable capability.

β€” Alibaba Cloud Β· Qwen Β· The Paper

πŸ”— Alibaba Cloud β€” Qwen3.8-Max Launch Β· Qwen Β· The Paper Coverage


OpenAI's GPT-Live Goes Full-Duplex: Listen and Speak at the Same Time

OpenAI published an engineering deep-dive on August 3 explaining how GPT-Live, its third-generation voice system, moved from turn-based to streaming, full-duplex interaction. The model listens and speaks simultaneously; the classic "turn detector" that decided when the assistant could start talking is gone. Audio streams directly into the model as a continuous signal, and when deeper reasoning or tool use is needed, GPT-Live hands off to frontier models like GPT-5.5 in the background without interrupting the live conversation.

The underlying engineering is where the post earns its length. The inference engine was rewritten in Go, replacing a Python asyncio pipeline β€” frame delivery p95 now matches the old p50. Transport sits on WebRTC with clock-drift handling that stretches or compresses audio to stay real-time. A new WARP protocol (WebRTC Abridged Roundtrip Protocol) cuts media session startup from six network round trips to one, and Instant Connect pre-negotiates SDP parameters so a session can start from a single UDP packet. The stack also supports stateful inference: warm replacement of model instances, prefill with context, and context compaction with no media interruption.

This architecture now powers computer control from the ChatGPT desktop app and agent coordination inside ChatGPT, and OpenAI says it will underpin the upcoming GPT-Live API. The practical read for builders: realtime voice is becoming a delegation surface β€” the cheap, fast model carries the conversation while the expensive reasoning happens off the critical path, and the user never notices the handoff.

β€” OpenAI

πŸ”— OpenAI β€” Continuous Voice Interaction with GPT-Live


Claude Mythos Finds Mathematical Flaws in HAWK and Reduced-Round AES

Anthropic published on July 28 the first results of Claude Mythos Preview attacking the math inside cryptographic algorithms, not just their implementations. Against HAWK, a NIST post-quantum signature candidate that survived two rounds of expert review over two years, Mythos found a nontrivial automorphism in the lattice structure and cut the scheme's effective key strength in half β€” in about 60 hours of work. Against a reduced seven-round version of AES, it invented a shortcut it named the MΓΆbius Bridge that eliminates a 256-way lookup, making the strongest known theoretical attack 200 to 800 times faster.

Neither result touches production systems. HAWK is not deployed, and the AES attack operates in a chosen-plaintext model that assumes around 2^105 chosen plaintexts β€” Anthropic itself calls it completely impractical. The HAWK team has since withdrawn the scheme from the NIST process. Each result cost roughly $100,000 in API spend to develop, mostly run autonomously: one researcher scaffolded a loop that let Mythos explore the AES attack over three days and roughly a billion output tokens.

The most interesting part of the post is what Anthropic admits about verification. Mythos found the AES attack in a week; two researchers then spent nearly a month convincing themselves it was correct, and the company says most of its research time lately has gone into verifying model output. To keep the field measurable, Anthropic built CryptanalysisBench with ETH Zurich, Tel Aviv University and the University of Haifa. The takeaway for anyone watching safety research: discovery is no longer the bottleneck β€” checking the discovery is.

β€” Anthropic Β· CyberScoop

πŸ”— Anthropic Research β€” Discovering Cryptographic Weaknesses Β· CyberScoop


NVIDIA Backs Ilya Sutskever's SSI With Vera Rubin Compute β€” an Order of Magnitude

Safe Superintelligence Inc. and NVIDIA announced a long-term strategic partnership on July 27. NVIDIA made an investment β€” Bloomberg pegs it around $5 billion, neither company confirms a figure β€” and granted SSI access to the next-generation Vera Rubin platform, which SSI says will multiply its compute by an order of magnitude within 12 months. The two companies will also collaborate on advancing NVIDIA's current and future compute platforms, using SSI's research insights.

The unusual part is what NVIDIA got in exchange. Jensen Huang said the company decided to enter the partnership "after obtaining rare access into the company's closely guarded research." SSI has spent two years in near silence since Ilya Sutskever and Daniel Levy founded it in 2024, with roughly $3 billion raised at a reported $32 billion valuation from a16z, DST Global, Greenoaks and Sequoia. Sutskever's comment was characteristically plain: "We have research that is worthy of scaling up, and having access to a big NVIDIA computer will let us do so."

NVIDIA's role in the industry keeps shifting with deals like this. It is no longer just selling chips to whoever shows up β€” it is choosing which labs get early access to the next platform, and investing in the ones it believes in. For a lab like SSI whose entire pitch is a single long-horizon bet on aligned superintelligence, the deal converts the scarcest resource in AI β€” next-generation compute β€” into a strategic asset rather than a line item.

β€” NVIDIA Β· SSI Β· TechCrunch

πŸ”— NVIDIA Newsroom β€” SSI Partnership Β· SSI Β· TechCrunch


Fireworks Crosses $1B ARR, Raises $1.505B at a $17.5B Valuation

Fireworks AI announced a $1.505 billion Series D at a $17.5 billion valuation, led by Atreides Management, Index Ventures and TCV, with participation from NVIDIA, Lightspeed, Bessemer, Menlo Ventures and others. The milestone figures attached to the round: annualized revenue run rate past $1 billion, five times the previous year, and more than 40 trillion tokens served per day, up from 15 trillion β€” with 95% of that volume coming from specialized, fine-tuned models rather than general API access.

The company, founded in 2022 by Lin Qiao and co-founders from Meta and Google Brain, runs an inference cloud where enterprises fine-tune open models on their own data and serve them in production. Its customer list runs from Uber and Shopify to GitLab, MongoDB, and the legal and coding tooling built on it β€” Harvey and Cursor. Qiao's cost framing for why enterprises switch: "Our cost compared with the equivalent-quality closed model is five to 10 times cheaper."

The round is the largest inference-layer funding in AI infrastructure history, and it is a direct rebuttal to the persistent bear case that optimization-layer companies get commoditized as model prices fall. The counter-evidence, on the numbers above: cheaper models pulled more enterprise workloads onto AI infrastructure, which made serving them well more valuable, not less. For the broader market, Fireworks' $1B ARR at 95% specialized volume is the clearest signal yet that the money in AI infrastructure is consolidating around owning and serving models β€” not renting frontier intelligence by the token.

β€” Fireworks AI Β· Yahoo Finance

πŸ”— Fireworks β€” Series D Announcement Β· Yahoo Finance


Mitsubishi Motors to Put UTokyo-Spin-Off Humanoids in Its Factories From 2027

Mitsubishi Motors signed a basic agreement with Highlanders, a robotics startup spun off from the University of Tokyo, to develop and deploy humanoid robots for manufacturing. The plan: humanoids first enter Mitsubishi Motors' own factories to accumulate real-world operational data and know-how, with production of the robots scheduled to begin at Mitsubishi's Kyoto Plant from 2027. The automaker brings manufacturing expertise; Highlanders contributes the robotics and AI stack, built around physical AI applications.

The stated driver is Japan's demographic reality. Mitsubishi positions the partnership as a practical response to labor shortages in industrial workforces, preserving and transferring skilled manufacturing techniques through robotic systems β€” and potentially opening a new business line in the process. It is one of the more concrete steps by a traditional Japanese automaker into volume-oriented humanoid robotics, following similar explorations by global peers who are also treating factories as the first real deployment surface for humanoids.

The agreement says nothing yet about robot specifications, production volumes, or which factory tasks come first β€” those are expected as the collaboration moves toward the 2027 target. What is already notable is the pattern: automakers stopped treating humanoids as research demos and started treating them as workforce infrastructure, with the factory floor as the proving ground where the economics can actually be calculated.

β€” Mitsubishi Motors Β· Highlanders Β· Humanoid Press

πŸ”— Mitsubishi Motors Β· Humanoid Press Β· Highlanders


avatarin Puts a 24/7 Voice Shopping Agent in Yamada Denki's Online Store

OpenAI published a customer story on July 30 about avatarin, an AI customer-service company spun out of ANA Holdings, and its "Kurashi-Marugoto AI Agent" for Yamada Denki's online store. Built on OpenAI's GPT-Realtime, the agent runs around the clock in multiple languages, answers voice questions about products, and β€” the design choice that stands out β€” asks follow-up questions rather than waiting for instructions. In a two-week public campaign, roughly 30,000 people used it and 92% of post-use survey responses were positive.

The architecture follows a pattern worth copying: RAG keeps product answers grounded in accurate catalog data while GPT-Realtime keeps the conversation responsive; Yamada Denki's sales knowledge is encoded into conversation flows, so the agent adapts when a shopper changes requirements or goes off topic; and the proactive-question design moves the interaction from Q&A toward guided discovery. avatarin CEO Akira Fukabori's framing: "Customers do not want a chatbot. They want intelligence. 'I need a refrigerator for a family of four, but my kitchen is small. Which one should I choose?'"

The interesting signal is where this sits in OpenAI's own narrative β€” an always-on, multilingual, voice-first agent that extends expert knowledge beyond store hours, with every conversation producing insight about what shoppers care about. For the agentic-commerce direction, the notable numbers are the ones nobody is talking about yet: what happens to the "retail interface" when the sales floor becomes a voice agent that never closes.

β€” OpenAI Β· avatarin Β· Yamada Denki

πŸ”— OpenAI β€” avatarin Case Study Β· avatarin Β· Dada3C Coverage

Top comments (0)