Google ships Gemini 3.8 Flash — and a cyber twin
On September 2, Google released Gemini 3.8 Flash — its third Flash-tier model in six weeks — alongside a restricted Gemini 3.8 Flash Cyber variant. The workhorse keeps the same introductory price as 3.7 Flash: $0.75 / $3.75 per million tokens through December 31, 2026 (doubling to $1.50 / $7.50 on January 1, 2027).
Google calls 3.8 Flash its "most intelligent workhorse," with real gains on software engineering and long agentic loops: ~71–73.7% on DeepSWE v1.1 and 90.8% on Terminal-Bench 2.1 (up from 81.6% on 3.7 Flash), ahead of GPT-5.6 Terra (87.4%) and Claude Sonnet 5 (80.4%) on that benchmark. The Cyber variant targets vulnerability discovery and automated patching, gated behind the new Fairwind Program for governments and trusted defenders, with a reported 70%+ real-world vuln-discovery rate across 20 languages.
Why it matters: Google is betting that cheap, fast, constantly-iterated Flash models win developer mindshare more reliably than occasional flagship leaps.
Meta unveils Muse Spark 1.3 — its biggest coding jump yet
Also on September 2, Meta released Muse Spark 1.3 into Muse Code and the Meta Model API. Zuckerberg framed it as Meta's largest improvement to date on coding and agentic work, with frontier-class performance "almost too cheap to meter." Headline numbers: 88.8% on Terminal-Bench 2.1 (matching GPT-5.6), 75.4% on DeepSWE v1.1, and a standout 98.5% on MRCR 256K–512K long-context understanding.
Meta also trailed two unshipped items: a larger unnamed model and — notably — open weights for the Muse Spark line. If that lands, it would be Meta's first open-weight frontier-class release of this generation.
Why it matters: the coding/agentic race is now the primary battleground, and Meta is signaling it intends to compete on both proprietary API and open weights.
Alibaba's Qwen3.8-Max-0902 tops the WebDev leaderboard
Alibaba refreshed its flagship Qwen3.8-Max to the -0902 build, post-trained specifically on Coding & Cowork. It carries a 2.4T-parameter headline and a 1M-token context window, priced at $2 / $6 per million tokens.
On Code Arena's WebDev leaderboard, the new build sits first at 1,691 — a 22-point move from the prior 1,669 — ahead of Claude Opus 5 Max (1,687) and Kimi K3 Max. One caveat: Arena labels the score preliminary with a ±19 spread, so the lead over Opus 5 Max is within noise rather than a settled separation.
Why it matters: Chinese labs keep pushing the open/affordable frontier, and Qwen's post-training focus on real enterprise coding tasks mirrors the exact same strategic pivot Google and Meta just made.
A quiet but busy 48 hours at the frontier: three labs shipped meaningful model updates, and the theme is unmistakable — coding and long-horizon agentic work, priced to move. More daily AI briefings at AI Nexus Daily.
Top comments (0)