DEV Community

HIROKI II
HIROKI II

Posted on

AI Daily Digest — September 3, 2026: Broadcom's AI Record, Gemini 3.8 Flash, Cognition Hits $47B

Cover

Broadcom posts a record quarter — $16.7B in AI revenue — and the stock drops anyway

Broadcom reported fiscal Q3 2026 results on September 2, and the headline numbers were enormous: total revenue of $29.59 billion, up 86% year over year and slightly ahead of the ~$29.36 billion consensus, with non-GAAP EPS of $3.32 beating the ~$3.24 estimate. The engine is custom AI silicon — AI semiconductor revenue hit $16.7 billion, up 221% year over year and 54% sequentially, ahead of the roughly $15.9–16.0 billion analysts modeled. The semiconductor solutions segment contributed $20.8 billion (+127%), infrastructure software $8.8 billion (+29%), and free cash flow came in at $13.7 billion, 46% of revenue. CEO Hock Tan's forward number was even bigger: Q4 AI semiconductor revenue guided to $21.7 billion, up 236% year over year.

So why did the stock fall about 5% in after-hours trading, to roughly $348? Because total Q4 revenue guidance of approximately $34.8 billion came in below the $35.0–35.05 billion consensus — a slim miss in a market that had already priced in perfection. The backdrop adds pressure: the 10-year Treasury near 4.80%, a stock up only ~1% year-to-date against a 64% gain in the PHLX Semiconductor Index, and structural questions about Broadcom's share of hyperscaler custom-chip spending after Marvell's Google deal (worth up to $120 billion through fiscal 2033) and Broadcom's own July manufacturing pact with Samsung worth over $200 billion. My read: this is the purest signal yet that the AI-capex trade has moved from "beat or bust" to "beat, raise, and still get sold" — the revenue is real and accelerating, but the market is now pricing the custom-ASIC competitive race, not just the quarter. Stanley Druckenmiller had already exited his entire position in Q2, a call that looks prescient.

— Broadcom (official, SEC 8-K) · Benzinga · Unite.AI
🔗 Broadcom Q3 FY2026 results (SEC 8-K) · Benzinga on the beat and the slide · Unite.AI on the record quarter

Google DeepMind ships Gemini 3.8 Flash — and a defender-only 3.8 Flash Cyber

Google DeepMind released Gemini 3.8 Flash on September 2, its third Flash model in six weeks (3.6 Flash on July 21, 3.7 Flash on August 13, now this). Pricing stays at the 3.7 Flash introductory rate of $0.75 per million input tokens and $3.75 per million output tokens through December 31, 2026, before doubling to $1.50/$7.50; the context window remains 1M tokens with 64K output. Google's positioning is that 3.8 Flash "works harder" — on hard tasks it executes extra reasoning steps and iterates on tool calls, which shows up in the benchmarks: 54.9% on HLE-Verified, 73.7% on DeepSWE v1.1 (versus 74.0% for Claude Opus 5), wins on Vals Finance Agent V2 (61.4% vs 58.6%) and Harvey's Legal Agent Benchmark (10.0% vs 6.7%), and a 13-point jump to 56.5% on hard-mode BioMysteryBench. But the honest picture is more mixed: on 14 benchmarks against Opus 5, 3.8 Flash wins eight and loses five — and the losses are the hardest agent tests, like Terminal-bench 4.0 (19.1% vs 51.8%) and OSWorld-2.0 (59.0% vs 75.4%). Third-party testing also notes per-task cost is up roughly 40% versus 3.7 Flash at high effort levels, because the model simply burns more tokens.

The second release is the more interesting one. Gemini 3.8 Flash Cyber, a cybersecurity-specialized variant, is not a public API drop — it is distributed only through the new Fairwind Program to trusted defenders: government authorities, critical-infrastructure operators and software maintainers. Its numbers: CWE-Bench pass@1 of 47.2% on Collinear's external benchmark (a leading frontier model scores 47.8% at far higher cost), a 20-language internal vulnerability suite above 70% success, Chrome's security team reporting 2.6 times more correct patches than much larger commercial models, and Google's Cloud Vulnerability Research team finding a critical foundational vulnerability in under two hours where discovery usually takes months. My read: DeepMind, now led by Koray Kavukcuoglu after Demis Hassabis stepped down, is running a frequency war — cheap Flash models shipped every few weeks to own the developer workflow, while the cyber variant shows frontier labs converging on the same gated, defender-only distribution model OpenAI used for Astra. The 3.8 Flash model card's own caveats are worth keeping: knowledge cutoff March 2026, non-English safety evals down 5.4 points, and the intro price expires December 31.

— Google DeepMind (official) · Google (The Keyword) · IT之家
🔗 DeepMind: Introducing Gemini 3.8 Flash and 3.8 Flash Cyber · Google: The Keyword · IT之家 on the quiet launch

Cognition nears a $1B round at a $47B valuation, roughly doubling in three months

Bloomberg reported on September 2 that AI coding startup Cognition is closing a new round of about $1 billion at a valuation near $47 billion — up about 81% from the $26 billion valuation it set in late May — with investor demand of nearly $10 billion, roughly ten times what the company is seeking, and the final raise may exceed $1 billion. The driver is revenue: Cognition's annualized revenue has passed $900 million, nearly doubling from the $492 million it disclosed in late May, growth widely credited to the Windsurf acquisition (July 2025) combining Devin's autonomous engineering agent with an IDE and developer workflows. CEO Scott Wu has denied SpaceX acquisition rumors outright, saying the company is not for sale — notable because the reported talks follow SpaceX's $60 billion acquisition of rival Cursor, which closed in mid-August and reset the entire AI-coding market.

The context matters as much as the numbers. Cognition's customers reportedly include Citi, Goldman Sachs and the US Navy and Army; the company was founded in November 2023 and shipped Devin in March 2024 as one of the first "AI software engineers"; a year ago it was valued at $10.2 billion, meaning the price has grown more than fourfold in twelve months. My read: the $47B number is less about Devin's current revenue than about where the market thinks agentic engineering is heading after the Cursor deal — SpaceX paying $60 billion for a code editor legitimized the category, and every private AI-coding company is now being repriced against that ceiling. The uncomfortable part is the same one that dogs all of these rounds: a ~900% revenue-to-valuation ratio is a bet that agentic coding compounds for years, and with Anthropic's Fable line, OpenAI's Codex and Google's Flash models all circling the same developers, the competition is only getting more crowded.

— Bloomberg · The Edge Singapore · 新浪財經
🔗 The Edge Singapore (Bloomberg): Cognition set to raise ~$1B at $47B · AIbars on the round details · 新浪財經: Cognition 估值升至 470 億美元

NVIDIA puts $3.5B into MediaTek and opens NVLink Fusion to the custom-XPU crowd

NVIDIA and MediaTek announced a deepened partnership on August 31, with NVIDIA investing $3.5 billion in convertible bonds issued by MediaTek — reported as NVIDIA's largest direct investment outside the US, and roughly 90% of MediaTek's $3.9 billion overseas convertible offering (Alphabet also participated, with an undisclosed amount). MediaTek's stock jumped 9.94% on September 1. The strategic centerpiece is not the money but the platform: MediaTek will adopt NVIDIA's NVLink Fusion platform, offering hyperscalers, cloud service providers and frontier model developers a prevalidated path to build custom XPUs — their own AI accelerators — and plug them into NVIDIA NVLink-connected, rack-scale AI factories. That means custom ASICs no longer need to sit outside the NVIDIA ecosystem; NVLink Fusion supplies the interconnects, the NVLink-C2C links and the new NVHBM memory scheme.

The collaboration spans three fronts: AI infrastructure (the NVLink Fusion custom-XPU work), local AI computing (continued multi-generation RTX Spark and DGX Spark PC chips pairing NVIDIA GPUs with MediaTek SoCs), and automotive (AI-powered software-defined vehicles for the physical AI era). It extends a history that runs from Dimensity Auto integrating NVIDIA tech in 2023 through the GB10 Grace Blackwell chip for DGX Spark in 2025 to the RTX Spark launch this June. My read: this is NVIDIA's answer to the custom-ASIC threat from Broadcom and Marvell — the same week Marvell's Google deal (up to $120 billion through 2033) made headlines, NVIDIA is betting it can keep the rack, the network and the system layer even when the compute chip isn't a GPU. For MediaTek, the deal converts a potential competitor into a foundry-and-design partner with NVIDIA's full weight behind it. The market's verdict on day one — MediaTek +9.94%, NVIDIA down over 2% at the open — says investors read the shift as NVIDIA defending rather than attacking.

— NVIDIA (official newsroom) · 证券时报 · 21世紀經濟報道
🔗 NVIDIA Newsroom: NVIDIA and MediaTek deepen partnership · 证券时报: 三十五億美元投資敲定 · 21世紀經濟報道: 爭奪機架級入口

Uber's software factory: 70% of PRs now come from agents, and the token bill went flat

Uber's engineering team published "Running a Software Factory Efficiently at Uber Scale" on August 27 after presenting at the AI Engineer conference, and it is the most detailed operations record of agentic development yet. The scale numbers: more than 70% of pull requests are attributed to local or cloud agents; engineers have built over 3,600 agent skills across the software development lifecycle; those skills execute more than 30,000 times a day. From February to August 2026, weekly active users across all agentic offerings grew 7x and weekly agentic requests grew 9.4x — while total AI spend stayed relatively flat since April. With the model held constant from February to July, cost per 1,000 model requests fell almost 34% from its peak and cost per session fell 52% from its June peak.

The methodology is the transplantable part. Uber decomposes agent spend into six multiplied terms — users, sessions per user, turns per session, requests per turn, tokens per request, price per token — and optimizes the middle terms where agents do autonomous work on top of what an engineer asked. A growing share of sessions are now initiated by machines, not humans: managed agents do code review (uReview), self-heal CI failures, complete end-to-end PRs with visual validation, triage on-call alerts and debug incoming bugs. Underneath sits the AI Context Graph — 24 million nodes and 80 million edges across 30+ internal systems — which agents query in natural language instead of spelunking through code: one grounded agent answered a data question in 38 seconds where the ungrounded version spent 20 minutes, spawned two subagents, hit three errors and concluded the dataset was unqueryable (it wasn't). My read: the headline "70% of PRs" gets the clicks, but the real lesson is Uber's insistence that cost control means eliminating waste, not downgrading models — every workload gets a Pareto-optimal model chosen by benchmarking real PRs, and managed agents beat interactive sessions on ROI. This is the template every engineering org scaling agents will copy for the next year.

— Uber Engineering (official) · InfoQ · AGI Hunt
🔗 Uber Engineering: Running a Software Factory Efficiently at Uber Scale · InfoQ: Uber 公開 AI 軟體工廠省錢方法 · AGI Hunt on the key metrics

Perplexity ships hybrid compute on Mac: cloud for thinking, local for your files

Perplexity launched hybrid compute for its Mac app on September 1, available to Pro, Max and Enterprise subscribers. The idea is a clean division of labor: each Perplexity Computer task starts in the cloud, where frontier models handle reasoning, web search and planning — but the moment a task touches a private file, sensitive data or an on-device action, it moves to a local model running on the Mac itself, so that data never leaves the machine. A "privacy gate" on the device decides what may leave: before anything transmits, an on-device classifier scans the task and can mask sensitive spans, keep the step local, refuse the action or ask for consent. Credentials, payment card numbers and government IDs get the strictest treatment. Three local models launch today — Gemma 4 E4B, Qwen3.6 35B-A3B and a Perplexity model post-trained for Computer — and local work consumes no cloud credits. The hardware bar is real: Apple silicon, macOS 15+ and at least 24GB of unified memory (32GB recommended), which rules out the 8GB and 16GB Macs many professionals own.

The credibility move is that Perplexity open-sourced the classifier that makes the routing decision: PII-Tracer, a 0.6B-parameter bidirectional encoder adapted from a Qwen3 backbone, which the company says recorded the highest character F1 among the 12 systems it evaluated on PII-TRACE, a new 13,148-conversation benchmark spanning 13 languages. CEO Aravind Srinivas framed it directly: hybrid compute is the answer to users who want frontier intelligence without surrendering client files, deal documents or privileged records to a cloud log. My read: this is a genuinely different privacy architecture from the "we delete your data after 30 days" arms race OpenAI and Anthropic are running — instead of promises about retention, Perplexity makes it structurally impossible for the sensitive half of the work to leave the device at all. For lawyers, finance teams and agencies working with confidential material, that distinction may matter more than any model benchmark. The catch is timing: the announcement lands days after a Forcepoint proof-of-concept showed AI assistants could be silently hijacked via invisible HTML in emails — local processing helps, but a compromised Mac is still a compromised Mac.

— Perplexity (official blog) · Unite.AI · 9to5Mac
🔗 Perplexity: Introducing Hybrid Compute on Mac · Unite.AI on the launch · 9to5Mac on hybrid compute

Hugging Face's $399 Microduck sells out — the Raspberry Pi moment for embodied AI

Hugging Face's robot subsidiary Pollen Robotics opened pre-orders for Microduck on August 27, and the little bipedal duck has become the surprise hardware hit of late summer. Microduck is 25 cm tall, weighs under 800 grams, and costs $399. It walks, sits, stands up after falling, kicks a ball, does a forward roll, picks objects up off the floor with its articulated beak — and, with an optional $39 roller-skate accessory, can skate around under a different reinforcement-learning policy. Inside: a Rockchip RK3566 with an AI accelerator, 15 motors driven at 50Hz, a camera, LiDAR, an 8×8 time-of-flight array, two IMUs, two NFC antennas, Wi-Fi/Bluetooth, and a removable battery good for about an hour. The software is the real product: everything is Apache-2.0, the training stack (microduck_rl) runs PPO in MuJoCo Warp at 4,096 parallel environments, a usable gait takes roughly one to two hours on a CUDA GPU, and the sim-to-real recipe models motor friction, battery sag, command delay and gear backlash rather than assuming an ideal servo.

The market response was the story. First-24-hour orders exceeded $2.6 million — more than 5,000 units, peaking at roughly one sale every four seconds — and Hugging Face soon posted that new orders face a four-to-six-month backlog, meaning orders placed today arrive in early 2027. Hugging Face's stated target is 20,000 units cumulative, based on co-founder Thomas Wolf's observation that the best-selling robots in history top out around that number. CEO Clem Delangue's line on launch day frames the ambition: "This is a $399 open-source robot you can teach new tricks with reinforcement learning — welcome to the era of affordable open-source robotics." My read: the "Raspberry Pi moment" label is deserved but incomplete — Microduck's real significance is that it compresses a full embodied-AI research pipeline (RL training, sim-to-real transfer, ONNX deployment, a Rust real-time runtime) into a consumer-priced box, following Reachy Mini's 10,000-plus shipped units. Whether it becomes a learning platform or a shelf ornament depends on whether the developer community actually builds on it, but the sellout is the strongest evidence yet that physical AI has a consumer demand curve. One caveat worth naming: a camera-equipped robot headed into bedrooms is a privacy question Hugging Face's open-source framing doesn't fully answer.

— Hugging Face @huggingface (official) · 極客公園 · AIBase
🔗 Hugging Face announcement on X · 極客公園: 399 美元的小黃鴨 · AIBase: Under $400, you can train it yourself

Top comments (0)