Two disclosures landed this week that answer the same question from opposite ends. Anthropic published numbers on how much of its own research Claude now performs. Google published a method for making agents spend far less compute to reach the same answer. One measures how fast the field is automating itself; the other shows where that automation currently wastes money.
Below: seven stories from September 17 to 19, 2026.
1. Anthropic says Claude now leads 26% of its own AI research
Anthropic published a set of metrics on Thursday for measuring how fast AI development moves inside a frontier lab. The first number is the one people will quote: as of August 2026, Claude leads 26% of Anthropic's internal AI research and development work. In February, that share was under 1%.
The index borrows a scale built by Epoch AI that runs from AL0, no AI involvement, to AL5, full autonomy with no human in the loop. AL3 means AI collaborates under close human direction. AL4 means the model completes most of a task from a high-level prompt while a human supervises. Anthropic says no measured R&D work has reached AL5, where a system would identify a problem, decide what work is needed, build and test a solution, and ship it without a person involved. More than 90% of measured work now sits at AL3 or above.
The method is worth reading before the number gets repeated. Each week in July, Anthropic randomly sampled 20% of staff in departments involved in model development. A Claude research agent reviewed their work and identified roughly 15,000 granular tasks across 542 categories covering model training, reinforcement learning, evaluation, and engineering. Tasks were weighted by person-time, so work that consumed more employee hours counted more. Claude's ratings matched human ratings exactly 59% of the time, against 35% agreement between two humans, and 97% of ratings landed within one automation level of each other.
The agent numbers are as striking as the R&D index. About 30,000 agents were running research and engineering work at any one time on Anthropic's most-used internal platform in August. Every action passes through an online monitor before it executes. Anthropic analyzed more than 1 billion agent decisions that month and blocked about 0.002% of them, roughly one in 47,000. A separate offline system flags around 100,000 transcripts a week for automated review, and about 50 of the highest-priority cases reach human reviewers each week.
On compute, a sampled week in July put about 6% of AI R&D compute toward safety work. For compute used specifically in AI-driven R&D, the figure was about 12%. Anthropic calls both estimates conservative, because compute that advances capability and safety at the same time is not counted.
Two limits are stated up front. There is no shared measurement standard across labs, so these numbers cannot be compared with anyone else's. And Anthropic used its own models to assess itself, which leaves open the possibility that the evaluator shares the failure modes of the thing being evaluated.
โ Anthropic
๐ Anthropic
2. Google's Dream-RSI replays old searches instead of rerunning them
A team of 17 researchers from Google, Google DeepMind, the University of Maryland and the University of Virginia published Dream-RSI, a framework that cuts the number of discovery-agent calls by up to 162x without touching the underlying model's weights.
The mechanism rests on a simple observation: every agent run already leaves a record of what it tried and what happened. Dream-RSI runs in three stages. In online explore, the current policy guides an agent to build a historical discovery tree that logs every decision and its outcome. In construct replay simulator, those trees become a reusable repository that stands in for the realized search space. In dreaming-based policy improvement, a separate model proposes alternative exploration policies and scores them against recorded history. Because outcomes are already on disk, testing a new policy means reading records rather than rerunning the discovery agent or its evaluator. The winning policy goes back online and the loop repeats.
On a Lasso path-solver task, the SimpleTES baseline needed 51,200 discovery-agent calls. Dream-RSI on Gemini 3.1 Pro needed 317, with average runtime falling from 3,587.1 milliseconds to 2,931.0 milliseconds. On Gemini 3.7 Flash the count dropped from 3,200 to 1,879. The solvers it produced beat sklearn and glmnet across six held-out datasets. On KernelBench, it matched existing methods while generating 2.43x fewer candidates on VGG16 and 1.79x fewer on LayerNorm.
One finding runs against intuition. The researchers tried summarizing past exploration into natural-language hints, the way most agent memory systems work, and feeding those hints into the next prompt. That performed worse. The summaries pushed the agent to eliminate directions early and narrowed the search, while replaying the full history preserved the diversity that made the improvement possible.
โ arXiv ยท VentureBeat
๐ arXiv:2609.14858 ยท VentureBeat
3. OpenRouter's anonymous Union Alpha scores high, then gets accused of being a router
On September 16, a model appeared in OpenRouter's catalogue as stealth/union-alpha, listed under a provider identified only as "Stealth". The spec sheet reads like a frontier release: 262,144-token context, 131,072-token maximum output, text and image input, tool calling and structured JSON output, and a price of zero for both input and output tokens. OpenCode's Zen service carries the same preview under the model id union-alpha.
The early numbers matched the spec sheet. Union Alpha posted roughly 74% on DeepSWE, a band where GPT-6 Astra and Claude Opus 5 sit about one point higher, at a fraction of the expected cost, plus 90.9% on GPQA Diamond. OpenRouter's chief executive charted it against Opus 5, Fable 5 and Astra on a cost-versus-score plot. First-day traffic reached about 1.96 billion tokens, and OpenCode advertised capacity of 5 trillion tokens per day.
Then developers started pulling it apart. Theo, the t3.gg creator, wrote that OpenRouter released Union Alpha "without telling us it was a crappy model router." The claim circulating among early users is that Union Alpha switches between Llama-3.3-70B, Qwen 3.6 35B-A3B and GLM 5.3 Flash, with DeepSeek v4 Flash reportedly acting as a judge when the system is under load. No reasoning effort setting is exposed, so every comparison runs at one fixed configuration. Over the first three days availability sat at 96.17%, with median latency at 9.5 seconds and the 99th percentile at 474 seconds. On September 17 the project's account said significantly more capacity would come online the next morning.
This is the second time OpenRouter has run the experiment. Its first stealth model, Ox Alpha, appeared on August 20 and was later identified as Z.ai's GLM-5.3 Flash. Stealth terms let the provider retain prompts and completions but prohibit using them for training. Union Alpha's identity remains unconfirmed, and the free window has no published end date.
โ OpenRouter ยท HuggingNews
๐ OpenRouter ยท HuggingNews
4. Cohere absorbs Aleph Alpha to build a sovereign AI stack outside US cloud law
Cohere and Aleph Alpha signed a definitive business combination agreement on Wednesday, creating what the two companies describe as the first transatlantic sovereign AI provider. The combined entity is valued at roughly $20 billion, with Cohere shareholders retaining about 90%. Schwarz Group, the parent company of Lidl and Kaufland, committed โฌ500 million, about $600 million, to lead Cohere's Series E.
The structure keeps both national roots. The company will be dual-headquartered in Berlin and Toronto, with Aleph Alpha's Heidelberg office continuing as a research center, and headcount passing 1,000 across both continents. Aleph Alpha co-CEO Ilhan Scheer becomes Cohere's chief operating officer, and co-founder Samuel Weinbach becomes chief research officer. Aidan Gomez remains chief executive. The transaction still needs regulatory approval and is expected to close later in 2026. Compute will be delivered on Schwarz Digits' STACKIT platform, which the companies position as independent of the US CLOUD Act.
The stack covers both sides. Cohere brings Command A for enterprise reasoning, Aya Vision for multilingual multimodal work, and Parse for document intelligence. Aleph Alpha contributes the Pharia model family, trained in English, German, French, Spanish, Italian, Portuguese and Dutch, plus PhariaAI for orchestration, compliance and source attribution. A joint Command-Pharia foundation model is planned for the fourth quarter of 2026. The companies cite McKinsey research putting AI services above $1 trillion annually, with sovereign AI needs accounting for nearly $600 billion of that.
Funding may go further. The Globe and Mail reported on September 11 that Cohere was in advanced talks to raise $2 to $3 billion in its Series E, beyond Schwarz Group's commitment, with participation from the Canadian government and possibly the German government. At those figures the round would nearly triple the $7 billion valuation Cohere reached in September 2025. PwC Legal confirmed on September 18 that it advised Aleph Alpha on the transaction.
โ Cohere ยท PwC Legal
5. Toyota will put 400,000 robots into its factories, including a two-fingered humanoid
Toyota said Friday it will deploy 400,000 robots across its production system, with 150,000 going into its own plants and 250,000 into group companies, covering roughly 60 factories worldwide. From 2028, Toyota and its group companies plan to spend about ยฅ1 trillion, around $6.7 billion, a year renovating and rebuilding global plants.
The robot doing the talking is ELEY, short for Embodied Learning robot for Enhanced Yield. Each hand carries two fingers, the unit weighs about 50 kilograms, and it moves on a wheeled base rather than legs, running on battery or a cord. ELEY is an upgrade of HSR, the life-support robot Toyota launched in 2012. Its height adjusts so it can pick items off the floor or reach high shelves, and an omnidirectional chassis moves it to a target without a human guiding the path. It runs on Large Behavior Models from Toyota Research Institute, a generative physical AI that learns behavior from sensor data instead of being programmed task by task.
The learning loop already runs on some lines. Workers wear tools modeled on ELEY's fingers and repeat their normal tasks, while the robot watches and picks up the motion. At a September investor session at Toyota's European headquarters, ELEY folded T-shirts at near-perfect accuracy after two weeks and about 1,500 repetitions. Once a skill is learned, the data can be pushed to robots in other plants so they pick up the same task. Toyota says the loop eventually runs both ways, with robots teaching new hires.
Executive vice president Hiroki Nakajima described the goal as a world where robots coexist with people rather than replace them. Toyota operates about 60 plants and employs roughly 18,000 veteran technicians it calls takumi, whose accumulated technique becomes training data. Group headcount was 391,000 at the end of March, about 344,000 of them in automotive. The same week, UBTech opened a 14,000-square-meter humanoid plant in Liuzhou that can produce one robot every ten minutes and 10,000 a year, and XPeng showed systems where robots build robots.
โ Toyota ยท Nikkei
6. Nvidia, Google and Emerald AI want data centers to behave like grid resources
Nvidia, Google and Emerald AI launched the AI Energy Management Alliance on September 16 with 18 founding partners drawn from AI, data centers and power. Anthropic is on the list, along with AES, Constellation, National Grid, NRG and RWE.
The bottleneck is interconnection. New US data centers can wait a decade or more for a grid connection because utilities have to guarantee capacity during peak demand. Emerald AI chief executive Varun Sivaram wrote in Fortune that the grid runs at only about 50% utilization on average, and that if AI data centers could reduce load during the worst hours, the US could unlock 100 gigawatts on the existing grid for flexible facilities. A Goldman Sachs study cited in coverage put the number at 76 gigawatts if maximum grid usage were capped at 90% for a few hours at a time.
The technical proposal is demand response applied to compute. Data centers would shift workloads, discharge on-site batteries, or run on-site generation when the grid comes under stress. Nvidia said in a blog post that US power infrastructure was built for flat, static demand rather than for facilities that can moderate their draw. The alliance says it will be technology-neutral and will measure response speed, curtailment and contingency obligations rather than endorsing particular hardware.
The next step is a demonstration. Nvidia, Emerald AI and Digital Realty plan to build a nearly 100-megawatt power-flexible AI facility in Virginia that adjusts its consumption to grid conditions. AEMA also plans to push state capitals and Washington for a simple trade: faster and larger grid connections for data centers that commit to flexibility, with obligations they are held to. Emerald AI recently raised $150 million in a Series A led by Energize Capital and DCVC.
โ NVIDIA ยท Emerald AI
๐ NVIDIA ยท Emerald AI
7. Manus seeks $500 million at a $4 billion valuation and weighs a Hong Kong IPO
Manus is in talks to raise about $500 million at a valuation near $4 billion, and is exploring a restructuring to prepare for a potential Hong Kong listing, according to the Wall Street Journal. Prospective investors include IDG Capital and Boyu Capital. CATL, the Chinese battery maker, has also held talks. Existing backers Tencent, HSG and ZhenFund are considering participating. The company did not comment.
The backstory explains the price move. Meta agreed to buy Manus last December for more than $2 billion, when the startup was reporting over $100 million in annual recurring revenue. China's National Development and Reform Commission blocked the deal in April 2026, the first publicly halted foreign acquisition in the AI sector since the 2021 foreign investment security review rules took effect. Founders and early backers, including Tencent, HSG and ZhenFund, bought the shares back at roughly the original $2 billion valuation. Tencent became the largest external shareholder, taking over the stake Benchmark had held. In August, Manus told users to export their own data because Meta-era data had to be deleted.
A $4 billion valuation would roughly double the buyback price. Manus builds agents that produce research reports, slide decks and vibe-coded apps, a product surface that overlaps with OpenAI, Lovable and Replit. The talks are described as preliminary, and the company has not confirmed either the round or the listing plan.
โ Wall Street Journal ยท Manus
๐ Wall Street Journal ยท Manus
What to watch next
The gap between Anthropic's 26% and Google's 162x reduction is the interesting part. One says agents are already doing a quarter of the work inside a frontier lab. The other says most of what agents spend on exploration is repetition of paths that already failed. If the second claim holds up outside benchmark tasks, the first number has room to move faster than the compute budget behind it.
KD Agentic publishes this digest daily. Previous editions cover model releases, agent frameworks, robotics and AI infrastructure.

Top comments (0)