DEV Community

RAXXO Studios
RAXXO Studios

Posted on Originally published at raxxo.shop

Anthropic Says Claude Now Leads 26% of Its AI Research

  • Anthropic said on September 17 that Claude now leads 26% of its AI research and development, up from effectively nothing at the start of the year

  • Claude also collaborates with staff on more than 90% of research work as of August, according to the company's own figures

  • The report introduces a framework for tracking AI acceleration inside frontier labs, including oversight metrics and compute spent on AI R&D itself

  • For a one-person studio running Claude Code daily, the honest reading is that the review step matters more as the tool gets more capable, not less

What Anthropic Actually Published

On September 17, Anthropic published a report defining what it calls "leading" an R&D task: Claude can take a high-level prompt and carry most of a task through to completion end to end, while a human stays in a supervising role rather than doing the work directly. By that specific definition, the company says Claude now leads 26% of its own AI research and development, a jump from close to zero when the year started. Multiple outlets that covered the report, including Bloomberg, the Washington Post, and the Associated Press, described the underlying figures consistently, and I checked the framing across those independent sources before writing anything here.

The company did not stop at the headline number. It also reported that Claude collaborates with Anthropic staff on more than 90% of research work as of August, a broader measure than the 26% "leads" figure because it counts any meaningful contribution, not just full end-to-end ownership of a task. The gap between those two numbers, 90% collaboration versus 26% full leadership, is itself informative. It suggests Claude is embedded almost everywhere in the research process already, but only recently capable of carrying a meaningful share of that work start to finish without a human doing the driving.

Coverage across the outlets I checked framed the 26% figure as new territory specifically because it started near zero. A number that climbs from roughly nothing to a quarter of an entire research organization's output inside a single calendar year is a steeper curve than most capability metrics this industry usually publishes, which tend to move by single-digit percentage points between model generations. That steepness is likely the actual reason Anthropic chose to publish a whole framework alongside the number rather than just the number itself. A slow-moving metric does not need a proposed oversight structure attached to it. A metric moving this fast arguably does.

It is also worth noting what the report did not claim, since the coverage was consistent on this point too. Nothing in Anthropic's own description suggests Claude is choosing what research to pursue, setting its own objectives, or operating with reduced human sign-off compared to earlier in the year. The 26% figure describes execution of tasks that were still scoped and assigned by people, carried through to completion by the model, and reviewed afterward. That is a large jump in throughput, not a change in who decides what gets built next.

Anthropic frames this as a self-measurement exercise rather than a capability announcement to celebrate. Alongside the R&D figures, the report proposes ways to track whether frontier AI development is accelerating responsibly: how much agent activity gets monitored, how long a flagged action takes to get reviewed, how often agent behavior actually gets flagged in the first place, and how much compute a lab devotes specifically to AI research and development rather than to other work. The intent, as described across the coverage, is to give outside observers a way to check whether the pace of AI improving itself is being watched carefully, not just reported after the fact.

Why This Builds on Something Anthropic Already Said in July

This is not the first time Anthropic has put a number on how much of its own work Claude is doing. I wrote in July about Pacing the Frontier, the open letter Anthropic and OpenAI both endorsed the same day it went live, which leaned on an earlier Anthropic report showing Claude had authored more than 80% of the code merged into Anthropic's own production codebase as of May. That was a narrower measurement, specific to code, and it was already a striking number on its own.

This new report measures something broader than code merges. Twenty-six percent "leading" R&D work end to end is a claim about the whole research process, not just what gets committed to a repository, and the 90% collaboration figure suggests Claude's presence in that process has become close to universal even where it is not yet fully autonomous. Read together, the two reports describe the same underlying trend from two different angles, two months apart: a company that keeps choosing to measure and publish exactly how much of its own advancement its own model is responsible for, rather than let outsiders guess.

What stands out to me is that Anthropic keeps publishing these numbers even though a more cautious company might treat them as awkward to admit. A 26% figure invites the obvious follow-up question, which is how fast that number moves next, and whether the oversight tools the report proposes can actually keep pace with it. Anthropic seems to be betting that transparency about the trend line is safer than silence about it, which is consistent with the argument behind Pacing the Frontier back in July.

What "Leading R&D" Actually Means in Practice

It is worth being precise about what this figure does and does not claim, because "AI builds itself" headlines tend to flatten an already careful distinction Anthropic drew in its own report. Leading a task, in this framing, means Claude can take a prompt like "improve this training pipeline's efficiency" and carry the work through multiple steps, decisions, and iterations without a human doing the intermediate work, while a human still reviews the outcome before it ships. That is meaningfully different from a model operating with no oversight at all, and it is also meaningfully different from a model that only ever executes narrow, fully specified instructions.

The 90% collaboration figure covers a much wider band of involvement, everything from Claude drafting an early idea a researcher then substantially reworks, to Claude executing most of a well-specified task with light review. Coverage of the report noted that this broader number has been climbing for a while and was already high earlier in the year, while the 26% "leads" figure is the one that moved sharply, which is the more informative signal about where Claude's actual capability jumped rather than where its presence merely expanded.

None of the coverage I checked described this as full autonomy, and Anthropic's own definition explicitly keeps a human in the supervising seat. The oversight metrics proposed alongside the headline number, how much agent activity gets monitored and how quickly a flag gets reviewed, only make sense as a concept if the company assumes human review remains part of the process. The report reads less like an announcement that Claude has taken over research and more like an attempt to put honest numbers on a trend that was going to happen with or without the measurement, so that the trend is at least visible while it happens.

What This Means for Someone Running a Studio Through Claude Code Alone

I do not do AI research, and Anthropic's internal R&D pipeline has nothing structurally in common with a one-person storefront and a handful of tools. But the shape of the finding still lands close to home, because the shift from typing every line myself to directing an agent that carries a task through multiple steps is the same shift Anthropic is measuring at a much larger scale. The 26% figure is a number for a research organization. The underlying behavior, an agent taking a high-level instruction and returning most of a finished task rather than a single completed line, is the same behavior I rely on every day building and maintaining RAXXO tools.

The part of this report worth sitting with is not the percentage itself, it is the pairing of a rising capability number with a proposed oversight framework published in the same breath. Anthropic is not saying capability growth alone is the story. It is saying capability growth without visible, checkable oversight is the actual risk, and the two need to be reported together or the number becomes meaningless on its own. That is a useful frame for a much smaller operation too. The value of directing an agent through more and more of a task is not just about speed on my end. It is about whether my own review step keeps pace with how much of the work I am handing over, the same trade Anthropic is trying to make visible at its own scale.

Bottom Line

Anthropic's new report says Claude now leads 26% of the company's AI research and development end to end, up from almost nothing at the start of the year, and collaborates in some form on more than 90% of research work as of August. It pairs those numbers with a proposed framework for tracking oversight and compute spent on AI R&D itself, an attempt to keep the trend visible rather than let it advance quietly. This follows a July report that measured a narrower slice of the same trend through code authored rather than research led, and together the two paint a consistent picture of a company that keeps choosing to publish exactly how much of its own progress its own model is now responsible for.

For me, the number that matters is not 26% specifically, since that figure belongs to a research operation with nothing else in common with a small storefront. What matters is the reminder built into how Anthropic framed it: rising capability is only a good trade when the review step around it grows just as deliberately. That has been true of every RAXXO tool I have shipped by directing Claude Code through most of the work, and it is not going to stop being true just because the agent gets better at carrying more of that work on its own.

Top comments (0)