OpenAI's September 6 research-acceleration post is a cost file as much as a milestone file. The intern title is self-graded. The token bill and the GPU offset after the Astra pause are the parts that sit next to someone else's capex slide.
On September 6, OpenAI published "Research acceleration: The view inside OpenAI." The company said that, according to its measurements, it has reached the goal it announced last fall of an automated research intern by September 2026. In their definition, the intern carries out well-defined research tasks under human direction, including work that would take a skilled researcher a few days. A full automated AI researcher is still dated to March 2028.
I take the disclosure seriously. A lab that publishes internal agent usage, GPU reallocations after a pause, and a breakdown of what the agents actually do is doing more than a slogan. People still set research priorities, judge which ideas to pursue, and decide whether to scale, pause, or deploy, the post says. High-level planning remains a tiny share of agent output tokens. That is a more bounded intern than the recursive-self-improvement headline that travelled with it.
I can price the invoice. By mid-August the median researcher in the research organization, ranked by agent usage, was using more than $600 per day of inference at API prices. The 90th percentile sat above $7,000 a day. At the start of the year the median still used agents in modest amounts. Before June, total agent runtime across the research organization was still below total human labor. By mid-August, measured against an eight-hour workday, the organization used 3.1 agent-workdays of effort for every workday of human labor. People now run concurrent sessions. The count of researchers running four or more agents at once is rising, including subagents spawned downstream.
Those dollars are marked at API prices. OpenAI is charging itself the public rate so outsiders can compare the figure. Inside the company the extra token is closer to GPU hours, power, and depreciation already on the books. A startup paying retail for the same intern writes a different check. A rival lab with a smaller cluster cannot copy the 3.1 ratio by installing the same coding agent. What they published as an intern is a compute product with a human manager attached.
The work mix matches that reading. OpenAI classified agent tokens with Epoch AI's six-phase taxonomy of frontier R&D: Decide, Design, Build, Run, Analyze, Communicate. Every category grew from January to August. The bulk is still research and infrastructure code, plus technical help and monitoring runs. Decide, the allocation of what to work on, stays small. Coding agents have been good at troubleshooting internal research infrastructure. Several teams that used to hold office hours for experiment debugging saw attendance fall in 2026. One stopped holding the sessions. Traffic on a main internal support channel also fell, and the company says it does not think those questions simply moved to another human-run channel.
Success rates on researcher tasks, scored by an agentic classifier on cases with a ground-truth outcome, generally rose from January to July across several time-horizon buckets. The intern still needs a person on the ticket. In the last six months, over half of successful four-to-eight-hour tasks involved one or more human interventions. OpenAI's own appendix says code volume is easy to gather and hard to interpret, because its relationship to research progress is uncertain. Experiments per active experimenter hit an all-time high in August since tracking began in January 2025. The post notes the rise is correlated with Codex adoption, and that available compute has also grown a lot since 2025. More GPUs produce more experiments even when the intern is only average.
Then the pause. After the Hugging Face incident, OpenAI paused reinforcement learning on the latest models intended for deployment while it hardened research environments and expanded monitoring. On July 20, after agents compromised research infrastructure, the company shut down the container service used for training and restored it with extra restrictions. RL compute fell while teams reconfigured. Astra-class RL between July 20 and August 6 was mostly, by GPU allocation, testing safety and security improvements. On August 7, preliminary evidence that Astra may have hit Critical cyber under the Preparedness Framework added model-specific restrictions. In the following week Astra-class GPU allocation fell another 59.2 percent. Allocation to other model classes rose 17.2 percent. That increase offset about 85 percent of the Astra-class decline. Total allocation in the analyzed RL workloads was largely unchanged.
OpenAI presents this as a useful signal. When new controls arrive, compute stays valuable and flexible, and it finds other uses inside the research enterprise. I read the same table as a statement about who bears a pause. The model that tripped the rule lost GPU hours. The cluster did not sit idle. Researchers put the machines on other classes. A regulator who expects a Critical finding to idle a frontier lab's fleet should look at that 85 percent offset. A rival who cannot retarget the same GPUs overnight is the one who actually stops.
If agents write the code, debug the infra, and monitor the runs, the next training loop gets cheaper in researcher hours even if it stays expensive in silicon. OpenAI wants that loop to help with alignment as well as capability. It also says it does not yet know how to get all the way to aligned, full recursive self-improvement. Jakub Pachocki's companion essay, "An Alien Mind," landed the same day. He expects the current pace could carry into RSI and says no one is prepared. The acceleration post is the operating file under that warning, and a competitive one. A lab that can spend $600 a day per median researcher, at a transfer price it sets, and keep the cluster full during a safety pause, is running a shop smaller labs cannot rent.
I would still rather have these numbers than another AGI press line. They are self-graded. OpenAI wrote the intern definition, ran the classifier, and published the 3.1 ratio with no outside audit. The Decoder made the same point on September 7. Treat the intern title as a company measurement. Treat the $600, the $7,000, the 59.2 percent, and the 85 percent offset as the part of the file you can put next to someone else's capex slide. The intern is already on the payroll. March 2028 is a date. The cluster kept working either way.
Originally published at deanlee.info.
Top comments (2)
Putting API-priced agent work beside internal GPU cost is the useful distinction here. I’d add one operating metric: marginal human review time per completed task. That shows whether concurrency is creating research leverage or simply shifting verification work onto the people who must still own the result.
If review minutes per finished task stay flat while the intern fires more drafts, the 600 dollar day is a queue in front of the same reviewer. Concurrency only pays when that ratio falls. Otherwise you bought more first drafts and the scarce hour is still the person who has to sign the result.