Disclosure: TechSifted has no affiliate relationship with OpenAI. This is editorial news coverage of a publicly announced model launch.
On September 6, 2026, Nvidia CEO Jensen Huang posted a message to X that sent the AI world into a predictable argument. "From ChatGPT to o1 to Astra in 4 years," he wrote. "AGI has arrived. Congratulations @OpenAI team. 400K GPUs coming online next." (Source)
That framing — AGI has arrived — is a claim worth examining carefully. OpenAI itself is more measured. Greg Brockman, the company's president, told Fortune: "It's not unreasonable to feel that we are now in the AGI era, and I think that if you want to say this [model is] the first one, I think it's reasonable." That's an attribution of possibility, not a declaration of fact. Sam Altman has previously described AGI as "a very poorly defined" and "irrelevant marketing term." (Source)
The benchmark organization that designed the test Huang and Brockman are referencing — ARC Prize — is more direct. ARC Prize co-founder Mike Knoop stated plainly: "we lack evidence to call this AGI yet." The organization's official position on ARC-AGI-3: "saturating the benchmark would not represent proof of achieving AGI."
So what actually happened when OpenAI launched GPT-6 Astra on September 3, 2026? And what does it mean for people deciding which AI tools to use today?
What GPT-6 Astra Is
GPT-6 Astra is OpenAI's flagship reasoning and agentic model, released September 3, 2026. The launch started with enterprise customers in OpenAI's Daybreak cybersecurity program and expanded to ChatGPT Plus, Pro, Business, and Enterprise subscribers — Pro and Business customers got access first (September 4-5), followed by Plus users. It is rolling out to the API as well, priced at $10 per million input tokens and $50 per million output tokens. For existing subscribers, Astra usage is included in current subscription allowances, with options to purchase additional credits.
The model's headline capability is computer use: the ability to navigate a computer the way a human would. According to Greg Brockman, Astra can "zip through spreadsheets, fill out forms, and navigate across web pages often at superhuman speed." OpenAI's own VP of Research Mia Glaese put it this way: "Computer use shows how far we've come from sort of aspirationally training for computer use to bringing real value to people every day."
The framing OpenAI and Nvidia are reaching for — "agentic AI era," "AGI era" — is pinned to this capability: the model doesn't just answer questions, it takes actions. It can operate a computer. It can complete a multi-step workflow without being explicitly guided through each step. OpenAI's own example is representative: cat-sitter research that previously took 30 minutes took Astra 5 minutes and 27 seconds. A job search task that previously took 5 hours took Astra 2 minutes and 51 seconds.
Those are OpenAI's own demonstrations, so treat them as best-case illustrations, not independent benchmarks.
If you are comparing Astra to ChatGPT's prior behavior, the practical difference is the gap between a capable text assistant and something that can actually do computer tasks on your behalf. That's a meaningful change in what the product is.
The ARC-AGI-3 Number: Which One Is Real?
Here's where it gets complicated, and where the "99.9%" headline number requires a correction.
There are two scores. The one you've seen in the press is 99.9% on ARC-AGI-3 — but that score came from OpenAI's own Provider Adapter harness, at a cost of $18,817. The score from ARC Prize's independent, provider-neutral standard harness is 62.7%, at a cost of $26,098. Both are on the same model.
The difference is the harness — the software wrapper around a model that controls what tools it can access, how context gets managed between requests, and what information it retains across steps. ARC Prize describes the Provider Adapter harness as "preserving the opaque reasoning state (which we don’t see) between requests," which the standard harness does not.
The result is striking: with reasoning disabled in OpenAI's adapter, Astra still scored 34 points higher than the same model at maximum reasoning in the standard harness. The adapter runs also used 49% fewer tokens and were 3.66 times faster. That's not a minor difference in setup — it's a fundamentally different evaluation environment.
The Next Web also reported that five metrics were revised after publication. Astra's hallucination rate changed from 4.2% to 2% — and then back to 4.2%.
Stanford researchers Anka Reuel and Mike Hardy termed the practice of re-running evaluations under different conditions until scores improve "benchmaxxing." Vincent Sunn Chen of Snorkel AI suggested the industry establish norms requiring companies to disclose what changed between published revisions.
The 99.9% number is real — on OpenAI's harness. The 62.7% number is also real — on ARC Prize's harness. Neither is fabricated. But reporting 99.9% without explaining the harness difference is selective framing.
For context on what 62.7% means: GPT-5.6 Sol scored 7.78% and Claude Opus 5 scored 30.16% on the same standard harness. Astra's improvement is significant regardless of which number you use.
Read the rest of this article on TechSifted →
Originally published at techsifted.com.
Top comments (0)