DEV Community

AgustaON
AgustaON

Posted on AI-assisted

The Week AI Shipped With Locks On: GPT-6 Astra and Claude Fable 5.1

GPT-6 Astra and Claude Fable 5.1 landed 48 hours apart. The benchmark charts are the least interesting part of either launch.

The first week of September 2026 compressed a year of AI industry direction into five days. On Monday, the European Commission formally designated ChatGPT a Very Large Online Search Engine, the first chatbot ever placed in the same regulatory category as Google Search. On Tuesday, Anthropic released Claude Fable 5.1. On Thursday, OpenAI released GPT-6 Astra.

The coverage has mostly been score tables, and the score tables will be stale by spring. Three quieter things shipped alongside them, and those will still matter in two years: the frontier now comes with locks, with labels, and with a new judge deciding questions that affect businesses which have never heard of either model.

What actually shipped

Claude Fable 5.1

Anthropic's release is really two releases. Fable 5.1 is the generally available model. Mythos 5.1 is the same model with looser safeguards, available only to vetted organisations working in cybersecurity and the life sciences. On capability, Anthropic reports 52.6% on Terminal-Bench-Science, roughly double its predecessor's 24.7%, and Artificial Analysis measured it at 66 on its Intelligence Index, the highest score it has recorded. The commercial headline is a 75% cut to cached input pricing, from $1.00 to $0.25 per million tokens, which VentureBeat notes works out to roughly 25% lower costs on typical workloads and up to 45% on long-running agent tasks.

And then there is the line most coverage skipped: Fable 5.1 and Mythos 5.1 are the first Claude models whose text output carries an invisible watermark.

GPT-6 Astra

Astra is OpenAI's first full version-number jump in over a year, and the company is not being modest about it. State of the art on computer use, browser use and software engineering. Asked whether it amounts to AGI, OpenAI president Greg Brockman told Axios "I do think we're there."

It is also the first OpenAI model rated Critical for cyber capability under the company's own Preparedness Framework. In OpenAI's internal testing it solved every task on its exploit-development benchmark and, given twenty recent high-severity vulnerabilities to work through, independently discovered and used two previously unknown ones along the way. Rollout is phased: a small set of organisations first, then ChatGPT Plus, Pro, Business and Enterprise accounts, the API and AWS over the following days. Enterprise administrators get it switched off by default.

The locks: trusted access is the new release model

Notice the shape both companies chose, independently, in the same week. Anthropic splits its flagship into a public version and a vetted-access version. OpenAI rates its flagship Critical, gates the advanced cyber capabilities behind a small tester group, and plans wider access through a vetted security programme. Neither company slowed its release down. Both put doors on it.

This is now the pattern, and it changes something subtle about the ecosystem: the most capable version of the model you can buy is no longer the most capable version that exists. There is a tier above the price list, and entry to it is an application process, not a credit card. For most teams that changes nothing tomorrow. For anyone whose planning assumed public frontier access would always equal frontier capability, it quietly stopped being true this week.

The labels: provenance became infrastructure

Fable 5.1's watermark deserves more attention than it got. It is a statistical signal woven into the text itself, one that estimates how likely it is that Claude was involved in writing a passage. It carries no information about the user, does not change the output, and is invisible without a detection API that sits in private preview for regulators, media organisations, fact-checkers, researchers and similar institutions.

Anthropic did not do this on a whim. It signed the EU AI Act's code of practice on transparency of AI-generated content in July, which requires watermarking for models released after 2 August. Google has been running SynthID through Gemini's text since 2024. Vendor by vendor, AI text is acquiring a receipt.

If you publish content for a living, the practical read is calmer than the headlines will be. No search engine uses a text watermark as a ranking signal, and Google, which has watermarked its own model's output for two years, would be demoting its own product daily if one did. What Google's spam policy actually targets is scaled, unedited content published to fill space. The era that is ending is narrower and fairer: publishing raw model output worked because nobody could prove anything, and the vendors themselves are now signing that deniability away. Edited work carrying knowledge only you have was never the target.

The judge: a new model is answering questions about your business

Here is where the week's three events converge into one story.

The EU designation was triggered by a number: 159.1 million people in Europe use ChatGPT search in an average month, more than three times the threshold that defines a very large search engine, and that figure counts the search feature alone. People ask these systems what to buy and who to hire, at regulated-infrastructure scale. And as of this week, the model doing the answering is being swapped out underneath them.

An AI answer is not a stored ranking. Nothing is retrieved from a league table when someone asks for the best accountant in their city; the shortlist is generated on the spot by whatever model is running that day. Change the model and you change the judgment: what it chooses to read, which sources it trusts, which names make the answer. A business's visibility can move without a single word changing on its website.

We have measured how differently these systems already judge. When we ran the same 50 buying questions across four assistants, ChatGPT named about 8.9 businesses per answer and drew 26% of its citations from directories rather than the businesses' own sites. Gemini named 9.1, used directories only 13% of the time, and 62% of the sources it cited were used by no other assistant. Same questions, four different juries. Astra is a fifth jury being seated this week, in front of the largest audience of them all.

There is recent precedent for how fast this layer moves. In August, ChatGPT changed how it gathers sources, and Reddit, one of the most cited sites on the internet, lost roughly 86% of its citations there. That was a behavioural adjustment to one model. A full model replacement is a bigger event, and it will not send anyone an email.

Two checks are worth doing while the rollout is still in progress. First, the boring one: confirm your site is not accidentally blocking the crawlers these systems read through. One stray line in a robots.txt file undoes everything else, and a free checker will tell you in seconds which of the AI crawlers can reach you. Second, ask an assistant the question your customers ask before they know your name, note who gets named, and ask again once Astra reaches your account. Ask more than once, because answers vary run to run. What matters is not the wobble but the names that disappear entirely.

For the people building on the APIs

Three developer-facing details worth pulling out of the launches. Fable 5.1's cache-read cut is the significant one for anyone running agents or long sessions: cached input at $0.25 per million tokens changes the economics of keeping large context resident, and Anthropic's own estimate of up to 45% savings on highly agentic workloads is plausible arithmetic rather than marketing. Astra arrives in the API and on AWS alongside the ChatGPT rollout, with usage inside existing subscription allowances and credits on top. And the watermark requires nothing from you: it does not alter outputs, there is no flag to set, and the detection API is not publicly available.

The part that lasts

By next spring, some other model will top the charts that these two currently top, and this week's score tables will be trivia. The other three things do not expire. Vetted-access tiers will still sit above the public models. The watermark plumbing will be in more vendors' output, not fewer. And the answers that decide which businesses get discovered will have been regenerated, again, by whichever judge is newest.

Watching that last layer is the whole of what we do at Greater Than Services: the same buying questions, asked week after week across five engines, so a business finds out it stopped being the answer before its revenue does. This was the busiest week that layer has ever had.

The frontier did not just get smarter this week. It got doors, receipts and a new judge. None of those come off with the next benchmark chart.


Sources: Anthropic on Fable 5.1 and Mythos 5.1 · OpenAI, Path to Astra · TechCrunch on the Astra launch · Axios, Brockman on AGI · VentureBeat on Fable pricing · European Commission designation notice · Search Engine Journal on the 159.1M figure · Artificial Analysis on Fable 5.1

Top comments (0)