Anthropic's Claude Opus 5.5 is a real capability and price release, but the model's effort setting is part of the specification rather than a cosmetic control. That matters because the same model can look cheaper, slower or stronger depending on how much reasoning the serving system allows it to use.
Key facts
- Anthropic launched Claude Opus 5.5 in September 2026.
- Anthropic reports 54.6% on FrontierCode at default or medium effort.
- Its “40% cheaper” language is scoped to typical workloads versus Opus 5.
- Primary source: Anthropic's Claude Opus 5.5 announcement.
Anthropic's release says Opus 5.5 is stronger at coding and agentic work and advertises a lower typical-workload cost. The company also reports 54.6% on FrontierCode at medium/default effort, against GPT-6 Astra's highest reported score of 53.3%, and says that run cost about one fifth as much per task. Those are legitimate announced results, but they are not a universal ranking of models or bills. Anthropic's effort documentation explains the hidden variable: effort controls the amount of reasoning work a request may receive.
A useful analogy is a legal researcher. “Use the same researcher” tells you little unless you also say whether they get ten minutes, an afternoon, access to a database and a reviewer. Higher effort may find a subtle dependency; it may also consume far more output tokens. That is the model-market version of test-time compute. It makes static price comparisons deceptively neat.
The independent counterargument is quantitative. In an Artificial Analysis comparison, medium-effort Opus scored 51 against Astra at xhigh's 52 on its broader Intelligence Index, and cost $1.34 versus $2.31 per task—roughly 42% less, not 80%. At higher Opus effort, the model can score more strongly on some comparisons while costing more per completed task. METR's analysis supplies the other necessary caution: its assessment is incremental capability, not a public demonstration of a new autonomy threshold.
Anthropic says Opus 5.5 is designed to handle complex work; the responsible interpretation is that it offers another strong point on the cost-quality frontier. The strongest community debate is not whether a benchmark cell is “true,” but whether providers disclose comparable effort, tools, retries and stopping rules. A team selecting it should run a fixed task suite at documented effort levels, measure success and review cost separately. The caveat is important: no public result can tell an organization how its proprietary repository, tool permissions or workflow will behave.
Originally published on Ground Truth, where every claim is checked against the primary source.
Top comments (0)