On September 22, 2026, Anthropic released Claude Opus 5.5 (API model id claude-opus-5-5). It is the first model in the new 5.5 family and it is positioned as Fable 5.1-level work at a meaningfully lower run cost. The three things that matter most to anyone shipping code with it:
- 1M token context by default, 128K max output on the synchronous Messages API (up to 300K on Batches with a beta header).
- Adaptive thinking is always on and cannot be disabled. You steer depth with an effort parameter (low / medium / high / xhigh / max, default medium) instead of toggling thinking off.
- Lower list price: roughly $4 / $20 per million input / output tokens, cache reads at $0.20, Batch at 50% off. Anthropic claims about 40% lower cost to run than Opus 5 on typical workloads. None of those numbers are set in stone, so treat the exact price as a launch reference and confirm against the official pricing page before you quote it anywhere. What actually breaks in existing code If you are migrating from Opus 5, four changes will fail requests rather than degrade gracefully:
- Thinking cannot be disabled. Sending thinking: {"type": "disabled"} or thinking: {"type": "enabled", ...} returns a 400. Omit the field entirely and use effort.
- Forced tool choice is rejected. tool_choice types any and tool return a 400, as on Fable 5.1. Use auto with strict tool use.
- Thinking blocks are bound to their model and conversation. You cannot lift a thinking block produced by one model and replay it elsewhere; treat thinking output as ephemeral.
- Old computer-use tool is gone on the API. computer_20251124 returns a 400 on Claude API and Google Cloud; the new computer_toolset_20260801 is required, and computer_toolset_20260801 is the one to pin. The practical takeaway: grep your codebase for thinking and tool_choice before you flip the model id. Most "it worked yesterday" bugs after a model switch are one of those four. Calling it through an OpenAI-compatible client The fastest way to experiment without rewriting your HTTP layer is the OpenAI Python SDK pointed at a compatible endpoint. This works whether you run it against Anthropic directly or through a unified gateway that exposes the OpenAI shape: import os from openai import OpenAI
point base_url at any OpenAI-compatible endpoint
client = OpenAI(
api_key=os.environ["OPENAI_API_KEY"],
base_url="https://easy88ai.com/v1",
)
resp = client.chat.completions.create(
model="claude-opus-5-5",
messages=[{"role": "user", "content": "Write a quicksort in Python"}],
extra_body={"effort": "medium"},
)
print(resp.choices[0].message.content)
Notes that save you a debugging session:
- Pass the model as claude-opus-5-5. The id has no date suffix.
- Do not send thinking at all. Control depth with extra_body={"effort": "medium"}.
- extra_body is the escape hatch for Anthropic-specific fields when you are on an OpenAI-shaped client. Wiring it into Claude Code If your agents run under Claude Code, you usually do not touch code — you set two environment variables so every call routes through your chosen endpoint: export ANTHROPIC_BASE_URL="https://your-endpoint.example/v1" export ANTHROPIC_API_KEY="sk-xxx" Or in the settings file: { "env": { "ANTHROPIC_BASE_URL": "https://your-endpoint.example/v1", "ANTHROPIC_API_KEY": "sk-xxx" } } Once that is set, the CLI, the SDK calls it spawns, and any sub-agents all inherit the same base url. That single line is the difference between "I reconfigured one project" and "my whole toolchain points at the model I want." Thinking and effort: what changes for prompting Because thinking is now always on, a few habits that worked on older models need rethinking:
- Stop trying to disable thinking. On Opus 5.5 a 400 is the only result. If your old prompt engineering relied on turning thinking off for speed, move that intent into effort instead. effort: "low" is the closest thing to "think fast," and it is far cheaper than max on routine steps.
- Let the model think before it acts. With adaptive thinking, the first tokens of a response are reasoning. If you stream and parse results mid-stream expecting immediate text, you may capture an empty text block at the default display setting. Consume the final message content, not the intermediate thinking, for downstream parsing.
- Effort is a budget, not a switch. Medium is tuned to be the default sweet spot; high and xhigh buy more reasoning for genuinely hard tasks (deep refactoring, novel algorithm design) but cost proportionally more tokens. Profile a sample of your real tasks at two effort levels before committing the whole pipeline to one.
- Prompt for structure, not for reasoning off-switch. You no longer need to say "think step by step" to force reasoning; you need to say "return JSON with these keys" or "stop after the diff" so the always-on thinking has a clear target. Clearer output contracts reduce the wasted tokens that thinking can otherwise spend wandering. This is the part of the upgrade that is easy to underestimate. The model id flip is one line; retraining your prompts and your parsers around always-on thinking is the real migration. Where the real savings come from The headline "40% cheaper" is a blend of two things: a 20% list-price cut and the model using fewer tokens to finish the same task. You cannot control the second one, but you can absolutely control the first lever — cache reads. On Opus 5.5, cache reads cost $0.20 per million tokens, which is 5% of the base input price. For agentic coding, where the same repository context is re-read on every turn, cache reads dominate the bill. Mark the stable prefix (system prompt, project background, constant definitions) as cacheable: system_prompt = "You are a senior Python engineer. Project background and constants:\n"
resp = client.chat.completions.create(
model="claude-opus-5-5",
messages=[
{"role": "system", "content": system_prompt},
{"role": "user", "content": "Refactor parse_config in utils.py to support nested dicts"},
],
extra_body={"cache_control": {"type": "ephemeral"}, "effort": "medium"},
)
A few more levers:
- Drop effort to medium, not max. Unless a task genuinely needs the deepest reasoning, medium shaves token usage on long runs with little quality loss for routine work.
- Batch the offline stuff. Anything that is not interactive (bulk summarization, test generation, code review of closed PRs) goes to Batch at half price.
- Route by difficulty. Simple transformations do not need Opus. A lightweight model at ~$2 / $10 per MTok finishes them for a fraction of the cost. Keep Opus for the hard, long-horizon tasks where its 1M context earns its keep. A quick cost mental model Take a coding job that feeds 2M fresh tokens, re-reads 8M cached tokens, and writes 1M tokens. On Opus 5.5 that is roughly $4×2 + $0.20×8 + $20×1 = $29.60. On Opus 5 the list price was higher on every line, so the same shape lands near $39. The gap widens the more you cache and the longer the agent runs. The lesson is not "Opus 5.5 is cheap" — it is "Opus 5.5 rewards you for caching and for not maxing effort on trivial steps." Error handling you will hit Error Likely cause Fix 401 Key invalid or unset Verify the env var is actually exported in the process; check the key prefix 429 Rate limit or quota Lower concurrency, add exponential backoff, or move heavy jobs to Batch 400 thinking disabled Old code disabling thinking Remove the thinking field, switch to effort 400 forced tool_choice any / tool not supported Use auto with strict schema region / timeout Unstable network path Use a stable endpoint, add request timeout and retries Wrap calls so a single failure degrades instead of aborting the whole run: def call_model(client, model, prompt, effort="medium"): try: resp = client.chat.completions.create( model=model, messages=[{"role": "user", "content": prompt}], extra_body={"effort": effort}, ) return resp.choices[0].message.content except Exception: # fall back to a lighter model on the same endpoint fallback = "claude-sonnet-5" if model != "claude-sonnet-5" else "claude-opus-5-5" resp = client.chat.completions.create( model=fallback, messages=[{"role": "user", "content": prompt}], ) return "[fallback] " + resp.choices[0].message.content Should you switch? If you run long-horizon agentic coding, the 1M context plus cheap cache reads are a real win — re-reading a large repo every turn stops being the dominant cost. If you mostly do short Q&A or already have a stable Sonnet-based pipeline, the move is less urgent; keep Opus 5.5 in reserve for the tasks that actually need it. A concrete example of where it pays off: a repo-audit agent that loads a 200K-token codebase, reasons over it, and writes a patch. On a model without cheap cache reads, every one of the dozens of tool-loop turns re-pays for that 200K context; on Opus 5.5 the re-reads fall into the $0.20 cache-read tier, and the run that used to cost a few dollars drops to pocket change. That is the workload shape to migrate first — not the quick one-off prompt. One migration order that avoids surprises: set effort, remove thinking, fix tool_choice, then flip the model id and watch the first hundred calls for 401/429 before trusting it in production. Keep the fallback wrapper in place for at least a week; the errors you will actually see in the wild are almost never the four breaking changes — they are quota and network flakes, and a boring try/except around the call saves more incidents than any model-level tuning. Why a unified endpoint helps here I run a unified OpenAI-compatible gateway at easy88ai.com that fronts Claude, GPT, and DeepSeek through one base url. The reason I reach for it as base_url is not marketing — caching, effort tuning, and the fallback above all live in one place, configured once instead of three times. When something breaks, I have one audit stream, not three dashboards.
Top comments (0)