Attributed compile (not original research)
Primary source: Introducing Claude Sonnet 5.5 (Anthropic), published 2026-09-28
Copilot availability: Claude Sonnet 5.5 in GitHub Copilot (GitHub Changelog), published 2026-09-28
Builder details (model IDs / effort / Claude Code):* Building with Claude Sonnet 5.5 by **Addy Osmani (claude.dev), published 2026-09-28
This post restates the launch packaging and first-request path in my own words, with a short decision checklist for builders. I did not run Claude Sonnet 5.5 for this article. Every price, model ID, and benchmark figure is taken from a named source and labeled as vendor-reported.
Anthropic shipped Claude Sonnet 5.5 on September 28, 2026, as the second model in the Claude 5.5 family after Opus 5.5. The same day, GitHub made it generally available in Copilot. If you write features, fix bugs, or draft docs from an IDE or Claude Code, that dual launch is the part worth acting on—not another Elo screenshot.
Below is a practical walkthrough: when Sonnet 5.5 is the right default versus Opus 5.5, how to send a first call three ways, what the effort dial actually does, and why an unchanged list price can still cut your bill.
What Anthropic is selling (vendor positioning)
Anthropic frames Sonnet 5.5 as a faster, lower-cost complement to Opus 5.5. Opus stays the pick for complex, open-ended work that needs sustained judgment. Sonnet 5.5 is positioned for well-scoped everyday work: features, bugs, polished documents, slides, and spreadsheets, plus a stronger eye for design than Sonnet 5.
Vendor-claimed upgrades over Sonnet 5 (Anthropic launch post):
- Speed: generates outputs 30%+ faster
- Task cost: typically up to ~30% less per task, even though list price is unchanged
- Agentic coding (vendor): Terminal-Bench 4.0 70.6% vs Sonnet 5’s 10.3%; CursorBench 4.0 55.5%
- Knowledge work (vendor): GDPval-AA nearly level with Opus 5.5
- Context: 1M tokens native (no beta header), per the builder guide
Treat those numbers as Anthropic-reported. They are useful for orientation. They are not a guarantee on your repo or your eval harness.
Same list price, lower task cost
This is the part builders miss when they only skim the pricing table.
| Per 1M tokens (Anthropic) | Sonnet 5.5 | Opus 5.5 |
|---|---|---|
| Input | $2 | $4 |
| Output | $10 | $20 |
| Cache read | $0.20 | $0.20 |
| Cache write (5 min) | $2.50 | $5 |
Sonnet 5.5 matches Sonnet 5 list rates. Anthropic’s claim is that it usually finishes the same work in fewer tokens and fewer steps (more batched tool calls in early testing), so the finished task costs less even when $/token is flat. That is the “list price lies” story: the invoice follows tokens used, not the sticker next to the model ID.
Effort is the other half of the bill. Levels are low, medium, high, xhigh, and max. Defaults differ by surface: medium in Claude Code / Claude apps, high on the Claude Platform API. Higher effort generally means more thinking tokens and higher cost per task. If you are tempted to pin xhigh/max on every Sonnet call, Anthropic’s own builder note says you may lose the quality/speed/cost balance that makes Sonnet useful—and should consider Opus 5.5 instead.
When to pick Sonnet 5.5 vs Opus 5.5
Start here (condensed from Anthropic’s own workload table):
| Workload | Start with |
|---|---|
| Well-scoped coding: bugs, feature iteration, verify against a clear checklist | Sonnet 5.5 |
| High-volume everyday development | Sonnet 5.5 |
| Docs / slides / spreadsheets where polish and design matter | Sonnet 5.5 |
| Repeated, well-defined agent tasks (investigate, review, draft) | Sonnet 5.5 |
| Long-horizon agentic coding or knowledge work that needs careful judgment | Opus 5.5 |
| Hardest open-ended problems | Opus 5.5 |
Rule of thumb: clear spec + checkable result → Sonnet. Ambiguous goal + long trajectory where early mistakes compound → Opus.
First call three ways
1. Claude API (Python)
Model ID: claude-sonnet-5-5. Bedrock ID: anthropic.claude-sonnet-5-5. Thinking is on by default (adaptive), so read content blocks by typedo not assume content[0].text.
import anthropic
client = anthropic.Anthropic()
response = client.messages.create(
model="claude-sonnet-5-5",
max_tokens=4096,
messages=[
{
"role": "user",
"content": "List three failure modes for a coding agent that claims a change is done.",
}
],
output_config={"effort": "medium"},
)
for block in response.content:
if block.type == "text":
print(block.text)
Migration gotcha: if you ran Sonnet 5 with thinking off (thinking: {"type": "disabled"}), that returns 400 on 5.5. Switch to thinking: {"type": "between_tools"} before you change the model ID. between_tools works at low/medium/high only—not xhigh/max.
2. Claude Code
From Claude Code v2.1.284 (Agent SDK for TypeScript v0.3.284+), the sonnet alias resolves to Sonnet 5.5 on the Claude API, at medium effort, with the 1M window native. The default model stays Opus 5.5—switch with /model sonnet for well-scoped tasks. You cannot turn thinking off for Sonnet 5.5 inside Claude Code; effort is how you dial token spend.
3. GitHub Copilot
Per GitHub’s changelog (2026-09-28), Claude Sonnet 5.5 is GA for Copilot Pro, Pro+, Max, Business, and Enterprise. Pick it in the model picker in VS Code, Visual Studio, JetBrains, Copilot CLI, the coding agent, github.com, GitHub Mobile, Xcode, and Eclipse. Rollout is gradual—if you do not see it yet, check again later. Enterprise/Business admins gate access via the Copilot model policy (new models auto-enable unless the global default is off or this model is explicitly disabled). Billing follows provider list pricing under usage-based billing.
A short decision checklist
- Scope first. Clear acceptance criteria → Sonnet 5.5. Open-ended architecture or multi-hour judgment → Opus 5.5.
-
Set effort on purpose. Latency-sensitive chat →
low/medium. Agentic coding with a tight spec startmedium, bump tohighwhen quality dips. Do not inherit Platformhighon every high-volume route. -
Migrate thinking correctly. Old “thinking off” →
between_toolsbefore swapping the model ID. Update the SDK ifbetween_toolsis missing. - Read blocks by type. Adaptive thinking means the first block may not be text.
- Measure tokens on your tasks. Anthropic’s ~30% per-task claim is vendor testing. Log input, output, cache hits, and tool-step counts on a 20-task golden set of your own.
- In Copilot, confirm the picker and the policy. Individual plans: select Sonnet 5.5 when it appears. Org plans: check admin model policy before assuming the team can use it.
- Do not mix harness rankings. Terminal-Bench, CursorBench, and GDPval-AA are different jobs. Compare like-for-like, or run your own trajectories.
Why this post exists
Model launches move weekly. The durable skill is knowing which knobsmodel tier, effort, thinking mode, IDE picker, org policy—change reliability and cost when you drop a new Sonnet into an existing loop. Sonnet 5.5 in the API, Claude Code, and Copilot is a packaging change you can act on today; the scoreboard will keep shifting.
If you try it, start with one production-like coding trajectory at medium and high, compare token counts and failure modes against Sonnet 5 (and against Opus on one hard task), then promote the winner. That experiment beats another screenshot of a vendor table.
Byline: YongBo Yu — Toronto AI engineer (agents, LLM workflows). GitHub: YongBoYu1.*
Top comments (0)