TL;DR: On Claude Haiku 5.5, the effort setting changes output volume far more than anything else in the request. Independent runs show max effort using about 14 times the output tokens of low for a 14-point gain on the index. Default is medium. Run a sweep on your own prompts, start simple routes at low, and watch for four failure modes: truncated replies, cache resets, early stopping, and empty replies at xhigh.
The engineering problem
Claude Haiku 5.5 changes three things that affect token spend, all at once:
- Thinking is on by default (adaptive). Haiku 4.5 only thought when you asked.
- Thinking tokens are output tokens: they count toward
max_tokensand bill at the output rate ($0.50 per million for prompts up to 100K tokens). - The default effort is
medium, one level below most current Claude models, which default tohigh.
The old control, thinking: {"type": "enabled", "budget_tokens": N}, now returns a 400. Effort replaces it, so there's no number to carry over. And prompting the model to "answer directly" doesn't reduce thinking; in Anthropic's testing it kept thinking anyway. Effort is the only real lever, so it needs measuring.
The five levels are low, medium, high, xhigh, max, set in output_config (see Anthropic's effort docs).
Prerequisites
- An AIHubMix API key in
AIHUBMIX_API_KEY. - A current Anthropic SDK:
pip install -U anthropic. - 20–50 real prompts from production, plus a way to grade the outputs (a rubric, expected labels, or tests).
The script uses the AIHubMix Claude native endpoint, because effort lives in the Messages API's output_config.
The sweep
import os
import time
import anthropic
client = anthropic.Anthropic(
api_key=os.environ["AIHUBMIX_API_KEY"],
base_url="https://aihubmix.com",
)
prompts = [
"Summarize this support ticket in one sentence: ...",
"Extract the invoice number and total from: ...",
]
for effort in ["low", "medium", "high"]:
out_tokens, seconds = 0, 0.0
for prompt in prompts:
start = time.time()
r = client.messages.create(
model="claude-haiku-5-5",
max_tokens=8000,
output_config={"effort": effort},
messages=[{"role": "user", "content": prompt}],
)
seconds += time.time() - start
out_tokens += r.usage.output_tokens
text = next((b.text for b in r.content if b.type == "text"), "")
# Save `text` next to the prompt and grade it against your own rubric.
print(f"{effort}: {out_tokens} output tokens, {seconds:.1f}s total")
Two details in that code matter:
-
max_tokens=8000leaves room for thinking. A small cap would measure truncation, not effort. - The text block is picked by
type, because a response can start withthinkingblocks.
Sanity check: if output tokens barely change between low and high, the setting is probably not reaching the model. Inspect what your gateway forwards. Per-token price on the AIHubMix Haiku 5.5 page is the same at every level, so the token totals convert directly into cost.
What the published curves look like
Use these to sanity-check your sweep's shape, not as your numbers. Index data: Artificial Analysis. FrontierCode data: Cognition's public result file, compiled by Kingy.ai.
Artificial Analysis Intelligence Index v4.3.2:
| Effort | Index | Output tokens (whole index) | Cost per task | TTFT |
|---|---|---|---|---|
| low | 29 | 32M | $0.02 | 9.9 s |
| medium | 34 | 54M | $0.05 | 13.4 s |
| high | 38 | 97M | $0.08 | 28.0 s |
| xhigh | 41 | 180M | $0.12 | n/a |
| max | 43 | 440M | $0.21 | n/a |
xhigh latency was not published; the max-effort latency figure looked anomalous and is omitted.
FrontierCode 1.1 Main, Haiku 5.5 in Claude Code:
| Effort | Composite | Cost per rollout | Output tokens per rollout |
|---|---|---|---|
| low | 34.8% | $0.06 | 23,868 |
| medium | 41.6% | $0.13 | 36,172 |
| high | 41.9% | $0.26 | 55,149 |
| xhigh | 45.8% | $0.63 | 100,847 |
| max | 46.4% | $1.33 | 181,387 |
What to take from them:
- low → medium is the cheapest step: +5 index points for ~1.7× tokens; +6.8 FrontierCode points for ~2× cost.
- medium → high can be flat: +0.3 on FrontierCode at ~2× cost, though +4 on the broader index.
- high → max is steep: ~4.5× tokens for +5 index points. On FrontierCode, max costs ~10× medium for +4.8 points.
Also note: Anthropic's launch benchmarks were run at max effort. The system card's medium-effort results are lower (GDPval-AA 1277 vs 1620). Compare your default-effort results against those.
Starting points
| Route | Start at |
|---|---|
| Classification, routing, short extraction | low |
| Chat and live support | low (test high for strict rules) |
| Summaries, compaction, subagent work | medium |
| Narrow agentic coding changes | medium |
| Knowledge work, long agent tasks | high |
| Anything needing xhigh or max | Compare with Sonnet 5.5 first |
These follow Anthropic's Haiku 5.5 guidance.
Failure modes and fixes
Symptom: stop_reason: "max_tokens" with no text block.
Cause: thinking consumed the cap.
Fix: raise max_tokens (a 50-token cap for a one-word label is too small now) or lower effort.
Symptom: cache hit rate drops after an effort change.
Cause: changing top-level effort between requests invalidates the messages cache.
Fix: for a single harder turn, use a per-message effort change (beta header mid-conversation-output-config-2026-07-01, Claude API and Google Cloud, thinking must be on). Confirm the gateway forwards the header.
Symptom: 400 with thinking: {"type": "disabled"}.
Cause: disabled thinking is accepted only at low, medium, high.
Fix: lower effort, or switch to adaptive thinking. Note that with thinking off, per-message effort changes also 400, and the model may skip a needed tool call when the request also asks for JSON output.
Symptom: a coding agent at low stops before finishing and hands the task back.
Cause: long agent system prompts at low effort. Anthropic measured that moving to medium roughly halved early stopping and more than doubled output tokens per attempt.
Fix: add the "keep working until done" instruction from the Haiku 5.5 prompting guide, or raise effort.
Symptom: the agent reports a code change as done without running tests.
Cause: happens at low and medium.
Fix: add the guide's verification paragraph; expect more tokens.
Symptom: empty visible reply in a multi-turn chat.
Cause: at xhigh, the model sometimes puts the whole answer in its thinking.
Fix: check for empty text blocks if you run xhigh.
Symptom: no reasoning before a tool call.
Cause: forced tool_choice skips thinking.
Fix: use auto and say in the prompt when to call the tool.
What this post can't verify
- The benchmark tables are third-party and vendor runs on their own workloads; your curve may be flatter or steeper.
- Whether a given gateway passes
output_config.effortand newer beta headers through unchanged. The sweep's token totals are the practical check. - Artificial Analysis's max-effort latency, which was left out as anomalous.
Keep reading: the Claude Haiku 5.5 series
- Effort decides how many tokens you generate. For how Haiku 5.5 compares with Haiku 4.5, GPT-6 Luna, and Sonnet 5.5 at the top of the range, read Haiku 5.5 Went From 0% to 39% on Terminal-Bench at a Tenth of the Price. Here's What That Means for Your Stack
- Thinking tokens are billed as output, and prompt length picks the rate card. To turn your sweep results into a monthly bill, read Claude Haiku 5.5 Has a Price Cliff at 100K Tokens. Here's How to Guard Against It
- The old thinking budget now returns a 400. For that fix and the rest of the switch from Haiku 4.5, read Haiku 4.5 to 5.5: The Request Diff That Fixes Five 400s
Sources
- Effort (Claude Platform Docs)
- Prompting Claude Haiku 5.5 (Claude Platform Docs)
- Claude Haiku 5.5 on AIHubMix
- Claude Haiku 5.5 Intelligence, Performance & Price Analysis (Artificial Analysis)
- Claude Haiku 5.5: Benchmarks, Specs, Sol & Luna Compared (Kingy.ai)
Top comments (0)