DEV Community

Cover image for Tuning Claude Haiku 5.5's Effort Setting: A Sweep Script and the Failure Modes to Watch
AIHubMix
AIHubMix

Posted on

Tuning Claude Haiku 5.5's Effort Setting: A Sweep Script and the Failure Modes to Watch

TL;DR: On Claude Haiku 5.5, the effort setting changes output volume far more than anything else in the request. Independent runs show max effort using about 14 times the output tokens of low for a 14-point gain on the index. Default is medium. Run a sweep on your own prompts, start simple routes at low, and watch for four failure modes: truncated replies, cache resets, early stopping, and empty replies at xhigh.

The engineering problem

Claude Haiku 5.5 changes three things that affect token spend, all at once:

  • Thinking is on by default (adaptive). Haiku 4.5 only thought when you asked.
  • Thinking tokens are output tokens: they count toward max_tokens and bill at the output rate ($0.50 per million for prompts up to 100K tokens).
  • The default effort is medium, one level below most current Claude models, which default to high.

The old control, thinking: {"type": "enabled", "budget_tokens": N}, now returns a 400. Effort replaces it, so there's no number to carry over. And prompting the model to "answer directly" doesn't reduce thinking; in Anthropic's testing it kept thinking anyway. Effort is the only real lever, so it needs measuring.

The five levels are low, medium, high, xhigh, max, set in output_config (see Anthropic's effort docs).

Prerequisites

  • An AIHubMix API key in AIHUBMIX_API_KEY.
  • A current Anthropic SDK: pip install -U anthropic.
  • 20–50 real prompts from production, plus a way to grade the outputs (a rubric, expected labels, or tests).

The script uses the AIHubMix Claude native endpoint, because effort lives in the Messages API's output_config.

The sweep

import os
import time
import anthropic

client = anthropic.Anthropic(
    api_key=os.environ["AIHUBMIX_API_KEY"],
    base_url="https://aihubmix.com",
)

prompts = [
    "Summarize this support ticket in one sentence: ...",
    "Extract the invoice number and total from: ...",
]

for effort in ["low", "medium", "high"]:
    out_tokens, seconds = 0, 0.0
    for prompt in prompts:
        start = time.time()
        r = client.messages.create(
            model="claude-haiku-5-5",
            max_tokens=8000,
            output_config={"effort": effort},
            messages=[{"role": "user", "content": prompt}],
        )
        seconds += time.time() - start
        out_tokens += r.usage.output_tokens
        text = next((b.text for b in r.content if b.type == "text"), "")
        # Save `text` next to the prompt and grade it against your own rubric.
    print(f"{effort}: {out_tokens} output tokens, {seconds:.1f}s total")
Enter fullscreen mode Exit fullscreen mode

Two details in that code matter:

  • max_tokens=8000 leaves room for thinking. A small cap would measure truncation, not effort.
  • The text block is picked by type, because a response can start with thinking blocks.

Sanity check: if output tokens barely change between low and high, the setting is probably not reaching the model. Inspect what your gateway forwards. Per-token price on the AIHubMix Haiku 5.5 page is the same at every level, so the token totals convert directly into cost.

What the published curves look like

Use these to sanity-check your sweep's shape, not as your numbers. Index data: Artificial Analysis. FrontierCode data: Cognition's public result file, compiled by Kingy.ai.

Artificial Analysis Intelligence Index v4.3.2:

Effort Index Output tokens (whole index) Cost per task TTFT
low 29 32M $0.02 9.9 s
medium 34 54M $0.05 13.4 s
high 38 97M $0.08 28.0 s
xhigh 41 180M $0.12 n/a
max 43 440M $0.21 n/a

xhigh latency was not published; the max-effort latency figure looked anomalous and is omitted.

FrontierCode 1.1 Main, Haiku 5.5 in Claude Code:

Effort Composite Cost per rollout Output tokens per rollout
low 34.8% $0.06 23,868
medium 41.6% $0.13 36,172
high 41.9% $0.26 55,149
xhigh 45.8% $0.63 100,847
max 46.4% $1.33 181,387

What to take from them:

  • low → medium is the cheapest step: +5 index points for ~1.7× tokens; +6.8 FrontierCode points for ~2× cost.
  • medium → high can be flat: +0.3 on FrontierCode at ~2× cost, though +4 on the broader index.
  • high → max is steep: ~4.5× tokens for +5 index points. On FrontierCode, max costs ~10× medium for +4.8 points.

Also note: Anthropic's launch benchmarks were run at max effort. The system card's medium-effort results are lower (GDPval-AA 1277 vs 1620). Compare your default-effort results against those.

Starting points

Route Start at
Classification, routing, short extraction low
Chat and live support low (test high for strict rules)
Summaries, compaction, subagent work medium
Narrow agentic coding changes medium
Knowledge work, long agent tasks high
Anything needing xhigh or max Compare with Sonnet 5.5 first

These follow Anthropic's Haiku 5.5 guidance.

Failure modes and fixes

Symptom: stop_reason: "max_tokens" with no text block.
Cause: thinking consumed the cap.
Fix: raise max_tokens (a 50-token cap for a one-word label is too small now) or lower effort.

Symptom: cache hit rate drops after an effort change.
Cause: changing top-level effort between requests invalidates the messages cache.
Fix: for a single harder turn, use a per-message effort change (beta header mid-conversation-output-config-2026-07-01, Claude API and Google Cloud, thinking must be on). Confirm the gateway forwards the header.

Symptom: 400 with thinking: {"type": "disabled"}.
Cause: disabled thinking is accepted only at low, medium, high.
Fix: lower effort, or switch to adaptive thinking. Note that with thinking off, per-message effort changes also 400, and the model may skip a needed tool call when the request also asks for JSON output.

Symptom: a coding agent at low stops before finishing and hands the task back.
Cause: long agent system prompts at low effort. Anthropic measured that moving to medium roughly halved early stopping and more than doubled output tokens per attempt.
Fix: add the "keep working until done" instruction from the Haiku 5.5 prompting guide, or raise effort.

Symptom: the agent reports a code change as done without running tests.
Cause: happens at low and medium.
Fix: add the guide's verification paragraph; expect more tokens.

Symptom: empty visible reply in a multi-turn chat.
Cause: at xhigh, the model sometimes puts the whole answer in its thinking.
Fix: check for empty text blocks if you run xhigh.

Symptom: no reasoning before a tool call.
Cause: forced tool_choice skips thinking.
Fix: use auto and say in the prompt when to call the tool.

What this post can't verify

  • The benchmark tables are third-party and vendor runs on their own workloads; your curve may be flatter or steeper.
  • Whether a given gateway passes output_config.effort and newer beta headers through unchanged. The sweep's token totals are the practical check.
  • Artificial Analysis's max-effort latency, which was left out as anomalous.

Keep reading: the Claude Haiku 5.5 series

Sources

Top comments (0)