DEV Community

Yong Yu
Yong Yu

Posted on Originally published at yongboyu.hashnode.dev

Claude Haiku 5.5 Just Landed: Migrating a Haiku 4.5 Classifier Without the 400s (temperature, Prefill, Thinking Blocks)

Attributed compile + one small offline smoke test

Primary sources (Anthropic, 2026-10-07): Introducing Claude Haiku 5.5 (Anthropic); the Claude Platform docs What's new in Claude Haiku 5.5, Migration guide, and Prompting Claude Haiku 5.5

Secondary: Introducing Claude Haiku 5.5 on AWS by Aamna Najmi, Dani Mitchell, Alfredo Castillo, and Sofian Hamiti (AWS Machine Learning Blog, 2026-10-07)

Prices, benchmark claims, and the "~30% more tokens" figure are Anthropic's. I have no Claude API key on this box, so I did not call Haiku 5.5. What I did run is an offline check of the current Python SDK and a small request linter against hand-built mock responses, described below as exactly that.


Anthropic shipped Claude Haiku 5.5 yesterday, and its launch post on X drew close to 5 million views within a day. The headline is price. Anthropic lists $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100K tokens ($0.50 / $2.50 above that), against $1 / $5 for Haiku 4.5, and says it costs "around 75% less to run" on average. It also gets a 1M-token context window, 128K max output, and, for the first time on a Haiku, an effort setting.

Most coverage is about benchmarks. My view as a backend engineer is narrower: Haiku is the model people put behind ticket routers, intent classifiers, and extraction jobs. Those are exactly the integrations written in the Haiku 3/4.5 style, with temperature=0, a prefilled {"label": ", and max_tokens=20. Every one of those habits now breaks.

What actually changed for a classifier

Anthropic's migration guide lists several breaking changes. Three hit a typical classifier directly, and three more change behavior without an error.

Hard errors (HTTP 400 per the docs):

  1. Sampling parameters. temperature must be omitted or exactly 1. top_p must be omitted or exactly 0.99 (even 1 fails), top_k is rejected at any value, and sending temperature and top_p together also fails. The "set temperature to 0 for deterministic labels" reflex is now an error.
  2. Assistant prefill. A final assistant turn for the model to continue is rejected, even with thinking off. The docs recommend structured outputs or a tool with an enum field for classification instead.
  3. Manual thinking budgets. thinking: {"type": "enabled", "budget_tokens": N} is gone. Use {"type": "adaptive"} and steer it with output_config.effort.

Silent changes:

  1. Thinking is on by default, so a response can start with a thinking block even if you never asked for one. By default that block arrives with an empty thinking field and only a signature. Code that reads content[0].text gets nothing.
  2. Thinking tokens count toward max_tokens. A max_tokens=20 request can stop with stop_reason: "max_tokens" after the thinking block and before any text.
  3. The tokenizer changed. Anthropic says the same text counts as about 30% more tokens than on Haiku 4.5. Token-based budgets, truncation limits, and cost dashboards all shift.

There is also a new failure path: Haiku 5.5 runs safety classifiers that can return stop_reason: "refusal" with a stop_details.category, and the docs say retrying usually refuses again. If you run Priority Tier on Haiku 4.5, the guide also notes Priority Tier isn't supported on 5.5.

The smoke test: what I could check without an API key

SDK check. I created a fresh virtualenv on a Linux box (Python 3.13) and ran pip install anthropic, which resolved to anthropic 1.12.1 today. Inspecting client.messages.create showed thinking, output_config, and tool_choice parameters, and no temperature, top_p, or top_k parameters at all. Calling it with temperature=0 raised:

TypeError: Messages.create() got an unexpected keyword argument 'temperature'
Enter fullscreen mode Exit fullscreen mode

That happens client-side, before any network call. So depending on which SDK version your service pins, the same legacy code fails either as a Python TypeError at call time or, per Anthropic's docs, as a server-side 400. Both look like "the classifier is down," and neither tells you which label you would have gotten. I didn't check older SDK versions or the extra_body escape hatch.

Request linter. I wrote a ~150-line preflight.py that encodes the migration-guide rules and points them at a classic Haiku 4.5 ticket classifier:

legacy = {
    "model": "claude-haiku-4-5",
    "max_tokens": 20,
    "temperature": 0,
    "system": "Classify the support ticket. Reply with JSON only.",
    "messages": [
        {"role": "user", "content": "I was charged twice this month, please refund one."},
        {"role": "assistant", "content": '{"label": "'},   # prefill
    ],
}
Enter fullscreen mode Exit fullscreen mode

The linter flagged the model ID, temperature=0, the prefill, and the 20-token max_tokens (as a warning, since thinking eats into it). The migrated version replaces the prefill with a forced enum tool:

fixed = {
    "model": "claude-haiku-5-5",
    "max_tokens": 1024,
    "system": "Classify the support ticket. Reply with JSON only.",
    "messages": [{"role": "user", "content": "I was charged twice this month, please refund one."}],
    "tools": [{
        "name": "set_label",
        "description": "Record the single best label for the ticket.",
        "input_schema": {"type": "object",
                         "properties": {"label": {"type": "string",
                                                  "enum": ["billing", "technical", "account", "other"]}},
                         "required": ["label"]},
    }],
    "tool_choice": {"type": "tool", "name": "set_label"},
    "output_config": {"effort": "low"},
}
Enter fullscreen mode Exit fullscreen mode

Why a forced tool? Per the migration guide, with a forced tool_choice the response starts with the tool call and carries no thinking block. For a single-label router, that's the behavior you want: no reasoning tokens to pay for, and an answer constrained to your enum. If the label needs judgment, use tool_choice: {"type": "auto"} and let adaptive thinking work.

Response reader. Finally, I built four mock responses by hand, shaped the way the docs describe: thinking-then-text, forced tool call, thinking-then-max_tokens, and a refusal. They validated against the SDK's own Message pydantic type. A type-aware reader handled all four:

def read_answer(resp):
    if resp["stop_reason"] == "refusal":
        return {"status": "refused", "category": (resp.get("stop_details") or {}).get("category")}
    for b in resp["content"]:
        if b["type"] == "tool_use":
            return {"status": "ok", "label": b["input"]["label"]}
    text = "".join(b["text"] for b in resp["content"] if b["type"] == "text")
    if resp["stop_reason"] == "max_tokens" and not text:
        return {"status": "truncated_after_thinking"}
    return {"status": "ok", "text": text}
Enter fullscreen mode Exit fullscreen mode

The Haiku 4.5-era content[0].text returned None on three mocks and raised IndexError on the refusal mock, where I'd given it an empty content list. These are mocks, not model output. They show what your parsing code does with the documented shapes, not how often Haiku 5.5 actually produces each one.

On the cost math

The pricing is compelling, but recompute it with the new token counts instead of reusing Haiku 4.5 numbers. As an illustration using Anthropic's list prices and its ~30% estimate: a prompt measured at 400 input tokens on 4.5 becomes roughly 520 on 5.5. Per million such requests, input cost goes from about $400 to about $52. That's still a large cut, but it's smaller than the per-token price drop suggests. Effort matters too. Anthropic's prompting guide says moving from low to medium roughly halved early stopping in long agent prompts and more than doubled output tokens per attempt. Medium is the default on the API, so a high-volume classifier that never sets effort pays for thinking it may not need.

Run count_tokens with model="claude-haiku-5-5" on a sample of real prompts rather than trusting the 30% figure for your content.

Where it fits, and where something else might

Anthropic and AWS both pitch Haiku 5.5 as the fast layer: routing, classification, extraction, summaries, and subagents under Opus or Sonnet. Anthropic says Sonnet and Opus remain the better choice for complex agentic coding. For pure bounded-choice decisions, it's worth knowing that OpenAI put a Decisions API into public beta the day before, which returns typed predicates, choices, and scores from gpt-6-luna. I haven't tested either. If your workload is "pick one of N labels," evaluate both on your own labeled set before committing.

Migration checklist

  1. Swap the model ID (claude-haiku-5-5; Bedrock uses anthropic.claude-haiku-5-5 or a regional inference profile).
  2. Delete temperature, top_p, and top_k, and check whether your pinned SDK still accepts them.
  3. Replace prefills with an enum tool or structured outputs.
  4. Replace budget_tokens with adaptive thinking plus an explicit effort. Start a classifier at low.
  5. Raise small max_tokens values, or force the tool call so there's no thinking block.
  6. Read content blocks by type, never by position.
  7. Handle stop_reason: "refusal" without blind retries.
  8. Recount tokens and redo cost estimates on real prompts.
  9. Before switching traffic, compare labels from 4.5 and 5.5 on a few hundred real inputs.

Anthropic's guide also mentions a /claude-api migrate skill inside Claude Code that applies these edits across a repo. I didn't try it, but the checklist above tells you what to check in its diff.

The SDK check and the response reader can be reproduced from the snippets above with pip install anthropic and no API key. Again, no live model calls were made.


YongBo Yu is an AI engineer in Toronto building LLM workflows and agent systems. More at yongbo-yu.vercel.app and GitHub.

Top comments (0)