DEV Community

Cover image for Haiku 4.5 to 5.5: The Request Diff That Fixes Five 400s
AIHubMix
AIHubMix

Posted on

Haiku 4.5 to 5.5: The Request Diff That Fixes Five 400s

TL;DR: swapping claude-haiku-4-5 for claude-haiku-5-5 breaks five request patterns with a 400: thinking budgets, sampling parameters, assistant prefill, the old computer use tool, and edited history under replayed thinking blocks. The fixes live in four places in your code: the request builder, the response parser, the conversation store, and the agent loop. Below is the diff for each, plus the changes that fail silently.

Anthropic says Haiku 4.5 prompts should work on Haiku 5.5 without changes. The prompts are fine. The code around them is what needs work. Reference: Anthropic's Haiku 5.5 migration guide.

Prerequisites

  • A current Anthropic Python SDK: pip install -U anthropic. Older versions may not expose output_config as a named argument.
  • An AIHubMix key in AIHUBMIX_API_KEY. Examples use the AIHubMix Claude native endpoint, which takes the Anthropic SDK with base_url="https://aihubmix.com".
  • A handful of real requests from your Haiku 4.5 logs to replay against the new model.

The request builder

Here is a typical Haiku 4.5 classifier call, carrying two of the five breaking patterns (a thinking budget would be the third, but Haiku 4.5 doesn't allow thinking together with temperature changes or prefill):

# Haiku 4.5: two patterns that now return 400
client.messages.create(
    model="claude-haiku-4-5",
    max_tokens=50,
    temperature=0,                                          # 400 on 5.5
    messages=[
        {"role": "user", "content": ticket},
        {"role": "assistant", "content": "{"},              # prefill: 400 on 5.5
    ],
)
Enter fullscreen mode Exit fullscreen mode

And the Haiku 5.5 version:

import os
import anthropic

client = anthropic.Anthropic(
    api_key=os.environ["AIHUBMIX_API_KEY"],
    base_url="https://aihubmix.com",
)

r = client.messages.create(
    model="claude-haiku-5-5",
    max_tokens=2000,                     # room for thinking plus the JSON
    output_config={
        "effort": "low",                 # thinking is on by default; keep it light
        "format": {                      # replaces the prefill and temperature=0
            "type": "json_schema",
            "schema": {
                "type": "object",
                "properties": {
                    "label": {"type": "string", "enum": ["billing", "bug", "other"]}
                },
                "required": ["label"],
                "additionalProperties": False,
            },
        },
    },
    messages=[{"role": "user", "content": ticket}],
)
Enter fullscreen mode Exit fullscreen mode

What changed, line by line:

  • Thinking budget → effort. Haiku 5.5 supports adaptive thinking only. Send thinking: {"type": "adaptive"} or omit it, and set output_config.effort. Where the old budget was small to save tokens, pick a low level. The default is medium; set it explicitly.
  • Sampling parameters → removed. Only the defaults pass: temperature of 1, top_p of 0.99. Any other value, any top_k, top_p of 1, or both temperature and top_p together returns a 400, with or without thinking. Grep your SDK wrappers and gateway configs for injected defaults too.
  • Prefill → structured output. A final assistant turn returns a 400 even with thinking off. Format control moves to output_config.format; a preamble-skipping prefill becomes a system instruction to answer directly; a continuation moves into the user message ("Your previous response ended with [text]. Continue from there.").
  • max_tokens raised. Thinking counts toward the cap. At 50, thinking alone can use it up, and you get stop_reason: "max_tokens" with no text.

The response parser

Thinking is on by default, so responses can start with thinking blocks. On Haiku 5.5 those blocks come back with empty text and only a signature (Haiku 4.5 returned summaries).

# Breaks: the first block may be thinking
answer = r.content[0].text

# Works
if r.stop_reason == "refusal":
    raise RuntimeError(f"declined: {r.stop_details}")
answer = next(b.text for b in r.content if b.type == "text")
Enter fullscreen mode Exit fullscreen mode

The refusal branch is new. Haiku 5.5 runs safety classifiers in four categories (cyber, bio, frontier LLM development, general harms) and declines with an HTTP 200, stop_reason: "refusal", and a category in stop_details. There is no server-side fallback for this model: a list of fallback models returns a 400, and the default fallback mode leaves the request declined. Your code decides whether to rephrase, escalate to a larger model, or fail.

If your UI showed reasoning summaries, request them with thinking: {"type": "adaptive", "display": "summarized"}.

The conversation store

Two rules for anything that stores and replays messages:

  1. Append only. A thinking block is valid only while everything before it (system prompt, tool list, earlier messages) is unchanged. Replay it after an edit and you get a 400. Enforcement is on by default for accounts created on or after August 31, 2026; older accounts get it only when a request opts in. Typical culprits: a timestamp in the system prompt, a tool list that grows when a plugin connects, client-side truncation, reminders injected and later stripped. Haiku 5.5 accepts role: "system" messages inside messages with no beta header, which is the append-only way to add per-turn instructions.
  2. Keep empty thinking blocks. Pass them back unchanged with tool results. A serializer that drops blocks with empty text drops the reasoning.

Also: thinking blocks are bound to the API account that produced them. Replay through a different account and they are silently dropped; the request succeeds without that reasoning. If you escalate a conversation to Sonnet 5.5 or Opus 5.5, check the preserved thinking docs to confirm whether the earlier blocks carry over.

The agent loop

Computer use. On the Claude API and Google Cloud, computer_20250124 returns a 400. Migrate to computer_toolset_20260801:

  • drop the computer-use-2025-01-24 beta header;
  • declare {"type": "computer_toolset_20260801"};
  • dispatch on each tool_use block's name and toolset_name, not input.action;
  • handle every such block in a turn and echo toolset_name on results;
  • zoom is on by default, so disable it in the toolset config if your environment doesn't implement it.

On Amazon Bedrock, check the tool's compatibility notes before picking a version. The same toolset family adds browser use, which Haiku 4.5 didn't have.

Mid-turn user input. Haiku 5.5 is trained to resist injection through tool results. User text placed inside a tool_result block can be ignored as untrusted. Put it in a text block after the last tool result; keep harness notices in a separate system message.

Forced tool calls. Forced tool_choice still works but skips thinking. For tools with side effects, use auto and say in the prompt when to call the tool.

Low-effort behaviour. With a long coding-agent system prompt at low, the model sometimes stops early and hands the task back; at low and medium it sometimes reports a change as done without running a test. The Haiku 5.5 prompting guide has short instructions for both.

Search tools. Put today's date in the system prompt or the tool description.

Failure modes that don't throw

Symptom Cause Fix
Empty or wrong answer text Thinking block comes first Select blocks by type
Reply cut off, no text Thinking used max_tokens Raise the cap or lower effort
Usage and bills up about 30% New tokenizer Recount on the new model
More output tokens than before Thinking on, default medium Set effort explicitly
Replayed reasoning lost Different API account Replay via the original account
Priority Tier traffic rejected Not supported on Haiku 5.5 Plan capacity separately

The tokenizer row deserves a number. The same text produces about 30% more tokens, which affects usage, count_tokens, context budgets, tuned max_tokens values, and pricing: the 100K-token rate-card line sits at roughly 77K tokens as Haiku 4.5 counted them. Recount real prompts with model="claude-haiku-5-5" before reading anything into a cost dashboard.

Separately, per Anthropic's migration guide, the minimum cacheable prompt drops from 4,096 tokens to 512, and earlier-turn thinking blocks stay in the cached prefix by default.

What we couldn't verify

  • Gateway passthrough. Whether a gateway forwards every output_config field and newer beta headers unchanged is not documented per route. Check on your first run: replay the same prompt at low and high and compare usage.output_tokens. If nothing moves, effort isn't arriving.
  • Refusal rates on your traffic. The four classifier categories are documented; how often they fire on a given workload isn't. Per the launch post, the cyber safeguards allow more defensive work than Sonnet 5.5's but block penetration testing. Log refusals by category for the first week.

Once the migrated routes pass your evals, the AIHubMix model list lets the same client point at Sonnet 5.5 for the task types that keep failing on Haiku.

Keep reading: the Claude Haiku 5.5 series

Sources

Top comments (0)