Claude Haiku 5.5 costs $0.10 per million input tokens up to 100K tokens of prompt, and $0.50 per million above it, a fivefold cliff that decides your agent bill. Anthropic released the model on 7 October 2026 on its own platform, AWS, Google Cloud and Microsoft Azure the same day. Output is $0.50 per million tokens at or below the threshold and $2.50 above it. The context window is 1M tokens with up to 128K tokens of output.
Short answer: for high-volume agent loops, Haiku 5.5 is the cheapest Anthropic model ever listed, and a loop that keeps every prompt under 100K tokens runs about $1,050 per million calls at 8,000 input tokens and 500 output tokens each. Add prompt caching and that falls to roughly $520. The risk is the 100K line. One careless context-stuffing change moves a call from the cheap tier to the expensive one. Here is the price sheet, the worked budget and the rules that keep you below the line.
The two-tier price sheet
The numbers below come from Anthropic's pricing page as reported by VentureBeat and MarkTechPost on launch day. Anthropic describes Haiku 5.5 as its "best lightweight model yet" and says it beats Haiku 4.5 in coding, computer use and knowledge work. It also says the model is on average 75% cheaper than Haiku 4.5. VentureBeat frames the headline as a 90% API price cut against Haiku 4.5's $1.00 input and $5.00 output, which is what you get when you compare the lower tier alone.
| Item | Prompt up to 100K tokens | Prompt above 100K tokens |
|---|
| Input, per million tokens | $0.10 | $0.50 |
| Output, per million tokens | $0.50 | $2.50 |
| Cache read, per million tokens | $0.01 | not listed in my sources |
| Cache write (5 minute), per million tokens | $0.125 | not listed in my sources |
Read the two columns as a step function, not a slope. A 99,000-token prompt with a 1,000-token answer costs $0.0104. A 101,000-token prompt with the same answer costs $0.053, which is 5.1 times more for 2% more input. That calculation assumes the higher rate applies to the whole request, which is how the tier is listed. Confirm on Anthropic's pricing page whether your region and provider bill it that way before you build a forecast on it.
The cache read price is worth a second look: $0.01 per million is one tenth of the base input price. If most of your prompt is a fixed system prompt, tool definitions and few-shot examples, you pay a tenth for those tokens on every hit. The cache write costs $0.125 per million, a 25% premium over base input, charged when the prefix is first stored.
A worked budget for one million calls
Assume a classification or routing agent loop. Each call sends 8,000 input tokens and returns 500 output tokens. Of the input, 6,000 tokens are a fixed prefix: instructions, tool schemas and examples. The other 2,000 vary per call. All numbers are mine, for illustration. Swap in your own.
| Scenario (1M calls, 8,000 in, 500 out) | Input cost | Output cost | Total |
|---|
| Haiku 5.5, no caching | 8,000M tokens x $0.10 = $800 | 500M x $0.50 = $250 | $1,050 |
| Haiku 5.5, 6,000-token prefix cached | 6,000M x $0.01 + 2,000M x $0.10 = $260 | $250 | $510, plus about $7.50 of cache writes |
| Haiku 4.5 ($1.00 in, $5.00 out) | $8,000 | $2,500 | $10,500 |
| Gemini 3.5 Flash ($1.50 in, $9.00 out) | $12,000 | $4,500 | $16,500 |
The cache write line assumes one prefix write per 100 calls, which is 10,000 writes of 6,000 tokens at $0.125 per million. That is a hot-cache assumption. If your traffic is bursty and the five-minute cache expires between bursts, you pay the write more often, and the saving shrinks.
The ratio is the point. Under identical traffic Haiku 5.5 comes out near one tenth of Haiku 4.5 and under one fifteenth of Gemini 3.5 Flash at its listed $1.50 and $9.00. Gemini 3.5 Flash launched on 19 May 2026 with a cached price of $0.15 and batch pricing of $0.75 and $4.50, so its gap narrows with batch jobs but does not close. A newer Gemini 3.8 Flash is reported at an introductory $0.75 and $3.75 through 31 December 2026, then $1.50 and $7.50. Sources disagree on those figures, so check Google's page before relying on them.
OpenAI's GPT-6 Luna is the direct competitor. VentureBeat describes Haiku 5.5 as priced in line with it. I do not have a verified Luna price sheet from this week, so I am not putting a Luna row in the table. Price your own workload on the AI Model Cost Calculator with both models side by side.
Staying under the 100K line
The 1M context window is a capability, not a budget. Using it flips every call into the higher tier. Three habits keep an agent loop in the cheap lane.
First, cap the prompt in code, not in hope. Count tokens before you send. If a request crosses 90,000 tokens, summarise or drop the oldest turns instead of sending it. A ten percent safety margin costs you little and removes the cliff from your daily risk.
Second, treat tool output as the usual culprit. Agents that read files, web pages or search results grow their context fastest through tool results. Truncate results at the tool boundary and return identifiers the agent can fetch on demand, not whole documents.
Third, route the rare big job elsewhere on purpose. If a task genuinely needs 400K tokens of context, it is a different workload. Price it at the higher tier deliberately or give it to a model built for long inputs. Do not let it arrive by accident through a loop that never trims history.
// guard before every call (pseudo-code)
const tokens = countTokens(messages)
if (tokens > 90_000) messages = compact(messages) // summarise old turns
send(messages)
The new Agent Run Cost Simulator models a multi-step run, not a single call, which is where the tier line bites: step ten of an agent run carries the history of steps one to nine. Run your expected step count through it with the 100K threshold in mind.
What the cheap tier does not fix
Four costs survive the price cut, and they are the ones that surprise teams after the first invoice.
Retries multiply everything. A loop that retries a failed step three times triples the spend on that step, and an agent that wanders can retry silently. Put a hard cap on steps and attempts per run, and log the count per run so you can see the tail. A run that costs 40 times the median is a bug, and it hides inside an average.
Output tokens cost five times input at the lower tier. At $0.50 against $0.10 per million, a chatty model is expensive in a way a terse one is not. In the budget above, output was 24% of the uncached bill but 49% of the cached one. Once you cache the prefix, output is where the money goes, so ask for structured, short answers and set a maximum output length on every call.
Cache misses are silent. The cache read price is one tenth of base input, but only a hit earns it. Change one character early in the prefix, reorder your tool definitions, or insert a timestamp near the top, and the next call pays full price and pays a write on top. Keep the stable part first and the volatile part last, and check the cache hit numbers your provider returns on each response.
Rate limits are a separate ceiling. A million calls a day is about 12 calls a second sustained. Your account limits, not the price, may decide how fast you can run that. Check the limits for your tier before you promise a batch finishes by morning.
None of this argues against the model. It argues for measuring one thing before you scale: cost per successful task, not cost per call. Divide your total spend by the number of tasks that finished correctly. That figure is the one that stays honest when prices, models and prompts change again next month.
Where Haiku 5.5 fits in an agent stack
A model this cheap changes the architecture question from "can I afford a call here" to "which calls do I still send to a bigger model". Haiku 5.5 is a fit for routing and triage, extraction from structured text, log and ticket classification, first-pass code review comments, and the many small steps in a computer-use or browser agent. It is the wrong default for the single hard step that decides whether a run succeeds. For that step, pay for a stronger model and keep Haiku on everything around it.
The honest limit: Anthropic's claim that Haiku 5.5 beats Haiku 4.5 is a vendor claim, and a lower price tells you nothing about your accuracy. Run 200 of your real inputs through both models before you switch, and compare the failures, not the average score. A cheaper model that needs a retry on 15% of calls is not 90% cheaper.
Cost control is the discipline that makes cheap models safe. The AI Agent Ops Bundle covers specs, observability and cost control for exactly this kind of loop, and the Agent Prompt Vault gives you 50 production prompts with stable prefixes you can cache. Fleets of agents also have a coordination bill, which the post The Multi-Agent Tax covers. If your volume is on the OpenAI side, the sibling post on Codex Cloud pricing shows how a rate limit differs from a rate.
Quick answers
How much does Claude Haiku 5.5 cost?
For prompts up to 100K tokens it is $0.10 per million input tokens and $0.50 per million output tokens. Above 100K tokens it is $0.50 input and $2.50 output. Cache reads are $0.01 per million and 5-minute cache writes are $0.125 per million.
What is the Haiku 5.5 context window?
One million tokens, with up to 128K tokens of output. Using more than 100K tokens of prompt moves the call to the higher price tier.
Is Haiku 5.5 really 90% cheaper than Haiku 4.5?
At the lower tier, yes: $0.10 and $0.50 against $1.00 and $5.00. Anthropic's own average figure is 75% cheaper, which is the number to use if some of your calls cross the 100K threshold.
Which models compete with Haiku 5.5 on price?
VentureBeat says OpenAI GPT-6 Luna is priced in line with it. Gemini 3.5 Flash lists at $1.50 input and $9.00 output, so it costs about 15 times more at the same token mix.
Every product mentioned is available at wowhow.cloud — pay once, ship forever.
Originally published at wowhow.cloud
Top comments (0)