Anthropic released Claude Haiku 5.5 on October 7, 2026, nine days after Sonnet 5.5. The company calls it "the cheapest, fastest, and most capable small model we've ever released." The API ID is claude-haiku-5-5, a fixed ID with no date suffix and no separate alias. It runs on the Claude Platform, Amazon Bedrock, Google Cloud and Microsoft Foundry.
The launch post sells it for summaries, classification, compaction and subagent work under Opus 5.5 or Sonnet 5.5. If you run Haiku 4.5 in a queue worker today, the price cut is the headline. The migration guide holds the part that breaks your code.
Two price tiers
Haiku 5.5 splits its price at 100,000 prompt tokens. Haiku 4.5 had one price.
| Item | Haiku 5.5 | Haiku 4.5 |
|---|---|---|
| Input | $0.10 / $0.50 | $1.00 |
| Output | $0.50 / $2.50 | $5.00 |
| Cache writes | $0.125 / $0.625 | $1.25 |
| Cache reads | $0.01 / $0.05 | $0.10 |
Anthropic puts the cut at 90% for requests up to 100K tokens and 50% above that. It also says about 90% of Haiku 4.5 requests fell under the threshold. The same announcement halved Sonnet 5.5's cache-read price to $0.10.
The tokenizer eats into those savings. Haiku 5.5 uses the tokenizer from Claude 4.7 and later, and the guide says the same text produces about 30% more tokens than on Haiku 4.5. Images cost more too: Haiku 5.5 uses the high-resolution tier, so a 2,000 by 1,500 pixel image costs about 2.5 times the visual tokens it did on 4.5.
A rough sketch of my own, for a 20,000-token prompt with a 1,000-token answer and no thinking: Haiku 4.5 charges $0.025. With 30% more tokens on both sides, Haiku 5.5 charges about $0.0033. Thinking tokens bill as output, so your number lands higher. Measure it.
The numbers
- Terminal-Bench 4.0, Haiku 4.5 scored 0.0%
- 39.2%
- OSWorld 2.1 offline subset, up from 15.7%
- 72.4%
- more tokens for the same text
- ~30%
Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0, and Anthropic still points you to Sonnet and Opus for complex agentic coding. Haiku 5.5 is also the first Haiku with effort levels, from Low through Max.
Five changes that return a 400
The migration guide lists eleven items for Haiku 4.5 users. These fail outright:
| Haiku 4.5 setting | Haiku 5.5 replacement |
|---|---|
| thinking type "enabled" with budget_tokens | thinking type "adaptive" plus output_config.effort |
| temperature other than 1, top_p other than 0.99, any top_k | remove them and steer with the prompt |
| a final assistant turn (prefill) | end with a user turn, use structured outputs for format |
| computer_20250124 tool | computer_toolset_20260801 on the Claude API and Google Cloud |
| edited system, tools or earlier messages with thinking blocks sent back | keep the conversation append-only |
Sending temperature and top_p together also fails, even at their defaults. The append-only check applies in full to accounts created after August 31, 2026. Older accounts get the error after they opt in through thinking.block_binding.prefix_mismatch_behavior.
The thinking change is the one most Haiku code needs:
{
- "model": "claude-haiku-4-5",
+ "model": "claude-haiku-5-5",
"max_tokens": 16000,
- "thinking": { "type": "enabled", "budget_tokens": 8000 },
+ "thinking": { "type": "adaptive" },
+ "output_config": { "effort": "low" },
"messages": [{ "role": "user", "content": "..." }]
}
Behavior that changes without an error
Adaptive thinking is on by default. A request with no thinking field can come back with thinking blocks first, so code that reads content[0].text breaks. Those blocks arrive with an empty thinking field and a signature. Set "display": "summarized" if you log them.
Thinking tokens count toward max_tokens. A classifier tuned with max_tokens: 50 can stop with stop_reason: "max_tokens" before it writes any text. Raise the limit or drop to Low effort.
Thinking blocks from earlier turns now stay in context as input tokens. Haiku 4.5 dropped all but the latest turn's. Long chats grow faster than the tokenizer alone explains.
Forced tool_choice still works on Haiku 5.5, unlike on Sonnet 5.5. The response starts with the tool call and skips thinking. If you want the model to reason first, switch to auto and tell it in the prompt when to call the tool.
Haiku 5.5 runs safety classifiers and can return stop_reason: "refusal" with no fallback model. Priority Tier doesn't cover it, and Bedrock doesn't offer structured outputs for it.
This week
- Grep for
budget_tokens,temperature,top_pand requests that end on an assistant turn. Those four fail first. - Run your prompts through token counting with
model: "claude-haiku-5-5"and recompute cost from those counts, not Haiku 4.5's. - Read responses by block
type, and handlerefusalas a stop reason. - Start your effort sweep at Low for classification and summaries, and raise it where your evals drop.
If you run Haiku 4.5 on Priority Tier and depend on that capacity, stay put until you've planned around it. In Claude Code, /claude-api migrate this project to claude-haiku-5-5 applies the mechanical changes and hands you a checklist for the rest.
Volodymyr Chornous

