Anthropic released Claude Sonnet 5.5 on September 28, 2026, the second model in the Claude 5.5 family after Opus 5.5. It keeps Sonnet 5's prices. Anthropic calls it "our fastest Sonnet model to date," with output more than 30% faster than Sonnet 5 and a cost per task up to 30% lower, because it needs fewer tokens for the same work. The API ID is claude-sonnet-5-5, with no date suffix.
The benchmark table gets the attention. For a team with Sonnet 5 in production, the migration guide matters more. It lists five settings that return a 400 on Sonnet 5.5: thinking budgets, sampling parameters, assistant prefill, forced tool choice and thinking: {"type": "disabled"}. Code written for Sonnet 5 trips on the last two.
The numbers
- Terminal-Bench 4.0, up from 10.3% on Sonnet 5
- 70.6%
- CursorBench 4.0, Opus 5.5 scores 57.8%
- 55.5%
- context window, 128K max output
- 1M
- input and output per million tokens
- $2 / $10
| Item | Sonnet 5.5 | Opus 5.5 |
|---|---|---|
| Input | $2 | $4 |
| Output | $10 | $20 |
| Cache writes | $2.50 | $5 |
| Cache reads | $0.20 | $0.20 |
On FrontierCode 1.1, which checks whether an agent's change would merge without human edits, Sonnet 5.5 scored 46.2% at Max effort against 54.4% for Opus 5.5. Anthropic's footnote says the model scored lower at Max than at Xhigh, because it more often ran a review skill that fanned out to subagents and made edits outside the task. Anthropic and its external testers still rate Opus 5.5 stronger at open-ended work that needs sustained judgment.
Effort levels run Low, Medium, High, Xhigh and Max. The Claude apps and Claude Code default to Medium, and the Claude Platform API defaults to High. Anthropic reports that Sonnet 5.5 at Low effort on CursorBench beats Sonnet 5's best score at under a tenth of the cost per task, so test the low settings first if you pay per call.
Thinking is on unless you say otherwise
A request with no thinking field runs with adaptive thinking on Sonnet 5.5. That matches Sonnet 5. Sonnet 4.6 and earlier ran without thinking, so a jump from 4.6 changes your token bill and the shape of the response.
Sonnet 5.5 also removed thinking: {"type": "disabled"}. To skip up-front thinking you send the new lowest setting:
{
"model": "claude-sonnet-5-5",
"max_tokens": 4096,
- "thinking": { "type": "disabled" },
+ "thinking": { "type": "between_tools" },
"output_config": { "effort": "medium" },
"messages": [...]
}
between_tools comes with rules. It works at low, medium and high effort and returns a 400 at xhigh or max. It accepts no other field, so display or budget_tokens beside it fails. With it set, a per-message effort that differs from the one in effect also fails. The model still writes short updates between tool calls, and those arrive as thinking blocks, so code that reads text blocks alone will miss them.
Forced tool use is gone
Sonnet 5.5 rejects tool_choice of type any or tool, including on the token counting endpoint. The error text reads tool_choice: type "tool" and "any" are not supported for this model. The guide's replacement is auto plus a strict tool:
-"tool_choice": { "type": "tool", "name": "extract_invoice" },
+"tool_choice": { "type": "auto" },
"tools": [{
"name": "extract_invoice",
+ "strict": true,
"input_schema": { "type": "object", "additionalProperties": false, ... }
}]
With auto, the model can answer without calling the tool, so your prompt has to say when to use it. Strict tools need additionalProperties: false on each object. On Amazon Bedrock, structured outputs aren't available for Sonnet 5.5, so you send auto without strict and validate the input in your own code.
Safeguards you might hit
Sonnet 5.5 is the first Sonnet to ship with cyber safeguards like the ones on Opus 5.5. Anthropic says routine bug finding and fixing stay unaffected. Requests it rates as higher-risk security work fall back to Sonnet 5, and the fallback shows in the response, so security tooling built on Claude should log which model answered. The release also expands preserved thinking, which ties thinking to the account that created it. Moving a conversation between accounts, including a mid-session account switch in Claude Code, now needs the steps in Anthropic's docs.
This week
- Grep your code for
"disabled"thinking andtool_choicewithanyortool. Those two will break first. - Switch the ID in a staging config, then read responses by block
typeinstead ofcontent[0].text. - Rerun your effort sweep. Start at Low or Medium, and raise effort where your evals drop.
- Re-baseline cost per task, since the per-token price stayed flat and the token count per task is what moved.
If you need forced tool calls and can't move to strict tools yet, stay on Sonnet 5 until you can.
Volodymyr Chornous

