On September 28, 2026, Anthropic released Claude Sonnet 5.5. According to the announcement, it is the second model in the Claude 5.5 family, following Opus 5.5 six days earlier, and the previous Sonnet 5 (released June 30, 2026) has moved to Legacy (older, but still available). The official model overview describes it as the model with the best combination of speed and intelligence.

The short version: pricing is exactly the same as Sonnet 5 ($2 input, $10 output and $0.20 cache reads per million tokens), and in Anthropic's comparison table it improves substantially on Sonnet 5 in most rows. The API behaves differently, though: there are five changes that make code that worked on Sonnet 5 fail with a 400 error. Most notably, thinking: {"type": "disabled"} no longer works for turning thinking off, and its replacement, between_tools, is accepted only at effort high or below. This article checks the official documentation and the announcement against each other in the original text and walks through, in the order you will need it, what changed, what breaks when you migrate, and how to choose between it and Opus 5.5.

Information as of September 29, 2026: written the day after launch, after reading the original text of the official Claude Platform documentation (Models overview, the Sonnet 5.5 page, What's new, the migration guide, Pricing, Thinking, Effort and Prompt caching), Anthropic's announcement, the Claude Code documentation and CHANGELOG, and the GitHub changelog. Claude Haiku 5.5, which the announcement says is coming within weeks, had not been released at this point.

CLAUDE SONNET 5.5 — 2026.09.28

Same price, five breaking changes

— Officially positioned as the best combination of speed and intelligence

Model ID
claude-sonnet-5-5
No date suffix (official model overview)
Pricing (per million tokens)
$2 / $10
Same as Sonnet 5 (official pricing page)
Default effort
high on the API
medium in Claude Code and the apps (official docs and announcement)
Watch out
5 breaking changes
Listed officially as breaking changes
Source: Anthropic's official announcement and the Claude Platform documentation (September 28, 2026)

1. Sonnet 5.5 at a glance: performance, pricing, migration pitfalls

① Performance

In Anthropic's comparison table it beats Sonnet 5 in every row and comes within 2 points of Opus 5.5 on GDPval-AA. The announcement itself, however, says Opus 5.5 remains clearly stronger on complex, open-ended work.

② Pricing

Every rate is identical to Sonnet 5. What changed is the minimum cacheable length, which dropped from 1,024 to 512 tokens. The announcement says it is up to 30% cheaper per task, but that is Anthropic's own measurement.

③ Migration pitfalls

Five changes: disabling thinking moves to between_tools / forced tool use returns a 400 / thinking blocks are bound to the model and the conversation / the legacy computer use tool is rejected / some advisor pairings are rejected. On top of that, text between tool calls comes back in thinking blocks.

In one sentence, Sonnet 5.5 is a Sonnet that moved up a tier at the same price, handled with almost the same conventions as Opus 5.5. Three of the five breaking changes (forced tool use, thinking-block binding and the legacy computer use tool) also apply to Opus 5.5 and Fable 5.1. What is specific to Sonnet 5.5 comes down to two things: a way to cut thinking survives as a separate value, between_tools, and there are restrictions on advisor tool pairings.

2. Core specs and availability

The comparison covers three models: Sonnet 5, which it replaces; Opus 5.5 above it; and Haiku 4.5 below it.

Item Sonnet 5.5 Sonnet 5 (Legacy) Opus 5.5 Haiku 4.5
API model ID claude-sonnet-5-5 claude-sonnet-5 claude-opus-5-5 claude-haiku-4-5-20251001
Pricing (input/output) $2 / $10 $2 / $10 $4 / $20 $1 / $5
Context / max output 1M / 128K 1M / 128K 1M / 128K 200K / 64K
Thinking Adaptive thinking by default (minimum is between_tools) Adaptive thinking by default (can be turned off with disabled) Adaptive thinking always on (cannot be turned off) Extended thinking (budget-based)
Default effort on the API high high medium Not supported
Reliable knowledge cutoff June 2026 January 2026 June 2026 February 2025
Minimum cacheable length 512 tokens 1,024 tokens 512 tokens 4,096 tokens
Speed (official relative label) Fast — Moderate Fastest
Retirement Not before September 28, 2027 Not before June 30, 2027 Not before September 22, 2027 Not before October 15, 2026

Sources: Anthropic, "Models overview", "Claude Sonnet 5.5", "Claude Sonnet 5" and "Prompt caching" (checked September 29, 2026). Speed is a relative label within the current lineup and is not listed for the Legacy Sonnet 5. Retirement dates are commitments for platforms Anthropic operates; Amazon Bedrock and Google Cloud set their own.

Only three rows differ from Sonnet 5: thinking, knowledge cutoff and minimum cacheable length. Context, max output and the API's default effort are the same, and the tokenizer is also the same as Sonnet 5, so the same text produces the same number of tokens (What's new). On the Message Batches API, adding the beta header output-300k-2026-03-24 raises the output limit to 300K tokens (the same as Sonnet 5). Also note that sending non-default values for temperature, top_p or top_k returns a 400. That was already true on Sonnet 5, so it only matters if you are moving directly from Sonnet 4.6 or earlier.

Where it's available

API / cloud

Claude API (claude-sonnet-5-5), Amazon Bedrock (anthropic.claude-sonnet-5-5), Claude Platform on AWS, Google Cloud and Microsoft Foundry. It shipped on every platform on launch day.

Apps / developer tools

claude.ai and the Claude apps (Anthropic has published the system prompt for Sonnet 5.5), Claude Code (v2.1.284 and later), and GitHub Copilot (Pro, Pro+, Max, Business and Enterprise).

What it doesn't have

Fast mode (the faster variant) covers only Opus 5.5, Opus 5 and Opus 4.8 on the pricing page; Sonnet 5.5 does not have it. On Bedrock, structured outputs (including strict tool use) are not available for Sonnet 5.5 (migration guide).

3. Pricing: the same as Sonnet 5, except for the cache minimum

The official What's new page says it is priced the same as Sonnet 5, with prompt caching and batch processing priced the same as well. Here are the detailed rates.

Per million tokens Sonnet 5.5 Sonnet 5 Opus 5.5 Haiku 4.5
Input $2 $2 $4 $1
Output $10 $10 $20 $5
Cache write (5 min) $2.50 $2.50 $5 $1.25
Cache write (1 hour) $4 $4 $8 $2
Cache read $0.20 $0.20 $0.20 $0.10
Batch API (input/output) $1 / $5 $1 / $5 $2 / $10 $0.50 / $2.50

Source: Anthropic, "Pricing" (checked September 29, 2026). The Batch API is 50% off standard rates for both input and output.

Two things that can change within "the same price"

Identical rates don't guarantee an identical bill. Two factors can change it.

The first is the minimum cacheable length. On Sonnet 5, prompts shorter than 1,024 tokens were not cached even with cache_control. On Sonnet 5.5 this floor has dropped to 512 tokens. For example, a job that sends an 800-token system prompt plus tool definitions on every call was not eligible for caching on Sonnet 5, but it is cached on Sonnet 5.5. You can tell whether something was cached from the response's usage: if cache_creation_input_tokens and cache_read_input_tokens are both 0, nothing was cached (official Prompt caching page). Falling short of the minimum doesn't raise an error, so while you migrate it's worth checking for jobs where caching was silently not taking effect.

The second is the number of tokens per task. The announcement says it needs far fewer tokens for the same work, that in Anthropic's testing it is up to 30% cheaper per task, and that it generates output more than 30% faster than Sonnet 5. Those are Anthropic's own measurements. At the same time, the official documentation says that the effort levels have been recalibrated, so the same level won't necessarily think as much as it did on Sonnet 5. Thinking tokens are billed as output tokens even when they aren't displayed. After migrating, the only way to know is to measure usage and compare on your own workload.

Worked example: one task with 10 million cache-read, 500K input and 300K output tokens (token counts are this article's assumption; write costs omitted)

  • Sonnet 5.5: $2.00 + $1.00 + $3.00 = $6.00 (the same on Sonnet 5)
  • Opus 5.5: $2.00 + $2.00 + $6.00 = $10.00 (about 1.7× Sonnet 5.5, not 2×)
  • Haiku 4.5: $1.00 + $0.50 + $1.50 = $3.00 (exactly half of Sonnet 5.5)

This compares rates at the same token counts; in practice each model uses a different number of tokens. Rates are from Anthropic's "Pricing."

The gap to Opus 5.5 isn't "2×" because cache reads alone cost the same $0.20 on Opus 5.5. The higher the cache share of an agentic workload, the less you save by choosing Sonnet 5.5. Pricing across all Claude models is covered in our Opus, Sonnet and Haiku pricing comparison.

4. Benchmarks: reading them within Anthropic's table

The announcement's comparison table has four columns: Sonnet 5.5, Sonnet 5, Opus 5.5 and GPT-6 Sol. Here are the table's conditions first.

  • The table comes from Anthropic's announcement, and most rows are Anthropic's own measurements. The exceptions are GDPval-AA v2.1 and AA-Briefcase v1.1, which were run by Artificial Analysis (footnote 3).
  • Artificial Analysis ran those tests in a pre-release environment that had a bug which could degrade responses to requests using structured outputs. Anthropic notes that any effect would be small and would make the scores look lower, and says the bug has been fixed.
  • Opus 5.5's Terminal-Bench 4.0 figure is at xhigh and is Opus 5.5's best score (footnote 1). Sonnet 5.5's FrontierCode has two figures: 46.2% at max and 52.1% at xhigh (footnote 2).
  • GPT-6 Sol's GDPval-AA, AA-Briefcase and Chartography scores carry a note that they may predate OpenAI's fix for an image-understanding bug (footnote 4).
Benchmark Sonnet 5.5 Sonnet 5 Opus 5.5 GPT-6 Sol
Terminal-Bench 4.0
Agentic coding in the terminal
70.6% 10.3% 66.4% (xhigh) —
FrontierCode 1.1 (Main)
Whether changes get merged
52.1% (xhigh)
46.2% (max)
42.4% 54.4% 49.3%
CursorBench 4.0
Ambiguous multi-file tasks
55.5% 34.1% 57.8% —
GDPval-AA v2.1 (Elo)
Real work across 44 occupations (run by Artificial Analysis)
1844 1449 1846 1487
AA-Briefcase v1.1 (Elo)
Long-horizon knowledge work (run by Artificial Analysis)
1811 1359 1822 1483
Humanity's Last Exam
Cross-disciplinary reasoning (with tools)
64.5% 54.9% 67.7% —
OSWorld 2.1
Computer use (the table notes "partial")
80.1% 57.0% 81.8% —
Chartography
Reading charts (without tools)
61.6% 15.6% 64.4% 53.6%

Source: the comparison table and footnotes in Anthropic's "Introducing Claude Sonnet 5.5" (September 28, 2026; checked September 29). Bold marks the highest value in each row; "—" means the table has no value. Details of the methodology are in the Sonnet 5.5 system card linked from the same announcement.

Three things stand out.

  • The jump from Sonnet 5 is large. Terminal-Bench 4.0 went from 10.3% to 70.6%, and Chartography (without tools) from 15.6% to 61.6%: several times higher within the same table. GDPval-AA rose by about 400 points.
  • The gap to Opus 5.5 is within a few points in most rows. GDPval-AA is 2 points apart, CursorBench 2.3 points and OSWorld 2.1 1.7 points. On Terminal-Bench 4.0, Sonnet 5.5 beats Opus 5.5's best score (66.4% at xhigh).
  • Anthropic still places Opus 5.5 above it. The announcement says that benchmarks capture only one side of capability, and that both internally and among outside testers, Opus 5.5 remains clearly stronger on complex, open-ended work that requires sustained judgment.

Cost by effort level: from the announcement's chart descriptions

The announcement also has charts plotting score against cost per task for each effort level. Here is what their descriptions say.

Terminal-Bench 4.0

At medium, the default in the Claude apps, it clearly beats Sonnet 5's best score at less than a tenth of the cost per task.

FrontierCode 1.1

At high, the Claude Platform default, it matches GPT-6 Sol's best score at about a fifth of the cost per task. It is 10 points above Sonnet 5 at the same high level, at about a fifteenth of the cost.

CursorBench 4.0

At low, the lowest level, it beats Sonnet 5's best score at less than a tenth of the cost per task.

AA-Briefcase v1.1

At medium, it beats Sonnet 5's best score at about a ninth of the cost per task.

Source: chart descriptions in Anthropic's "Introducing Claude Sonnet 5.5." The Terminal-Bench and CursorBench charts show GPT-5.6 Sol because GPT-6 Sol's scores have not been published (note on the same page).

The two FrontierCode figures deserve attention. Sonnet 5.5 scores lower at max (46.2%) than at xhigh (52.1%). According to the footnote in Anthropic's announcement, at max it more often ran Claude Code's code-review skill and split work across many subagents, which in some cases led to timeouts and edits outside the task's scope. Raising effort doesn't always raise the score, so pick the level by comparing on your own work (Section 7). GPT-6 Sol's numbers are covered in our GPT-6 Sol and Luna release overview.

Don't mix this with the Opus 5.5 announcement's table: even for the same Opus 5.5, Chartography is listed at 89.0% "with tools" in the Opus 5.5 announcement and at 64.4% "without tools" in this table, under different conditions. If you pick numbers from the two announcements and put them side by side, you end up comparing values measured under different conditions. Keep comparisons to columns within the same table.

5. Five breaking changes that return a 400 when you move from Sonnet 5, and the fixes

The official What's new in Claude Sonnet 5.5 lists five breaking changes that affect code working on Sonnet 5. Each one fails with a 400 invalid_request_error.

① To turn thinking off, use between_tools instead of disabled

On Sonnet 5, thinking: {"type": "disabled"} turned thinking off at any effort level. On Sonnet 5.5, disabled returns a 400 with this message:

"thinking.type.disabled" is not supported for this model. Use "thinking.type.between_tools" for the lowest thinking setting, or "thinking.type.adaptive" and "output_config.effort" to control thinking behavior.

The replacement is thinking: {"type": "between_tools"}, the lowest thinking setting on this model, which stops the "upfront thinking" the model does before answering. For requests without tools, the response contains only text, just as with disabled on Sonnet 5. No beta header is needed, and it works on every platform that offers Sonnet 5.5. But if you swap it in as a drop-in for disabled, three things will trip you up.

  • It returns a 400 at effort xhigh and max. between_tools is accepted only at low, medium and high. For jobs that ran "xhigh with thinking off" on Sonnet 5, you have to either lower effort to high or below, or stop turning thinking off.
  • You can't send other fields with it. Sending display, budget_tokens or block_binding together with between_tools returns a 400.
  • You can't change effort mid-conversation. Sending a different level with per-message effort (beta) returns a 400. If you want to vary the level per turn, use adaptive thinking (omit thinking or send {"type": "adaptive"}).

Here is the migration guide's before-and-after example in Python. Note that because the "before" example uses xhigh, the "after" example lowers it to high.

# Before: works on Sonnet 5, returns 400 on Sonnet 5.5
client.messages.create(
    model="claude-sonnet-5",
    max_tokens=16000,
    thinking={"type": "disabled"},
    output_config={"effort": "xhigh"},
    messages=[{"role": "user", "content": "..."}],
)

# After: stop upfront thinking; effort must be high or below
client.messages.create(
    model="claude-sonnet-5-5",
    max_tokens=16000,
    thinking={"type": "between_tools"},
    output_config={"effort": "high"},
    messages=[{"role": "user", "content": "..."}],
)

Even with between_tools, the short progress notes the model writes between tool calls come back as thinking blocks with summary text. Don't discard them; send them back unchanged along with the rest of the assistant turn. From the blocks you send back, the model receives the full text of the notes it wrote (What's new). A manual budget, {"type": "enabled", "budget_tokens": N}, still returns a 400, as on Sonnet 5.

② Forced tool use returns an error

Setting tool_choice to {"type": "any"} or {"type": "tool", "name": "..."} returns a 400. The token counting API applies the same check. Only auto (the default) and none are allowed.

tool_choice: type "tool" and "any" are not supported for this model.

Fix: keep tool_choice at auto and add strict: true to the tool definitions (strict tool use), or move the schema to structured outputs. With auto, the model may answer in text without calling the tool, so state in the prompt when it should use that tool. Strict tool use accepts only a subset of JSON Schema, and every object in the schema needs additionalProperties: false. You can make up to 20 tools strict per request, and strict can't be set on MCP, computer use or browser use toolsets.

# Before: works on Sonnet 5, returns 400 on Sonnet 5.5
client.messages.create(
    model="claude-sonnet-5",
    max_tokens=1024,
    tools=tools,
    tool_choice={"type": "tool", "name": "get_weather"},
    messages=[{"role": "user", "content": "What's the weather in Paris?"}],
)

# After: auto + strict; say in the prompt when to use the tool
client.messages.create(
    model="claude-sonnet-5-5",
    max_tokens=1024,
    tools=[{**tool, "strict": True} for tool in tools],
    tool_choice={"type": "auto"},
    messages=[{"role": "user",
               "content": "What's the weather in Paris? Use the get_weather tool."}],
)

Amazon Bedrock is different. Structured outputs (including strict tool use) aren't available for Sonnet 5.5 on Bedrock, so send auto without strict, state in the prompt when to call the tool, and validate tool inputs in your own code (migration guide).

③ Thinking blocks are bound to the model and the conversation

Each thinking block records which model produced it. Sonnet 5.5 can read thinking blocks from Sonnet 5, Opus 4.8, Haiku 4.5 and earlier models, but not from Opus 5, Opus 5.5, Fable or Mythos. And no other model can read Sonnet 5.5's thinking blocks.

  • Switching from Sonnet 5 to Sonnet 5.5 carries the earlier reasoning over.
  • Switching from Sonnet 5.5 to any other model means the turns after the switch run without Sonnet 5.5's reasoning. The request itself succeeds, and the dropped blocks are not billed.

Fable 5.1 and Mythos 5.1 can read Opus 5.5's thinking blocks, but nobody can read Sonnet 5.5's. If you route "only the hard parts up from Sonnet 5.5 to Opus 5.5," test on the assumption that Sonnet 5.5's reasoning disappears at the model you escalate to.

There is one more check: whether everything before a Sonnet 5.5 thinking block (system, tools and earlier messages) has changed since the block was produced. For accounts created on or after August 31, 2026, 00:00 UTC, this is enforced by default on the Claude API, Amazon Bedrock and Google Cloud, and sending a block with a history that was rewritten along the way returns a 400.

Fix: advance the conversation by appending only. When you want to change instructions or tools, use a mid-conversation system message instead of rewriting history (a feature newly available on Sonnet 5.5 that Sonnet 5 doesn't have). If rewriting is unavoidable, add the beta header thinking-binding-controls-2026-08-01 and set thinking.block_binding.prefix_mismatch_behavior to "drop_block" to have the affected blocks discarded instead of raising an error. But block_binding works only with adaptive thinking, so with between_tools, either append only, or remove the thinking blocks from the rewritten turn onward yourself.

④ The legacy computer use tool is rejected on the Claude API and Google Cloud

Sonnet 5 also accepted computer use through the legacy tool computer_20251124 with a beta header. Sonnet 5.5 on the Claude API and Google Cloud accepts only the toolset computer_toolset_20260801, and declaring the legacy tool returns a 400. On the Claude API, the message starts with:

'claude-sonnet-5-5' does not support tool types: computer_20251124.

Fix: remove the beta header, replace tools with [{"type": "computer_toolset_20260801"}], and update your agent loop to the toolset's shape (tool_use blocks for members, multiple actions at once, and toolset_name in results). If you send the fine-grained-tool-streaming-2025-05-14 beta header, remove it. Sending it together with the toolset returns a 400, so add eager_input_streaming: true to each tool that needs it. On Amazon Bedrock the legacy tool still works, so no change is needed.

⑤ Some advisor tool pairings are rejected

The advisor tool (beta) lets an executor model ask a stronger model for advice. When Sonnet 5.5 is the executor, specifying Opus 4.8, Opus 4.7, Opus 4.6, Sonnet 5 or Sonnet 4.6 as the advisor returns a 400. The allowed advisors are Opus 5, Opus 5.5, Sonnet 5.5, Fable 5, Fable 5.1, Mythos 5 and Mythos 5.1.

In addition, every advisor that Sonnet 5.5 accepts returns its advice in an encrypted advisor_redacted_result block, so the client can't read the text of the advice. Anything that logged the advice or displayed it on screen needs rebuilding. Cheap pairings such as "Sonnet 5 as executor, Opus 4.8 as advisor" can't carry over as-is to Sonnet 5.5.

6. Changes that happen without an error

The five breaking changes raise a 400, so you'll notice them. The tricky ones are changes in behavior that don't raise any error.

Change What happens / what to do
Text between tool calls goes into thinking blocks Notes longer than a sentence or two, such as "Next I'll check X," which were text blocks on Sonnet 5, now come back in progress thinking blocks (a short one-liner stays as text). With the default display: "omitted" their content is empty, so a UI that streamed progress to users goes quiet during tool calls. With adaptive thinking, set display to "updates" (beta, header thinking-display-updates-2026-08-18) or "summarized", and show non-empty thinking blocks before the tool_use that follows. With between_tools, the text comes back with no configuration.
Effort levels were recalibrated The same level won't necessarily think as much as on Sonnet 5. The official guidance is not to carry settings over but to re-tune effort (Section 7).
Refusals now have five categories cyber, bio, frontier_llm, reasoning_extraction and general_harms. A refusal returns HTTP 200 with stop_reason: "refusal", so read stop_details and handle it. The server-side fallback (fallbacks: "default", beta, Claude API only) retries only cyber and frontier_llm on Sonnet 5.
Thinking blocks are bound to the account that created them Sonnet 5.5's thinking blocks can be used only by the account that created them or accounts linked to it. If sent from another account, the blocks are dropped and the request succeeds. The announcement says this includes switching accounts mid-session in Claude Code.
The cache minimum dropped Prompts of 512 to 1,023 tokens are now cached. That adds a write cost (1.25× input for 5 minutes), while from the second call on you pay only the read rate (Section 3).

Sources: What's new in Claude Sonnet 5.5, Migrating to Claude Sonnet 5.5, Anthropic's announcement

The announcement also describes changes to safeguards. Sonnet 5.5 is the first Sonnet to ship with cyber safeguards and a fallback, because its capabilities usable for cyberattacks have reached the level of Opus 5; high-risk cybersecurity work visibly switches to Sonnet 5. Finding and fixing bugs in everyday development is not affected. Biology safeguards are the same as on Sonnet 5.

On the other hand, some features were added. Per-message effort (beta; lets you change the level while keeping the prompt cache), mid-conversation system messages and mid-conversation tool changes (beta) were all unavailable on Sonnet 5. What's new also lists compaction, which summarizes the conversation whenever you ask (beta, header compact-2026-09-04), and defining tools inside messages (beta, header inline-tools-2026-09-15).

7. Default effort: high on the API, medium in Claude Code

Opus 5.5 lowered its default effort to medium even on the API, but Sonnet 5.5 keeps high as the API default. Meanwhile, the announcement says the default in Claude Code and the apps is now medium. The same Sonnet 5.5 thinks one level differently by default on the API and in Claude Code.

The official effort page recommends these starting points for Sonnet 5.5:

Type of work Starting level (official recommendation) Notes
General work (other than the two below) high Same as the API default
Agentic coding and multi-step tool use Start at medium medium for well-specified work; high for hard or long work
Latency-sensitive work such as chat medium or low Prioritize speed
xhigh and max Only when evals show higher quality Can't be combined with between_tools

Sources: "Recommended effort levels for Claude Sonnet 5.5" in Anthropic, "Effort", and the migration guide

The same page also recommends setting max_tokens large enough for both thinking and the response, and, for agentic coding, setting max_tokens to this model's limit of 128,000 and receiving the response via streaming. Thinking tokens count toward max_tokens even in settings where the thinking content isn't returned. How effort itself works is covered in our guide to the effort setting, and adaptive thinking in adaptive vs. extended thinking.

8. Choosing between Sonnet 5.5, Opus 5.5 and Haiku

The official model overview still says to start with Opus 5.5 if you're unsure; Sonnet 5.5 has not become the starting point. With that said, the announcement describes Sonnet 5.5's strengths as well-scoped everyday work, bug fixes, and creating documents, slides and spreadsheets, and Opus 5.5's as complex work that needs careful judgment. It also says that Sonnet 5.5 complements Opus 5.5 best at lower effort, and that at higher levels the two can reach similar performance at similar cost. In other words, if you run at xhigh or max, choosing Sonnet 5.5 gives you little cost advantage.

Choose Sonnet 5.5
  • You want to move fast on well-specified implementation and bug fixes at low to high.
  • You create a lot of documents, slides and spreadsheets.
  • You want to cut upfront thinking to shorten wait times (between_tools).
  • You use Sonnet 5 today (it's a same-price replacement).
Choose Opus 5.5
  • Open-ended design or research needs sustained judgment.
  • You can't decide between them (the official starting point).
  • You plan to run at high effort, where the cost gap with Sonnet 5.5 narrows.
  • You want fast mode (Opus only).
Choose Haiku
  • You have high volume and per-item cost comes first (Haiku 4.5 is half the price of Sonnet 5.5).
  • A 200K-token context is enough.
  • The announcement says Haiku 5.5 is coming within weeks.
  • Haiku 4.5's retirement is not before October 15, 2026, so check your plans ahead.

When in doubt, go in this order: ① run Sonnet 5.5 at medium and high on your own work → ② if that falls short, compare with Opus 5.5 at medium → ③ if Sonnet 5.5 at xhigh costs about the same as Opus 5.5, choose Opus 5.5. Opus 5.5 is covered in detail in our Opus 5.5 release overview, and the full list of current models in major AI models and their knowledge cutoffs.

9. Using it in Claude Code and GitHub Copilot

Claude Code added Sonnet 5.5 in the v2.1.284 CHANGELOG (September 28, 2026) and made it the default Sonnet on the Anthropic API.

The default model is still Opus 5.5

default remains Opus 5.5 on Pro, Max, Team, Enterprise and the API. To use Sonnet 5.5, pick /model sonnet (or claude --model sonnet at launch). Versions older than v2.1.284 can't use it, so run claude update.

What sonnet points to depends on the provider

It resolves to Sonnet 5.5 only on the Anthropic API. On Claude Platform on AWS it points to Sonnet 4.6, and on Bedrock, Google Cloud's Agent Platform and Microsoft Foundry to Sonnet 4.5. There, pick the full model name or set ANTHROPIC_DEFAULT_SONNET_MODEL.

Effort starts at medium; thinking can't be turned off

Sonnet 5.5 starts at medium. The old-style top-level effortLevel in user settings has no effect on Opus 5.5 and later models. The thinking toggle, alwaysThinkingEnabled and MAX_THINKING_TOKENS=0 also have no effect on Sonnet 5.5.

Cyber falls back to Sonnet 5; biology is refused

When a safeguard classifier fires, cybersecurity work is re-run on Sonnet 5 and the session continues on that model (use /model to switch back). For biology, Sonnet 5.5 has no fallback model, so it ends in a refusal.

Sources: Claude Code, "Model configuration" and the CHANGELOG (v2.1.284; checked September 29, 2026)

On the Anthropic API, Sonnet 5.5 always has a 1M-token context, with no [1m] suffix and no extra charge. Auto-compact runs at about 967K tokens by default. Through an LLM gateway (with ANTHROPIC_BASE_URL set), the context is treated as 200K tokens, so pick "Sonnet 5.5 (1M context)" (sonnet[1m]) in the model picker.

opusplan, which splits planning to Opus and execution to Sonnet, runs on Opus 5.5 and Sonnet 5.5 respectively on the Anthropic API. How to use it is covered in our opusplan guide. Claude Code itself advances conversations by appending only, so it doesn't run into the history check in Section 5 ③ (per the Opus 5.5 migration guide).

GitHub Copilot

On the same day, September 28, GitHub announced the general availability of Sonnet 5.5 in its changelog. It is available on Copilot Pro, Pro+, Max, Business and Enterprise, and can be selected from the model picker in VS Code, Visual Studio, Copilot CLI, Copilot coding agent, github.com, JetBrains IDEs, Xcode and more. The rollout is gradual, so it may not appear right away. Billing is usage-based at the provider's list price. On Business and Enterprise, administrators decide whether it's available through the model policy in Copilot settings.

10. Migration steps (for API users)

From the migration guide's checklist, here are the items that apply when moving from Sonnet 5, in working order. Claude Code also has /claude-api migrate to help with this (a bundled skill pointed to by the migration guide; it confirms the scope with you before editing).

  1. Change the model ID from claude-sonnet-5 to claude-sonnet-5-5 (anthropic.claude-sonnet-5-5 on Bedrock).
  2. Replace thinking: {"type": "disabled"} with {"type": "between_tools"} and set effort to high or below.
  3. Replace any and tool in tool_choice with auto plus strict tool use (on Bedrock, use auto only and validate inputs yourself).
  4. If you rewrite system, tools or past messages mid-conversation, switch to appending only.
  5. If you use computer use on the Claude API or Google Cloud, move to computer_toolset_20260801 and update the loop.
  6. If your advisor tool uses Opus 4.8, Opus 4.7, Sonnet 5 or similar as the advisor, switch to one Sonnet 5.5 accepts, and assume the advice comes back encrypted.
  7. If you show progress on screen, set display to "updates" or "summarized" (not needed with between_tools).
  8. Handle stop_reason: "refusal" and set up a fallback.
  9. If you route to other models, test on the assumption that Sonnet 5.5's reasoning isn't carried over.
  10. Re-tune effort and re-measure cost and latency. Also check caching for prompts under 1,024 tokens.

Source: "Every starting model" and "Migrating to Claude Sonnet 5.5 from Claude Sonnet 5" in Anthropic, "Migrating to Claude Sonnet 5.5", reordered into working order

If you are coming directly from Sonnet 4.6 or earlier, you also need to deal with thinking running even on jobs that didn't specify thinking, the 400s for budgets and for temperature and similar parameters, and about 30% more tokens for the same text. If you use it through Claude Managed Agents, nothing needs to change other than the model name (note in the migration guide).

Summary

Claude Sonnet 5.5 is a release that raised performance substantially while keeping the price unchanged, and in exchange moved its API conventions toward Opus 5.5. In Anthropic's own table it improved on Sonnet 5 in every row and came within a few points of Opus 5.5 in most. On the other hand, disabled, forced tool use, the legacy computer use tool, rewritten history and some advisor pairings now cause 400 errors.

The easiest things to miss are that between_tools, the replacement for turning thinking off, is accepted only at high or below, and the mismatch that default effort is high on the API but medium in Claude Code. On the API, set effort explicitly; in Claude Code, select it with /model sonnet and then check the effort level. Whether it actually got cheaper should be measured with usage, not unit prices.

Finally, benchmark numbers are values measured in that table under those conditions. Anthropic itself writes that Opus 5.5 remains clearly stronger on complex, open-ended work. Start by comparing Sonnet 5.5 at medium and high on your own work.

FAQ

Q. Is Sonnet 5.5 more expensive than Sonnet 5?

A. The rates are identical across every item ($2 input, $10 output, $0.20 cache reads, and half price on the Batch API). Anthropic says it is up to 30% cheaper per task, but that is its own measurement. Because the effort levels have been recalibrated, check with usage after migrating.

Q. Is there a way to turn thinking off?

A. On the API, thinking: {"type": "between_tools"} stops upfront thinking. For requests without tools, the response contains only text. However, it returns a 400 at effort xhigh and max, and progress notes between tool calls come back in thinking blocks. In Claude Code, settings that turn off thinking have no effect on Sonnet 5.5.

Q. I changed the model ID and got a 400 error.

A. You can tell which change from the error message. "thinking.type.disabled" means the thinking setting, tool_choice: type "tool" and "any" means forced tool use, and computer_20251124 means the legacy computer use tool. If it happens while using between_tools, check whether effort is xhigh or above, or whether you're sending display or similar fields with it. If you use the advisor tool, suspect the advisor model; if you rewrite history mid-conversation, suspect the thinking-block check (Section 5 ③).

Q. Claude Code isn't using Sonnet 5.5.

A. The default model is Opus 5.5, so select it with /model sonnet. If it's still Sonnet 5, your version is older than v2.1.284, so run claude update. On Bedrock, Google Cloud, Claude Platform on AWS and Microsoft Foundry, sonnet points to an older Sonnet, so pick the full model name or set ANTHROPIC_DEFAULT_SONNET_MODEL.

Q. Should I make Opus 5.5 or Sonnet 5.5 my default?

A. The official starting point is Opus 5.5. If your work is mainly well-scoped implementation, bug fixes and document creation, run at low to high effort, Sonnet 5.5 costs half as much per input and output token. If you run at xhigh or max, the cost gap narrows, as the announcement says, so compare both on your own work.

* The figures in this article are based on Anthropic's official announcement, "Introducing Claude Sonnet 5.5" (benchmarks from that page's comparison table and chart descriptions), the official documentation "Models overview," "Claude Sonnet 5.5," "What's new in Claude Sonnet 5.5," "Migrating to Claude Sonnet 5.5," "Pricing," "Thinking," "Effort" and "Prompt caching," Claude Code's "Model configuration" and CHANGELOG, and GitHub's changelog (all checked September 29, 2026). Specifications and prices may change, so confirm the final details in the official documentation.

Related articles: Claude Opus 5.5 release overview, Claude Fable 5.1 breaking changes and migration, Claude pricing comparison.