AI Catchup

Prompting Claude Sonnet 5.5: Effort, between_tools, and the Five Breaking Changes

By 18 min read

Run Claude Sonnet 5.5 at high effort for general API work, medium for well-specified agentic coding, and medium or low for chat. Thinking is on by default; turn it off with between_tools, not disabled. Five API changes break Sonnet 5 code, and cache reads now cost $0.10 per million tokens.

Claude Sonnet 5.5 needs three changes from you before it pays off: re-run your effort sweep, replace thinking: {"type": "disabled"} with between_tools, and fix the five API changes that turn working Sonnet 5 requests into 400 errors. Once you have, delete more prompt text than you add. Anthropic says existing Sonnet 5 prompts should perform well without changes, and recommends removing Sonnet 5 workarounds and re-running your evals (Prompting Claude Sonnet 5.5; claude.dev).

This guide condenses Anthropic's September 28 developer post, the Sonnet 5.5 prompting guide, and the migration docs into settings and prompt changes. For launch coverage, see Claude Sonnet 5.5 launches. For the larger model, see prompting Claude Opus 5.5. For techniques that apply to every model, see the prompt engineering playbook.

Key Takeaways

  • Pick Sonnet 5.5 for checkable work. Anthropic recommends it for well-scoped coding, documents, and repeated agent tasks, and Opus 5.5 for complex work requiring careful judgment.
  • Cache reads cost $0.10 per million tokens since October 7, half the rate in Anthropic's launch post. Input and output stay at $2 and $10.
  • Re-run your effort sweep. Levels are recalibrated. Start at high for general work, medium for well-specified agentic coding, and medium or low for chat.
  • Turn thinking off with between_tools. Sending disabled returns a 400 error, and between_tools works only at high effort or below.
  • Fix five breaking changes: disabled thinking, forced tool choice, edited history before a thinking block, the old computer use tool, and three advisor models.
  • Delete Sonnet 5 workarounds and any instruction asking the model to show its reasoning in the response.
  • Read every response by block type. Notes between tool calls now arrive as thinking blocks, so a text-only interface goes silent.

Sonnet 5.5 or Opus 5.5: Which to Use

Use Sonnet 5.5 when the task has a clear spec and a way to check the result, and Opus 5.5 when it needs long-horizon judgment. That is Anthropic's own split: Opus 5.5 is built for complex work requiring careful judgment, while Sonnet 5.5 handles well-scoped everyday tasks and is fast enough for quick iteration (claude.dev).

Your workloadStart with
Fixing bugs, iterating on features, verifying against requirementsSonnet 5.5
High-volume everyday developmentSonnet 5.5
Documents, slides, and spreadsheets where design mattersSonnet 5.5
Repeated agent tasks: investigation, review, draftingSonnet 5.5
Long-horizon agentic coding and knowledge workOpus 5.5
The hardest problemsOpus 5.5

Workload mapping from Anthropic's developer post (claude.dev).

The benchmark gap is small on scoped work. On Anthropic's launch table, Sonnet 5.5 scores 55.5% on CursorBench 4.0 against Opus 5.5's 57.8%, and 70.6% on Terminal-Bench 4.0 against Sonnet 5's 10.3% (Anthropic). Anthropic's announcement adds that Opus 5.5 remains clearly stronger at complex, open-ended work. For high-volume subagents and summaries, look one tier down: Claude Haiku 5.5 launched on October 7, 2026 (release notes). Our Haiku 5.5 coverage compares the two.

What It Costs Now

Sonnet 5.5 costs $2 per million input tokens and $10 per million output tokens, with cache reads at $0.10 per million (pricing). Anthropic cut the cache-read price from $0.20 to $0.10 on October 7, 2026 (release notes). The September 28 developer post and the launch announcement still show $0.20, so use the pricing page.

Per million tokensSonnet 5.5Sonnet 5Opus 5.5
Input$2$2$4
Output$10$10$20
5-minute cache write$2.50$2.50$5
1-hour cache write$4$4$8
Cache read$0.10$0.20$0.20

Prices from Anthropic's pricing page, checked October 9, 2026 (pricing).

Batch requests cost $1 input and $5 output per million tokens, and US-only inference through inference_geo applies a 1.1x multiplier (pricing). Your bill per task should fall even at Sonnet 5's per-token rates: Anthropic reports Sonnet 5.5 costs up to 30% less per task than Sonnet 5 because it needs fewer tokens (Anthropic).

Three details change the math. Coming from Sonnet 4.6 or earlier, images cost more: Sonnet 5.5 uses the high-resolution tier, so a 2000×1500 image costs about 2.5 times as many tokens as on Sonnet 4.6 (migration guide); downscale images when you don't need the detail. Text counts rise on the same move: against Sonnet 4.6, Sonnet 4.5, or Haiku 4.5, the same text produces about 30% more tokens (migration guide). In your favor, the minimum cacheable prompt drops to 512 tokens from 1,024 on Sonnet 5, so shorter system prompts now cache (What's new).

Choosing an Effort Level

Start at high unless your workload is agentic or latency-sensitive. Effort is the main control for how much Sonnet 5.5 thinks, and its levels are recalibrated, so the setting you used on Sonnet 5 won't carry over (Prompting Claude Sonnet 5.5).

WorkloadStart atMove to
General API workhighLower if evals hold
Agentic coding, well-specifiedmediumhigh for harder or longer tasks
Multistep tool usemediumhigh for harder tasks
Chat and latency-sensitive workmedium or lowRaise only if quality needs it
Hardest tasksxhigh or maxOnly where evals show a gain

Starting points from Anthropic's effort guidance (Effort).

The default is high on the Claude API and medium in Claude Code (claude.dev). Four rules keep effort predictable:

  • Size max_tokens for thinking. Thinking counts toward max_tokens. For agentic coding, Anthropic recommends 128,000, the model's maximum, with streaming (Effort).
  • Lower effort to get less thinking. From medium up, the model thinks briefly before almost every reply, even a greeting, and telling it in the system prompt to think less doesn't reliably reduce thinking (Prompting Claude Sonnet 5.5).
  • Change effort per message, not per request. Changing the top-level effort between requests invalidates the prompt cache. On the Claude API and Google Cloud, a per-message effort change (beta) keeps it; Amazon Bedrock doesn't offer per-message effort for Sonnet 5.5 (claude.dev; Effort).
  • Watch the low end. At low and medium, the model is more likely to stop and check in on long agentic tasks, and at low it can skip verifying a change (Prompting Claude Sonnet 5.5).

Turning Thinking Off With between_tools

Send thinking: {"type": "between_tools"} to turn off up-front thinking. A request with no thinking field runs with adaptive thinking, and thinking: {"type": "disabled"} returns a 400 error (claude.dev). With between_tools, the model thinks only between tool calls.

It comes with limits you need to code around (What's new):

  • Effort cap. It's accepted at low, medium, and high. At xhigh or max it returns a 400 error, so use adaptive thinking there.
  • No other fields. Sending display, budget_tokens, or block_binding alongside it returns a 400 error.
  • Fixed effort. A per-message effort that differs from the level in effect returns a 400 error. To vary effort per turn, use adaptive thinking.
  • Thinking blocks still appear. Longer notes between tool calls come back as thinking blocks with summary text. Pass them back unchanged with the rest of the assistant turn.

Two follow-ups matter. Without tools, between_tools means the model answers without thinking, so use adaptive thinking for requests that need a few steps of working out. And remove any instruction telling the model not to think: with between_tools, it makes internal XML tags more likely to show up in visible output (Prompting Claude Sonnet 5.5).

The Five Breaking Changes and the Errors They Return

Change the model ID to claude-sonnet-5-5, then fix these five changes, each of which can return a 400 error on code written for Sonnet 5 (What's new).

ChangeWhat failsFix
Thinking offthinking type disabledSend between_tools at high effort or below
Forced tool usetool_choice type any or tool, including on token countingSend auto, mark the tool strict: true, and say in the prompt when to use it
Thinking blocks bound to the conversationReplaying a Sonnet 5.5 thinking block after editing the system prompt, tools, or an earlier messageKeep history append-only; change instructions with mid-conversation system messages
Computer usecomputer_20251124 on the Claude API and Google CloudMove to computer_toolset_20260801; Amazon Bedrock still accepts the old tool
Advisor toolOpus 4.8, Opus 4.7, or Sonnet 5 as advisorPair with Opus 5.5, Opus 5, Sonnet 5.5, Fable 5, Fable 5.1, Mythos 5, or Mythos 5.1

Error behavior per Anthropic's What's new page and migration guide (What's new; migration guide).

The history check is enforced by default for accounts created on or after August 31, 2026, on the Claude API, Amazon Bedrock, and Google Cloud (What's new). Older accounts can opt in, so build append-only either way. Thinking blocks also stay with the account that produced them, or a linked account; another account's blocks are dropped before the model sees them.

On which models read Sonnet 5.5's thinking blocks, Anthropic's sources disagree. The developer post says no other model reads them. The What's new page says Opus 5.5 reads them on the Claude API and Google Cloud (What's new). The docs are the authoritative reference, so a conversation that moves up from Sonnet 5.5 to Opus 5.5 there keeps its reasoning. Moving to any other model drops it.

A sixth change fails nothing but breaks interfaces: notes longer than a sentence or two between tool calls now arrive as progress-update thinking blocks, empty at the default display (migration guide). The progress section below covers the fix.

Coming from older models, expect more 400 errors: thinking budgets, non-default temperature, top_p, or top_k, and assistant prefill all fail on Sonnet 5.5 (migration guide). Claude Sonnet 4.5 retires from the Claude API on November 30, 2026, so plan that move now (release notes). To automate the swap, run /claude-api migrate this project to claude-sonnet-5-5 in Claude Code; the bundled skill applies the model ID change and the breaking parameter changes, then lists what to verify by hand (migration guide).

Prompts to Add

Add instructions only when you see the symptom they fix. Each row below paraphrases a system-prompt pattern from Anthropic's Sonnet 5.5 prompting guide; the guide has the exact wording (Prompting Claude Sonnet 5.5).

SymptomWhat to add
Stops to check in before a coding task is done, at low or mediumKeep working until everything asked for is done; stop only when blocked or before a risky step
Adds tests, docs, or files you didn't ask forWhen the requested work is done and checked, stop and report; suggest extras at the end instead of building them
Starts its own review rounds or reviewer subagents at xhigh or maxOnce checks pass, stop; don't launch review passes unless asked
Reports code changes done without test or build output, at lowRun a real check that exercises the change (tests, type-checker, build), and say which check was skipped and why if none can run
Answers from training data on facts that changeUse the search tool for anything that may have changed, such as what is allowed, required, or charged
Builds a deliverable when you wanted ideasWhen asked for ideas or a plan, give them and stop until told to proceed

Two of these have measured effects. The stop-and-report instruction at max effort stopped reviewer subagents and cut session cost by about a third with no change in quality, and the real-check instruction at low made skipped checks rare at only a slightly higher cost per task (Prompting Claude Sonnet 5.5).

Two harness changes help more than prompts. Accept tool calls whose name differs only in letter case, or return an is_error result naming the exact expected tool, because Sonnet 5.5 occasionally writes bash for Bash. And for dense charts, give the model a crop or zoom tool: with tools at high effort, it read charts more accurately than without tools at max, at a fraction of the cost (Prompting Claude Sonnet 5.5).

Prompts to Delete

Delete before you add. Anthropic tells you to remove Sonnet 5 workarounds, such as refusal steering, tool-call retry shims, and "do not be lazy," then re-run your evals before tuning anything else (claude.dev).

DeleteWhy
Sonnet 5 workarounds (refusal steering, retry shims, "do not be lazy")Existing prompts should perform well without them
Requests to write out reasoning in the responseThey invite reasoning_extraction declines; read summarized thinking instead
Instructions to think less or not at allThey don't reliably reduce thinking, and with between_tools they leak internal tags
Instructions to minimize tool calls or use tools only when strictly necessaryThey push the model to answer from training knowledge
Instructions to save all findings for the final responseThey fight the progress updates the model now writes

Sources: the developer post and Anthropic's prompting guide (claude.dev; Prompting Claude Sonnet 5.5).

You can still ask for a short explanation of the answer or a summary of the actions taken. What triggers a decline is asking the model to reproduce its internal reasoning in the response text.

JSON and Structured Outputs

Use structured outputs, and add one line to the end of the system prompt: "Think the problem through before you answer." On tasks that need a few steps of working out, such as totaling figures or ranking items, Sonnet 5.5 often answers without thinking first, especially at low and medium (Prompting Claude Sonnet 5.5). At high effort, that line brings accuracy close to xhigh for a modest increase in output tokens. Running at xhigh with adaptive thinking gives the highest accuracy without the line.

Three rules round it out:

  • Use adaptive thinking, not between_tools, for these requests. Without tools, between_tools skips thinking and the line has no effect.
  • Treat a max_tokens stop as a failure even when the text holds valid JSON, and retry.
  • Without structured outputs, parse the last JSON value in the text blocks, not everything from the first brace to the last, because the model sometimes writes a draft first.

Structured outputs aren't available for Sonnet 5.5 on Amazon Bedrock; there, describe the format in the prompt and validate the output in code (migration guide).

Progress Updates and Mid-Turn User Messages

Set thinking.display to "updates" or "summarized" if your interface shows the model's notes between tool calls. Sonnet 5.5 returns notes longer than a sentence or two as progress-update thinking blocks, which are empty at the default display, so a client that renders only text blocks looks silent during long agentic turns (migration guide). The "updates" value is in beta and needs the thinking-display-updates-2026-08-18 header; with between_tools, the notes come back without any display setting. Render each non-empty thinking block before the tool_use block that follows it.

For exact text partway through a turn, such as a code snippet or a question, give the model a simple tool for messaging the user and declare it in the first request. If turns still go quiet, have your harness count consecutive silent tool-calling steps and, after about five, append a one-turn reminder as a turn-scoped system message; stop after the second or third reminder (Prompting Claude Sonnet 5.5).

Handle messages users type mid-task carefully, because Sonnet 5.5 resists prompt injection and can mistake a genuine user message for one:

  • Never put user text inside a tool_result block. That placement causes the misread most often.
  • Append the user's words as a text block in the user message, after the last tool_result.
  • Keep harness notices in a separate system message, never in the same block as the user's words.
  • Drop your own token or budget countdowns after tool results in interactive sessions.

Refusals and Fallback

Check stop_reason on every response. A declined request returns HTTP 200 with stop_reason: "refusal", and stop_details names one of five categories: cyber, bio, frontier_llm, reasoning_extraction, or general_harms (claude.dev). Anthropic's developer post calls Sonnet 5.5 the first Sonnet model with cybersecurity safeguards similar to those on its most capable models, and says most routine software development is unaffected.

Server-side fallback (fallbacks: "default", beta, Claude API) retries cyber and frontier_llm declines on Sonnet 5 and doesn't retry the other three (What's new). Since September 24, 2026, refusals that arrive before any output in the bio, frontier_llm, and reasoning_extraction categories are billed (release notes). For legitimate security work, Anthropic says the Cyber Verification Program will soon expand to Sonnet 5.5 (claude.dev). For life sciences work blocked by the bio classifier, organizations can apply to the Life Sciences Verification Program (Prompting Claude Sonnet 5.5).

Sonnet 5.5 in Claude Code

Run /model sonnet for well-scoped tasks; Claude Code's default model stays Opus 5.5 (claude.dev). Sonnet 5.5 requires Claude Code v2.1.284 or later (Claude Code docs).

SettingSonnet 5.5 in Claude Code
Minimum versionClaude Code v2.1.284 or later
Default effortmedium, versus high on the Claude API
ThinkingCan't be turned off; effort sets how much it thinks
Fast modeNot available
Context window1M tokens, native
Cyber-flagged requestsRe-run on Sonnet 5
Bio-flagged requestsEnd with a refusal

Settings from Anthropic's developer post and the Claude Code model configuration docs (claude.dev; Claude Code docs).

The sonnet alias resolves to Sonnet 5.5 only on the Anthropic API. On Claude Platform on AWS it resolves to Sonnet 4.6, and on Amazon Bedrock and Google Cloud's Agent Platform to Sonnet 4.5, so set the full model name or ANTHROPIC_DEFAULT_SONNET_MODEL there (Claude Code docs). A good split: let Opus 5.5 plan the architecture and Sonnet 5.5 implement it, and cap delegation with our subagent patterns. For the larger model's habits, read prompting Claude Opus 5.5 next.

Sources

More practices and workflows

Frequently Asked Questions

What effort level should I use with Claude Sonnet 5.5?

Start at high, the Claude API default, for general work. For agentic coding and multistep tool use, start at medium on well-specified tasks and move to high for harder or longer ones. For chat, use medium or low. Use xhigh or max only where your evals show a quality gain.

How do I turn off thinking on Claude Sonnet 5.5?

Send the thinking type between_tools instead of disabled, which now returns a 400 error. It works at low, medium, and high effort only, accepts no other thinking field, and still returns longer notes between tool calls as thinking blocks. In Claude Code, thinking can't be turned off for Sonnet 5.5.

How much does Claude Sonnet 5.5 cost?

Input costs $2 and output $10 per million tokens, the same as Sonnet 5. Cache reads cost $0.10 per million tokens since Anthropic halved them on October 7, 2026. Batch requests cost $1 input and $5 output, and US-only inference applies a 1.1x multiplier.

Why does my Sonnet 5 code return 400 errors on Sonnet 5.5?

Five changes break Sonnet 5 code: disabled thinking, forced tool_choice of any or tool, replaying thinking blocks after editing earlier history, the older computer_20251124 tool on the Claude API and Google Cloud, and Opus 4.8, Opus 4.7, or Sonnet 5 as an advisor. Each has a direct replacement.

Should I use Claude Sonnet 5.5 or Opus 5.5?

Use Sonnet 5.5 for well-scoped work with a clear spec and a way to check the result: bug fixes, feature iteration, documents and slides, and agent tasks you run repeatedly. Use Opus 5.5 for long-horizon work that needs careful judgment. Sonnet 5.5 costs half Opus 5.5's per-token price.

Get the weekly AI Catchup

Tools, practices, and what matters, in your inbox every week.