Prompting Claude Opus 5.5: Delete 'Think Carefully', Name the Finish Line, Start at Medium
Prompt Claude Opus 5.5 by handing over the whole task with a clear finish line, deleting “think carefully” lines, and starting at medium effort. Thinking is always on, so effort is your depth control. Add a CLAUDE.md rule that names when to stop, and set effort explicitly instead of carrying Opus 5 settings over.
Claude Opus 5.5 needs less prompting than Opus 5, not more. Hand it the whole task with a finish line, delete the lines that tell it to think hard, start at medium effort, and spend your prompt on when it should stop. This guide combines Anthropic's September 22 playbook for Claude apps and Claude Code (claude.dev) with the API guidance in Prompting Claude Opus 5.5.
The verdict: keep your Opus 5 prompts, then delete. Anthropic expects existing Opus 5 prompts to work well unchanged. The wins come from removing instructions the model no longer needs, setting effort on purpose, and telling it which stops you want on a long run. For benchmarks, pricing, and the full list of API breaking changes, see our Claude Opus 5.5 launch coverage; for the cross-model fundamentals, see the prompt engineering playbook. If you also run the smaller model, our Sonnet 5.5 prompting guide covers it.
Key Takeaways
- Start at medium effort and set it explicitly. Medium is the Opus 5.5 default, one level below Opus 5's, and it matches or beats Opus 5 at high on Anthropic's coding and knowledge-work evaluations.
- Delete "think carefully" lines. Opus 5.5 always thinks and decides how much; effort is the control.
- Say what "done" looks like. Give the whole task in one message with the finish line and the one condition for stopping to ask.
- Name the stops you want. A short CLAUDE.md rule keeps long Claude Code runs moving; keep permission prompts on for destructive commands.
- Treat a text-only turn as a report, not as done. Unattended agents need a continuation step, capped at two or three.
- Ask what it couldn't confirm. Research and review prompts work better when the model must say where it looked and what it could not verify.
- Know the flag path. Flagged messages move to an older model;
/modelswitches back.
What Changed From Opus 5 and What Carries Over
Most of what you wrote for Opus 5 still works. The differences that change how you prompt are thinking, effort, and how the model reports on long work. Anthropic says Opus 5.5 produces output more than 30 percent faster than Opus 5 and usually needs fewer tokens for the same task.
| Behavior | Claude Opus 5 | Claude Opus 5.5 | What to do |
|---|---|---|---|
| Default effort | high | medium | Set effort explicitly and re-run your sweep |
| Thinking | On, and can be disabled at high or below | Always on; disabling it is rejected | Use lower effort where you used to disable thinking |
| Thinking per turn at the same effort | Baseline | More, most of all at xhigh and max | Leave room in max_tokens |
| Notes between tool calls | text blocks | thinking blocks, empty by default | Set thinking.display if users should see them |
| Reports on long work | Less plain | Says what it did, found, and needs; some reports end the turn | Name the stops you want |
| Safety classifiers | Cybersecurity, reasoning extraction (Claude Code docs also list its own biology classifier) | Adds Fable 5.1's biology safeguards | Handle refusals and fallback |
The Opus 5 prompting guide remains a reasonable starting point for verbosity, scope, and subagent caps. What no longer applies is anything that disables thinking.
Pick an Effort Level: Medium First, Then Sweep
Start at medium and set it explicitly rather than relying on the default. Opus 5.5 supports all five levels (low, medium, high, xhigh, max), and medium is the default, so a request that omits effort runs one level lower than it did on Opus 5. In Anthropic's testing, Opus 5.5 at medium matched or beat Opus 5 at high on coding and knowledge-work evals, and low came close on several coding evals for far less.
| Workload | Start at | Move when |
|---|---|---|
| Simple, high-volume calls with checkable output | low | Quality drops on your evals |
| Coding agents, code review, research | medium | Your evals show a measured gain at high |
| Long-horizon coding where accuracy pays | high | A sweep shows xhigh earns its cost |
| Anything at xhigh or max | Only after measuring | Never as a carried-over habit |
Anthropic's own sweep shows the trade. On Anthropic's SWE-bench Pro subset, measured against high, Opus 5.5 scored about 2.5 points lower at medium for about 70% of the cost, about 8 points lower at low for about a third of the cost, and about 1.4 points higher at xhigh for 2.5 times the cost of high. When outcomes are checkable (here, by the benchmark's tests), a cheaper policy matched high for a little over half the cost: with Opus 5.5 at low, 13% of tasks failed, and re-running those at high passed about 97% for about $0.17 each, against 95.3% for $0.29 running everything at high (optimizing for cost and intelligence). Use it for the saving, not the lift: each failed task takes a second run's worth of time.
Three settings decide whether a level behaves the way you expect:
- Thinking per turn rises at a given level. Opus 5.5 usually thinks more per turn than Opus 5 at the same setting, especially at xhigh and max, so an effort value carried over from Opus 5 buys longer turns and more output tokens.
- Size
max_tokensfor thinking, and note that Anthropic's pages differ. The prompting guide says amax_tokensof 128,000, the model's maximum, has worked well for long agentic coding turns. The cost guide says to setmax_tokensto 64,000 for agentic work, or to 128,000 when a single cut-off attempt is costly, and the migration guide checklist says to start at 64k at xhigh or max. Pick 64,000 by default and 128,000 for long unattended turns you cannot afford to lose. - Change effort per message, not per request. Switching the top-level effort value from one request to the next breaks the prompt cache. The per-message effort change (beta) keeps the cache, and it requires the beta header
mid-conversation-output-config-2026-07-01(Effort).
In Claude Code, Opus 5.5 requires v2.1.280 or later and starts at medium. One trap: a top-level effortLevel in your user settings file doesn't count for Opus 5.5, so choose a level for it with /effort or the /model picker (model configuration).
Delete "Think Carefully" and Reasoning-in-Reply Requests
Remove "think carefully," "think step by step," and similar lines from prompts, system prompts, and saved instructions. Opus 5.5 thinks before every reply and sets its own depth. Anthropic tested this in a chat product: taking such a line out made replies start sooner, with no clear loss of quality. For a quick answer to a simple question, tell it to answer directly; to change depth, change effort.
Also remove requests to write out its internal reasoning in the reply. A prompt that pushes the model to reproduce its reasoning in the response text may be declined with the reasoning_extraction refusal category, and server-side fallback returns those declines to you instead of retrying them. Ask for what you actually need instead, such as a three-sentence explanation of why it chose an approach. On the API, read the reasoning from summarized thinking blocks (display: "summarized").
If your Opus 5 integration ran with thinking disabled, three more changes go with the upgrade:
- Start at low effort and measure. If time to first token still matters, a system prompt line asking it to answer without deliberating can cut thinking further, but check quality when you add it.
- Drop the old thinking-off workarounds. Anthropic's Opus 5 mitigations for leaked tags and text-only tool calls addressed artifacts that appeared only with thinking off; re-test whether you still need them, and remove any rule that tells the model not to think.
- Read responses by block type. A response may begin with a
thinkingblock whose text is empty at the default display setting, so never assume the first block is text.
Hand Over the Whole Task and Name the Finish Line
Give Opus 5.5 the whole task in one message, say what "done" means in checkable terms, and say when to stop and ask. Anthropic says the model is at its best on multistep work inside a real repository, like pushing a change across a large code base until the tests pass. A clear finish line is how it knows it has arrived.
A prompt in that shape, written for your own project:
Move every caller of the legacy billing client to the new client.
Done means: no imports of the legacy client remain, the legacy client is deleted, and npm test passes.
Ask me only if a test fails for a reason you cannot explain.
If you remember something mid-run, type it and press Enter while Claude works. Claude Code queues the message, and if Claude is running tool calls, passes it to Claude as soon as those tool calls finish, within the same turn (interactive mode). Runs are longer now, so a follow-up is cheaper than a restart.
Keep Long Runs Going: CLAUDE.md Stop Rules and Unattended Agents
Opus 5.5 reports as it goes, and on long tasks it sometimes ends a turn to report rather than keep working: it announces a next step without starting it, offers to carry on, or lists decisions that block nothing. It responds well to instructions that name these stops, so name them.
In Claude Code, put a short rule in CLAUDE.md. Write it in your own words; the shape is:
If a step does not need my input, keep going, and put any status note in the same message as your next action.
Stop and ask only when you cannot continue without me, or before deleting data, force-pushing, or touching anything outside this repo.
Fewer stops also means fewer chances to catch a bad step, so the rule's last line keeps a check before risky actions; leave permission prompts on for destructive commands as well. If a run still ends with an offer to continue, reply "continue." For pair programming, write the opposite rule instead: a one-line plan before it starts and a short recap at the end. Opus 5.5 follows either.
In your own agent loop, some progress updates end the turn with text rather than a tool call (stop_reason: "end_turn"), and a loop that reads that as completion stops early. Anthropic's harness advice:
- Treat a text-only end of turn as a report, not proof the task is done. Keep the task's parts in a checklist the model updates, in a to-do tool or a file.
- If a turn ends with open items and no stated blocker, send a short user message naming the open items. Alternatively, have a smaller model check the conversation against your completion condition.
- Stop after two or three automatic continuations on the same task, so a genuinely stuck run ends and can be reviewed.
- If a background command or subagent is still running, wait for its output and return it to the model before treating the task as done.
Anthropic also publishes a system prompt paragraph for fully unattended agents that lists the four early-stop patterns and tells the model to carry on. Add anything like it at the end of the system prompt from the first request: adding it partway through changes the system prompt and invalidates the conversation's earlier thinking blocks. Leave it out of human-in-the-loop products, and expect somewhat more tool calls and output tokens per task.
Show Progress Updates to Your Users
If your product streams the model's notes between tool calls, they go silent on Opus 5.5 until you change one setting. These notes now arrive as progress-update thinking blocks rather than text blocks, and their text is empty at the default thinking.display. Set display: "updates" (beta, thinking-display-updates-2026-08-18 header) to receive a short summary of each note.
Three more levers control what users see:
- A message tool for verbatim content. If the model may need to hand the user a snippet mid-turn, give it a simple send-message tool and declare it from the first request, because adding a tool later invalidates earlier thinking blocks.
- Ask for the cadence. A system prompt line asking for a one-line statement of intent before the first tool call and a short recap at the end works, most of all in human-in-the-loop work.
- Nudge long silences. Count consecutive tool-calling steps with nothing for the user to read; after several (five, for example), append a one-line reminder as a turn-scoped system message, and stop after two or three reminders. Anthropic says this roughly halved the share of agentic coding tasks with a long silent stretch, at no measurable extra cost.
Subagents, Task Files, and Time Budgets
For an audit, a migration, or a review across a large code base, ask Opus 5.5 to split the work across subagents and to check each subagent's evidence before accepting it. Anthropic says early testers had Opus 5.5 coordinate parallel subagents on long audits and migrations with little oversight. Ask for one closing table (item, result, evidence) so the fan-out ends in something you can check. Our Claude Code subagent patterns cover the definitions and caps.
Keep the task list in a file. Long runs fill the context window, at which point Claude Code summarizes older turns; a checklist in a file such as TASKS.md survives that and shows you at a glance what is done and what is left. Read the file, not the scrollback. Our context management guide covers the rest of that discipline.
Give multiagent harnesses a clock. Opus 5.5 pays close attention to elapsed time. If you can estimate how long a task should take, have your harness append the elapsed time against a budget, in seconds, to each message it sends back; the model paces itself and usually finishes well before the budget. If you cannot predict a budget, show the elapsed time and add one system prompt sentence saying that time matters. In Anthropic's tests, a single Opus 5.5 agent at medium effort, told that time matters and shown the elapsed time, finished the typical DRACO research task in 47% less time at 60% lower cost, scoring 4.1 points lower; on HLE and a physics set the time saving was 14% or less. The budget is advisory, so keep your own timeout.
For multi-app automation, tell it to look around first. Across email, documents, spreadsheets, and CRM records, Opus 5.5 tends to get to work quickly. One system prompt sentence telling it to list and open every source that could be relevant before acting helped it complete noticeably more multi-app tasks in Anthropic's testing, at slightly more tool calls. Because it acts on what it finds, keep untrusted content out of those sources.
Review and Research Prompts That Mark What Wasn't Confirmed
Use Opus 5.5 as the first reviewer on every diff, and make research answers state their gaps. One early tester found Opus 5.5 on its lowest effort caught more bugs than Opus 5 on high, with fewer false alarms. On Anthropic's launch page, Deloitte puts numbers on it: at its lowest effort setting, Opus 5.5 caught 72% of known bugs in Deloitte's code reviews against Opus 5's 56% at high effort (Anthropic).
- Review prompt: ask it to review the branch against main and list only problems you would block the merge for, each with file and line, why it is wrong, and how to show it fails.
- Research prompt: add an instruction to flag anything it could not confirm and name the places it checked. Knowing what it could not find is useful, and asking for it puts the gaps where you will see them.
- End-of-run summary: Opus 5.5's final summary says what it did, what it found, and what it needs from you. Read the part that is waiting on you first. To fix the format, say so in CLAUDE.md, for example headings for what is blocked on you, what changed, and what it found.
Charts, Screenshots, Documents, and Frontend Design
Attach the chart, diagram, or screenshot itself and ask a specific question; don't retype the numbers. In Anthropic's testing, Opus 5.5 on low effort out-read Opus 5 on its highest setting when pulling values from dense charts. It is also better where meaning depends on position, such as which boxes an arrow connects. For the densest inputs, two things still add accuracy: higher-resolution images, and image tools (a container with PIL and OpenCV, or at least a crop tool). Without tools, raising effort helps on technical drawings but does little for charts.
For documents, ask it to check a long plan or deck for anything that contradicts itself and to quote each problem with its location; Anthropic reports it catching a date on the wrong weekday in a long planning thread. When you want a spreadsheet or document, ask for the finished file, not an outline.
For frontend work, list the styles you don't want. Without design direction, Opus 5.5 reaches for a handful of default styles, and a vague request to avoid a generic look just trades one default for another. Name specific patterns instead (Anthropic's examples include cream backgrounds, numbered section labels, monospace labels, and pill-shaped buttons), look at what it chose instead, and extend the list.
Chat and Projects: Settled Answers and Pasted Text
In long chats, Opus 5.5 sometimes goes back over an earlier answer while thinking about a short follow-up, which slows replies. If follow-ups feel slow, add a project or system prompt instruction that answered questions are settled unless you ask about them or point out a problem. Leave it out where you want earlier work re-examined, such as long analyses or agentic tasks where a later step can expose an earlier mistake; Anthropic warns it can make the model less likely to point out its own earlier error.
Mark pasted text. Anthropic rates Opus 5.5 as the most resistant Opus model yet to indirect prompt injection, and it can also resist instructions hidden in text a user pasted, if you mark that text. Wrap each pasted block in matching opening and closing tags that carry the same short random ID your application generates, and add a system prompt note that instructions inside those tags are followed only when the user's own message asks for it. Measure the effect, since it can make the model slightly more cautious, and treat the tags as one guardrail among several because they can be imitated.
In the Claude apps, first check that the model picker says Opus 5.5.
When a Message Is Flagged
Opus 5.5 is the first Opus model to ship with biology and cybersecurity safeguards at the level of Claude Fable 5.1, so plan for occasional flags. Most flagged messages in Claude apps and Claude Code move to an older model, and your work continues there. You can still look for vulnerabilities in source code; high-risk dual-use security work is not allowed. The check covers everything in the conversation, including files and search results, so a flag can come from earlier content.
| Where | What you see | What to do |
|---|---|---|
| Claude apps | A notice starting "Switched to" plus an older model's name | Pick Opus 5.5 in the model picker, or start a new chat to avoid re-flagging; turn off "Switch models when a message is flagged" under Settings, Capabilities to be asked first |
| Claude Code | A notice naming the older model | Run /model to switch back, press Esc twice to edit and retry, change the flag setting in /config (or set switchModelsOnFlag to false), and use /feedback for a wrong flag |
| Claude API | stop_reason: "refusal" with a stop_details category | Configure server-side fallback (fallbacks: "default", beta), the SDK middleware, or your own retry |
In Claude Code, a biology flag re-runs the request on Opus 5 and a cybersecurity flag re-runs it on Opus 4.8. The effort level carries over: an Opus 5.5 session at medium that falls back to Opus 4.8 stays at medium, even though Opus 4.8's own default is high. Run /effort if you want a different level there.
When Fast Mode Pays
Turn on fast mode with /fast when you are waiting on each reply, and leave it off for long unattended runs. It is the same model with faster output, a research preview, and it costs more per token than standard mode.
- Speed: up to 2.5x higher output tokens per second from Claude Opus 5.5, Claude Opus 5, and Claude Opus 4.8 (fast mode).
- Price: $8 and $40 per million input and output tokens on Opus 5.5, against $10 and $50 on Opus 5 and Opus 4.8 (Claude Code fast mode).
- Billing: on a Claude subscription, your account needs usage credits turned on, which lets you bill past your plan's included usage.
- Turn it on early: enabling it the first time in a conversation bills the whole existing context once at the fast mode uncached input price, so it is cheapest at the start.
- Platforms: on the API, fast mode for Opus 5.5 runs on the Claude API only; it is not available on Amazon Bedrock, Claude Platform on AWS, Google Cloud, or Microsoft Foundry.
API Checklist From Opus 5
Swap the model ID to claude-opus-5-5, then fix the request settings Opus 5.5 rejects. In Claude Code, /claude-api migrate runs the bundled Claude API skill to do most of this across a code base.
| Setting | On Opus 5.5 | Replace with |
|---|---|---|
thinking: {"type": "disabled"} or a manual budget_tokens | Rejected | Omit thinking, set effort |
tool_choice of any or a named tool | Rejected | auto plus strict tool use or structured outputs, and say in the prompt when the tool applies |
| Editing the system prompt, tools, or earlier turns mid-conversation | Replayed thinking blocks rejected on newer accounts | Append-only conversations; mid-conversation system messages |
computer_20251124 on the Claude API or Google Cloud | Rejected | computer_toolset_20260801 |
Non-default temperature, top_p, top_k, or an assistant prefill | Rejected | Prompting, structured outputs, or system prompt instructions |
Each rejected setting returns an error: a request that disables thinking or sets a manual budget returns a 400 invalid_request_error, and so does forced tool choice. Anthropic enforces the thinking-block check by default on accounts created on or after August 31, 2026 (00:00 UTC), on the Claude API and cloud platforms. Opus 5.5 does not support Priority Tier. Opus 5.5 costs $4 and $20 per million input and output tokens, below Opus 5's $5 and $25, and batch processing is half price at $2 and $10 (What's new in Claude Opus 5.5). The launch coverage has the cost-per-task charts, and our model picks explain why Opus 5.5 at medium is the top pick.
Sources
- Anthropic (Addy Osmani), "Getting the most out of Opus 5.5 in Claude and Claude Code" (September 22, 2026): https://claude.dev/blog/getting-the-most-out-of-opus-5-5/
- Anthropic, "Prompting Claude Opus 5.5": https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5-5
- Anthropic, "What's new in Claude Opus 5.5": https://platform.claude.com/docs/en/models/opus-5-5/whats-new-opus-5-5
- Anthropic, "Claude Opus 5.5 migration guide": https://platform.claude.com/docs/en/models/opus-5-5/migration-guide
- Anthropic, "Effort": https://platform.claude.com/docs/en/build-with-claude/effort
- Anthropic, "Optimizing for cost and intelligence": https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence
- Anthropic, "Fast mode" (Claude API): https://platform.claude.com/docs/en/build-with-claude/fast-mode
- Claude Code docs, "Model configuration": https://code.claude.com/docs/en/model-config
- Claude Code docs, "Fast mode": https://code.claude.com/docs/en/fast-mode
- Claude Code docs, "Interactive mode": https://code.claude.com/docs/en/interactive-mode
- Anthropic, "Introducing Claude Opus 5.5" (September 22, 2026): https://www.anthropic.com/claude-opus-5-5
Read next
More practices and workflows- Claude
Claude Opus 5.5: Fable 5.1-Level Scores at $4 and $20, With Benchmarks, Effort Costs, and Migration Guide
Claude Opus 5.5, released September 22, 2026, beats Claude Fable 5.1 on every benchmark Anthropic published and costs $4 and $20 per million tokens, 60% less than Fable 5.1. It scores 66.4% on Terminal-Bench 4.0 against 57.9% for GPT-6 Astra. Thinking is always on, the default effort is medium, and four API changes can break Opus 5 code.
- Claude
Prompting Claude Opus 5: Trim the Verbosity, Delete the Verification
Claude Opus 5 needs opposite prompting from its predecessors: you prompt for conciseness because effort no longer controls visible length, and you delete verification and double-check instructions because the model already does both. Constrain scope on narrow tasks, cap subagent spawning for cost, and keep thinking enabled at low effort rather than disabling it.
- Claude
Prompt Engineering in 2026: The Playbook That Works Across Claude and GPT
Prompt engineering in 2026 has two layers: cross-model fundamentals (clear direct instructions, motivated rules, 3-5 structured examples, XML boundaries, long documents before the question) and model-specific overrides that matter more than ever, because the newest models need instructions deleted as often as added. Start here, then apply the guide for your model.
- Claude Code
Claude Code Subagent Patterns: 10 Reusable Agent Definitions
A Claude Code subagent is a delegated worker with its own context window, defined as a markdown file with YAML frontmatter in .claude/agents/. Subagents can edit files when you grant Edit or Write, nest three layers deep by default, and run 20 at a time. These 10 definitions cover the highest-value delegations.
Frequently Asked Questions
What effort level should I use for Claude Opus 5.5?
Start at medium and set it explicitly. In Anthropic's testing, Opus 5.5 at medium matched or beat Opus 5 at high on coding and knowledge-work evals. Reserve xhigh and max for work where you have measured a quality gain, and re-run your own effort sweep instead of reusing Opus 5 settings.
Should I still tell Claude Opus 5.5 to think carefully?
No. Opus 5.5 thinks before every reply and sets its own depth. When Anthropic removed a think-carefully line in a chat product, replies started sooner and quality showed no clear drop. To change how much it thinks, change the effort level instead.
Can I turn off thinking on Claude Opus 5.5?
No. Adaptive thinking is always on, and a request that disables thinking or sets a manual thinking budget returns a 400 error. If your Opus 5 integration ran with thinking off, start at low effort and measure, because lowering effort cuts thinking more reliably than prompt instructions.
Why does Claude Opus 5.5 stop partway through a long task?
It reports progress as it works, and some of those reports end the turn without a tool call. In Claude Code, add a CLAUDE.md rule that says when to keep going and when to stop. In your own agent loop, treat a text-only turn as a report and cap automatic continuations at two or three.
What happens when a Claude Opus 5.5 message is flagged?
Most flagged messages move to an older model and your work continues there. In Claude Code, a biology flag re-runs the request on Opus 5 and a cybersecurity flag re-runs it on Opus 4.8. Run /model to switch back, or change the flag setting in /config so Claude Code asks first.
Is fast mode worth it on Claude Opus 5.5?
Use it for back-and-forth work where you read every reply. On Opus 5.5 it costs $8 and $40 per million input and output tokens and delivers up to 2.5x the output speed. On a Claude subscription, Claude Code needs usage credits turned on before /fast works.