AI Catchup

Anthropic Releases Claude Haiku 5.5, Its Fastest Model and First Haiku With Effort Levels

By 12 min read

Anthropic released Claude Haiku 5.5 on October 7, 2026, its fastest model to date and the first Haiku with effort levels. It scores 72.4% on OSWorld 2.1 against Haiku 4.5's 15.7%, and costs $0.10 per million input and $0.50 per million output tokens for prompts up to 100,000 tokens, 90% less than Haiku 4.5.

Anthropic released Claude Haiku 5.5 on October 7, 2026, the small model in its Claude 5.5 family and, by Anthropic's account, its fastest model to date. It is built for high-volume work: summaries, compaction, database queries, classification, and subagents working under Opus 5.5 or Sonnet 5.5. It is the first Haiku-class model with an adjustable effort setting, and it is available now on all platforms, including Amazon Web Services, Google Cloud, and Microsoft Azure (Anthropic).

The capability jump over Haiku 4.5 is large: 72.4% on OSWorld 2.1 against 15.7%, and 39.2% on Terminal-Bench 4.0 against 0.0%. It also costs less. For prompts up to 100,000 tokens it is $0.10 per million input tokens and $0.50 per million output tokens, 90% less than Haiku 4.5, and Anthropic says it runs around 75% cheaper on average once its larger token counts are included (Anthropic).

The verdict: move every Haiku 4.5 workload to Haiku 5.5, and give it the subagent, summarization, and compaction work you were paying Sonnet prices for. Do not make it your main coding agent: on Terminal-Bench 4.0 it scores 39.2% to Sonnet 5.5's 70.6%, and Anthropic itself says Sonnet 5.5 and Opus 5.5 remain the better choices for complex agentic coding (Anthropic).

Key Takeaways

  • Released: October 7, 2026, available now on the Claude Platform, Amazon Web Services, Google Cloud, and Microsoft Azure (Anthropic).
  • Speed: Anthropic's fastest model to date at standard speed, though Opus models in Fast Mode run quicker (Anthropic).
  • Model ID: claude-haiku-5-5, a fixed ID with no date suffix and no separate alias (migration guide).
  • Price: $0.10 input, $0.50 output, and $0.01 cache reads per million tokens up to 100,000 tokens; five times that above the line (Anthropic).
  • Effort levels: the first Haiku-class model with an adjustable effort setting; the default is medium (Anthropic; models overview).
  • Benchmarks: 72.4% on OSWorld 2.1 (offline subset) against 15.7% for Haiku 4.5 and 48.9% for GPT-6 Luna (Anthropic).
  • Claude Code: the haiku alias resolves to Haiku 5.5 on the Anthropic API from v2.1.293 (Claude Code docs).
  • Breaking changes: thinking budgets, sampling parameters, assistant prefill, and the old computer use tool all return 400 errors (migration guide).
  • Also announced: Sonnet 5.5 cache reads drop to $0.10 per million tokens, and Max and Team plans get monthly API credits (Anthropic).

Benchmarks: Haiku 5.5 vs Haiku 4.5, GPT-6 Luna, and Sonnet 5.5

Haiku 5.5 beats Haiku 4.5 on every benchmark where Haiku 4.5 has a score and beats GPT-6 Luna on all six where Luna has a score. It trails Sonnet 5.5 on all of them. The biggest jump over Haiku 4.5 is in computer use, 72.4% against 15.7% on OSWorld 2.1, followed by visual reasoning (46.4% against 6.4% on Chartography) and Terminal-Bench 4.0, where Haiku 4.5 scored 0.0% (Anthropic).

BenchmarkHaiku 5.5Haiku 4.5GPT-6 LunaSonnet 5.5
GDPval-AA v2.1 (Elo)162073514371840
AA-Briefcase v1.1 (Elo)157861413361824
OSWorld 2.1, offline subset72.4%15.7%48.9%83.9%
Humanity's Last Exam, no tools45.9%10.2%n/a56.9%
Humanity's Last Exam, with tools57.4%18.7%n/a64.5%
Terminal-Bench 4.039.2%0.0%16.4%70.6%
FrontierCode 1.1 (Main)46.4%n/a42.4%52.1%
Chartography, no tools46.4%6.4%29.1%61.6%

Scores are Anthropic's; bold marks the best of the three small models. Anthropic lists Sonnet 5.5 for reference and notes its FrontierCode score was run at Xhigh effort. Methodology is in the Haiku 5.5 system card.

Cursor publishes its own numbers: Haiku 5.5 scores 48.4% on CursorBench at max effort and 30.9% at low effort with thinking on. With thinking off it scores 22.1% to 26.2% (Cursor).

Pricing and the 100,000-Token Line

Haiku 5.5 is priced by prompt length: a prompt over 100,000 tokens pays the higher rates. Other Claude 4.6 and later models bill their full 1M context window at standard pricing (Anthropic pricing).

Per million tokensHaiku 5.5, up to 100KHaiku 5.5, over 100KHaiku 4.5Sonnet 5.5
Input$0.10$0.50$1.00$2.00
Output$0.50$2.50$5.00$10.00
Cache writes$0.125$0.625$1.25$2.50
Cache reads$0.01$0.05$0.10$0.10

Prices are from Anthropic's announcement. The Batch API halves the base rates to $0.05 input and $0.25 output up to 100,000 tokens, and $0.25 and $1.25 above it (Anthropic pricing).

Anthropic says prompts up to 100,000 tokens made up around 90% of requests to Haiku 4.5, which is where the 90% discount applies. Above the line the discount is 50%. The roughly 75% average saving also accounts for Haiku 5.5's newer tokenizer, which uses more tokens for the same work (Anthropic).

Watch the line, not just the rate. The migration guide says the same text produces approximately 30% more tokens on Haiku 5.5 than on Haiku 4.5, so a prompt that sat just under 100,000 tokens on Haiku 4.5 can cross into the higher price on Haiku 5.5. Recount prompts with model set to claude-haiku-5-5 before you budget (migration guide).

Specs and Effort

Haiku 5.5 has a 1M-token context window, 128K maximum output tokens, adaptive thinking, and a medium default effort. Its reliable knowledge cutoff is June 2026, and Anthropic lists its retirement as not sooner than October 7, 2027 (models overview).

Effort is the main cost lever. Adaptive thinking is on by default, so a response can start with thinking blocks even when the request does not ask for them. Where Haiku 4.5 ran without thinking or with a small budget, the migration guide says to choose a lower effort level; at lower levels the model can skip thinking entirely on simple requests (migration guide).

Anthropic calls Haiku 5.5 its fastest model to date at each model's standard speed, though it runs less quickly than the Opus models in Fast Mode (Anthropic).

What Changes in Claude Code

On the Anthropic API, Claude Code's haiku alias resolves to Haiku 5.5 from v2.1.293. The same alias is the model Claude Code uses for background functionality, set by ANTHROPIC_DEFAULT_HAIKU_MODEL. Run claude update if you are on an older version (Claude Code docs).

  • Pick it directly: run /model claude-haiku-5-5 in a session, or start with claude --model claude-haiku-5-5.
  • Context: Haiku 5.5 runs with the 1M context window on every plan, with no [1m] suffix, and sessions auto-compact at about 967K tokens by default.
  • Other providers: Claude Platform on AWS, Amazon Bedrock, Google Cloud's Agent Platform, and Microsoft Foundry still resolve haiku to Haiku 4.5.
  • Old sessions: a session saved on Haiku 4.5 under the haiku setting resumes on Haiku 5.5.

All four points are from the Claude Code model configuration docs, which also note that a Haiku 5.5 request costs more per token once its prompt passes 100K tokens. If you pay API rates, set a smaller window with /autocompact to compact earlier.

Haiku 5.5 in Cursor

Cursor lists Claude Haiku 5.5 in its model picker; add it from Cursor Settings > Models. It supports context windows up to 1M tokens, has access to all of Cursor's agent tools, and draws from the third-party Other Models usage pool at $0.10 input and $0.50 output per million tokens up to 100k. Cursor notes that thinking can be turned off at low, medium, and high effort, at a cost in quality (Cursor).

Breaking Changes for API Users

The migration guide lists ten changes for code that calls Haiku 4.5. If you use Claude Managed Agents, only the model name changes (migration guide).

Haiku 4.5 codeOn Haiku 5.5Fix
thinking: {"type": "enabled", "budget_tokens": N}400 errorUse {"type": "adaptive"} and set output_config.effort
temperature, top_p, top_k400 unless you send only temperature 1 or only top_p 0.99; any top_k, or both together, failsRemove them and steer with prompting
Assistant prefill as the last message400 error, even with thinking offEnd with a user turn; use structured outputs for format
computer_20250124 tool400 error on the Claude API and Google CloudMove to computer_toolset_20260801
Reading the first content block as the answerThinking blocks can come firstSelect blocks by type
Summarized thinkingEmpty thinking field by defaultSet "display": "summarized"
Priority Tier commitmentNot supportedPlan capacity separately

Three more points from the guide: requests can stop with stop_reason: "refusal" with no server-side fallback; thinking blocks work only in the account that produced them or a linked one, and are dropped elsewhere; and a thinking block sent back after a change to system, tools, or earlier messages returns a 400 error, which accounts created before August 31, 2026 see only if they set thinking.block_binding.prefix_mismatch_behavior. In Claude Code, /claude-api migrate this project to claude-haiku-5-5 applies the model ID swap and parameter changes for you (migration guide).

Haiku 5.5 also supports the new browser use tool (browser_toolset_20260801) on the Claude API and Google Cloud; Haiku 4.5 does not. Anthropic is adding computer use and browser use support in beta to its Python and TypeScript SDKs alongside this launch (migration guide; Anthropic).

Do not wait on Haiku 4.5. Anthropic's deprecations page lists Haiku 4.5's retirement as not sooner than October 15, 2026 (model deprecations).

What Early Testers Report

Anthropic published six customer reports from early testing. They describe speed and cost gains on short, high-volume work (Anthropic):

  • Asana: over a 30% reduction in task-completion latency and up to 2.5x faster inference per agent turn, against its current model.
  • HubSpot: 92.8% on its CRM suite, averaged over three runs, the best score it has seen there.
  • AlphaSense: 0.84 against Haiku 4.5's 0.76 on 400 document questions, for a feature that runs about 8M calls a week.
  • Box: 11 points higher than Haiku 4.5 at about half the latency.
  • Cognition: Devin Fusion holds a FrontierCode score of 66.2 with Haiku 5.5 as the sidekick; Cognition offers it in the Devin CLI with Opus 5.5 as the lead.

Sonnet 5.5 Price Cut and Max and Team API Credits

Anthropic cut Sonnet 5.5's cache-read price by half the same day, from $0.20 to $0.10 per million tokens. Because cache reads are a large share of agentic token use, Anthropic says this makes Sonnet 5.5 around 20% cheaper on most agentic tasks (Anthropic).

Max and Team subscribers also get a monthly credit for the Claude Platform, rolling out this week (Anthropic; Claude Help Center):

PlanMonthly API credit
Max 5x$100 a month
Max 20x$200 a month
Team$20 per Standard seat, $100 per Premium seat, pooled and capped at $500
Free, Pro, EnterpriseNot eligible

The credit covers the Claude API, the Console Playground, Claude Managed Agents, and the Claude Agent SDK. It does not cover interactive Claude Code, extra usage, or Claude on Amazon Bedrock, Google Cloud Vertex AI, or Microsoft Foundry. Unused credit expires at the end of each billing cycle. To claim it, open Settings > Billing on claude.ai (Team owners use Organization settings > Billing) and link one Console organization; you cannot change the link yourself later (Claude Help Center).

Safety and Safeguards

Anthropic reports far fewer instances of misaligned behavior than Haiku 4.5 and a lower willingness to cooperate with misuse. Its cybersecurity safeguards are more restrictive than Haiku 4.5's but somewhat less restrictive than those on other recent models: they permit a wider range of defensive tasks than Sonnet 5.5's safeguards and still block penetration testing (Anthropic).

When to Choose Haiku 5.5

Choose by task scope, not by price alone. Anthropic positions Haiku 5.5 for narrowly scoped work that was cost-prohibitive on earlier Claude models (Anthropic).

  • Use Haiku 5.5 for: subagents under an Opus 5.5 or Sonnet 5.5 lead, compaction, summaries, classification, database queries, live support, and browser use.
  • Use Sonnet 5.5 for: well-scoped coding and feature work, now cheaper with $0.10 cache reads.
  • Use Opus 5.5 for: long-running, open-ended agentic coding.
  • Keep prompts under 100,000 tokens wherever you can; above it every rate is five times higher.

Sources

Related coverage

Frequently Asked Questions

What is Claude Haiku 5.5?

Claude Haiku 5.5 is the small model in Anthropic's Claude 5.5 family, released October 7, 2026. Anthropic built it for high-volume work such as summaries, compaction, classification, and subagents, and calls it its fastest model to date. It is available now on all platforms, with the model ID claude-haiku-5-5 (anthropic.claude-haiku-5-5 on Amazon Bedrock).

How much does Claude Haiku 5.5 cost?

For prompts up to 100,000 tokens, Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output tokens, with cache reads at $0.01. Prompts over 100,000 tokens cost five times as much: $0.50 input and $2.50 output.

Is Claude Haiku 5.5 cheaper than Haiku 4.5?

Yes. Anthropic prices it 90% below Haiku 4.5 for requests up to 100,000 tokens and 50% below for longer ones. After counting the new tokenizer, which uses more tokens for the same text, Anthropic puts the average saving at around 75%.

Can Claude Haiku 5.5 replace Sonnet 5.5 for coding?

Not for complex agentic coding. Haiku 5.5 scores 39.2% on Terminal-Bench 4.0 against 70.6% for Sonnet 5.5, and Anthropic says Sonnet and Opus remain the better choices there. Use Haiku 5.5 for subagents, summaries, compaction, and other narrowly scoped work.

Does Claude Code use Haiku 5.5?

Yes, on the Anthropic API from Claude Code v2.1.293. The haiku alias, which also handles background functionality, resolves to Haiku 5.5 there. Claude Platform on AWS, Amazon Bedrock, Google Cloud, and Microsoft Foundry still resolve haiku to Haiku 4.5.

What breaks when I move from Haiku 4.5 to Haiku 5.5?

Thinking budgets, non-default temperature or top_p, any top_k, assistant prefill, and the computer_20250124 tool all return 400 errors. Token counts rise about 30% for the same text, Priority Tier is not supported, and requests can now stop with a refusal.

Is Claude Haiku 5.5 available in Cursor?

Yes. Add it from Cursor Settings > Models. Cursor lists support for context windows up to 1M tokens and a CursorBench score of 48.4% at max effort, and bills it from the third-party Other Models usage pool.

Get the weekly AI Catchup

Tools, practices, and what matters, in your inbox every week.