Skip to content AI News: Qwen-Image-2.1-Turbo: Alibaba's open image model now draws in 8 steps →
AI News · Oct 10, 2026 · Claude

Claude Haiku 5.5 is out: a 1M-token context and up to 90% lower prices

Anthropic released Claude Haiku 5.5 on October 7, its new small model. It reads up to a million tokens, costs up to 90% less than Haiku 4.5, and on Anthropic's own tests lands much closer to Sonnet 5.5 than the old Haiku did. Free Claude users can pick it today. Here's what changed and how to prompt it.

ModelClaude Haiku 5.5 (claude-haiku-5-5)
WhereClaude apps (free too), API, Copilot, Claude Code
API pricefrom $0.10 / $0.50 per 1M tokens
Try it2 prompts
Claude Haiku 5.5 is out: a 1M-token context and up to 90% lower prices

What happened

On October 7, Anthropic released Claude Haiku 5.5, calling it the cheapest, fastest and most capable small model it has released. It replaces Haiku 4.5 as the small model in the lineup, under Sonnet 5.5 and Opus 5.5. Anthropic pitches it for high-volume work: sorting, routing, pulling data out of documents, customer support, and subagents that a bigger model hands small jobs to.

Free, Pro, Max, Team and Enterprise users can select it in Claude on the web, iOS and Android. The same day, GitHub added it to Copilot, rolling out gradually.

What changed

  • Context. Up to 1M tokens of input, up from 200K on Haiku 4.5, and up to 128K tokens of output, up from 64K.
  • Price. Anthropic says Haiku 5.5 costs 90% less than Haiku 4.5 for requests up to 100,000 tokens and 50% less above that, around 75% less on average.
  • Thinking. It is the first Haiku with effort levels: low, medium (the default), high, xhigh and max. It decides on its own how much to think at each level.
  • Knowledge. Training data runs to June 2026. It reads text and images and writes text.
  • Tokens. It counts text with Anthropic's newer tokenizer, so the same text comes to roughly 30% more tokens than on Haiku 4.5. The real saving is smaller than the price list suggests, which is why Anthropic quotes about 75% rather than 90%.

The price, for API users

Per 1M tokensHaiku 4.5Haiku 5.5, request up to 100K tokensHaiku 5.5, over 100K
Input$1$0.10$0.50
Output$5$0.50$2.50
Cache read$0.10$0.01$0.05

The Batch API takes half off again. The model ID is claude-haiku-5-5 on the Claude API, Google Cloud and Microsoft Foundry, and anthropic.claude-haiku-5-5 on Amazon Bedrock.

How close it gets to Sonnet

Anthropic's launch page compares it with Haiku 4.5 and Sonnet 5.5. These are Anthropic's own runs, and the page doesn't say at which effort level.

TestHaiku 4.5Haiku 5.5Sonnet 5.5
OSWorld 2.1 (using a computer)15.7%72.4%83.9%
Humanity's Last Exam, with tools18.7%57.4%64.5%
Terminal-Bench 4.0 (agentic coding)0.0%39.2%70.6%
Chartography (reading charts), no tools6.4%46.4%61.6%

The gap that's left is in long coding jobs. Anthropic says so itself: Sonnet 5.5 and Opus 5.5 remain the better choice for complex agentic coding of the kind Terminal-Bench measures. GitHub says that in its own testing Haiku 5.5 matched Claude Sonnet 5 on many coding tasks while using fewer tokens and steps.

Where to use it

  • Claude apps. Pick Haiku 5.5 in the model menu, on the free plan too.
  • GitHub Copilot. In the model picker on Copilot Pro, Pro+, Max, Business and Enterprise; Copilot Free isn't listed. It works in VS Code, Visual Studio, JetBrains, Xcode, Eclipse, the Copilot CLI and github.com. Business and Enterprise admins can switch it off in the model policy.
  • Claude Code. Since version 2.1.293 it is the default Haiku when Claude Code talks to Anthropic's API, so /model haiku picks it. Through Bedrock, Google Cloud or Foundry, the haiku alias still means Haiku 4.5 for now.
  • API and clouds. The Claude API, Amazon Bedrock, Claude Platform on AWS, Google Cloud and Microsoft Foundry.

What it means for your prompts

Anthropic says prompts written for Haiku 4.5 should work without changes. A few habits are worth dropping, and developers have hard changes to make:

  • Don't ask it to "answer directly" to skip thinking. Anthropic found the line didn't stop it. In the API, use a low effort level instead; in the apps, keep the question short and say what format you want.
  • Pick effort by job. Anthropic suggests low for chat and simple bulk calls, medium for most work including agentic coding, and high for knowledge work and strict instruction following. If you need xhigh or max, test Sonnet 5.5 against it first.
  • Give it the date when it can search the web, so it doesn't reason from June 2026.
  • For developers: temperature, top_p and top_k now return an error, prefilling the assistant's reply is gone, and budget_tokens is replaced by effort. Thinking is hidden unless you ask for a summary.

Our Claude prompts run on Haiku 5.5 as they are. Short, structured jobs suit it best: extraction, sorting, rewriting to a fixed format.

Try it

Sort a pile of messages in one go
Here are [number] messages from [customers / my inbox / a feedback form], separated by ---. Sort each into exactly one of these categories: [category 1], [category 2], [category 3], other.

Return a table with three columns: message number, category, and the one phrase from the message that decided it. Then list any message you weren't sure about, and why. Don't summarize the messages and don't add categories of your own.

---
[paste the messages]
Pull facts out of a long document
I'm attaching [a contract / a report / a set of meeting notes]. Pull out every [deadline / amount / commitment / name of a person responsible] it contains.

For each one, give: the item, the exact wording from the document in quotation marks, and where it appears (section or page). If two places in the document disagree, list both and say so. Don't include anything that isn't written in the document. If there's nothing to find, say that plainly.

Sources

Weekly

New prompts in your inbox

One email a week: the best new prompts and one short guide. Unsubscribe any time. Privacy policy