Claude Haiku 5.5 Launches at $0.10 Input and $0.50 Output per Million Tokens Under 100K, 90% Cheaper Than Haiku 4.5, With Adjustable Effort for the First Time

On October 7 Anthropic released Claude Haiku 5.5 (model ID claude-haiku-5-5), available in the API and on AWS, Google Cloud and Microsoft Azure. Pricing is tiered by prompt length: $0.10 per million input tokens and $0.50 per million output tokens up to 100K tokens, and $0.50 and $2.50 above that; against Haiku 4.5's $1 and $5, that is 90% cheaper under 100K and 50% cheaper above, about 75% cheaper on average by Anthropic's count, though an updated tokenizer uses slightly more tokens per task. It is the first Haiku model with five adjustable effort levels from Low to Max, and Anthropic calls it its fastest model at standard speed. In Anthropic's benchmarks it scores 72.4% on the OSWorld 2.1 offline subset (Haiku 4.5: 15.7%) and 39.2% on Terminal-Bench 4.0, below Sonnet 5.5's 70.6%. Sonnet 5.5 cache-read prices were also halved.

Price is the biggest change

On October 7 Anthropic released Claude Haiku 5.5, API model ID claude-haiku-5-5, also available on AWS, Google Cloud and Microsoft Azure. Pricing has two tiers by prompt length:

  • Up to 100K tokens: $0.10 per million input tokens, $0.50 per million output tokens, $0.01 for cache reads;
  • Over 100K tokens: $0.50 input and $2.50 output.

Haiku 4.5 was $1 and $5, so short prompts are 90% cheaper and long prompts 50% cheaper, about 75% cheaper on average according to Anthropic. Note that an updated tokenizer makes the same task use somewhat more tokens, so real savings will be a little smaller than the list prices suggest.

It is also the first Haiku with adjustable effort, in five levels: Low, Med, High, Xhigh and Max. Anthropic calls it its fastest model yet at standard speed (excluding Opus in Fast Mode); customer Asana reports more than 30% lower latency, and Box says latency is about half that of Haiku 4.5.

Where its capability sits

In Anthropic's benchmarks it is a generational jump over Haiku 4.5: 72.4% vs. 15.7% on the OSWorld 2.1 offline subset, and 39.2% vs. 0% on Terminal-Bench 4.0. It still trails Anthropic's own Sonnet 5.5 by a wide margin: 39.2% vs. 70.6% on Terminal-Bench 4.0, and 72.4% vs. 83.9% on OSWorld. Against OpenAI's lightweight GPT-6 Luna, it is higher on every item in Anthropic's table, for example 72.4% vs. 48.9% on OSWorld.

Anthropic positions it clearly: high-volume, cost-sensitive work such as summaries, context compaction, database queries, classification and subagents, while recommending Sonnet 5.5 and Opus 5.5 for complex agentic coding.

What it means

For teams already running classification, routing or summarization pipelines on Haiku 4.5, this is close to a cost cut by swapping the model name, and worth a regression pass on your own data soon, paying particular attention to the real token bill after the tokenizer change. The same day, Sonnet 5.5's cache-read price dropped from $0.20 to $0.10, which Anthropic says makes most agentic work about 20% cheaper. Splitting work so Sonnet runs the main agent and Haiku 5.5 handles subtasks will lower costs noticeably. All benchmarks and customer figures are as published by Anthropic.

via: Anthropic release page