
Anthropic shipped Claude Haiku 5.5 on October 7, 2026, and the headline is not the benchmark table. It’s the price sheet. For prompts up to 100k tokens, input costs $0.10 and output $0.50 per million tokens. Haiku 4.5 was $1 and $5. Anthropic says the average saving is about 75%, with 90% on prompts under 100k and 50% above that.
Cheap models are only interesting if you know where to put them. Here is where I’d put this one.
The numbers#
Anthropic’s own comparison, Haiku 5.5 vs Haiku 4.5 vs GPT-6 Luna vs Sonnet 5.5:
| Benchmark | Haiku 5.5 | Haiku 4.5 | GPT-6 Luna | Sonnet 5.5 |
|---|---|---|---|---|
| Terminal-Bench 4.0 | 39.2% | 0.0% | 16.4% | 70.6% |
| FrontierCode 1.1 (Main) | 46.4% | n/a | 42.4% | 52.1% |
| OSWorld 2.1 (offline subset) | 72.4% | 15.7% | 48.9% | 83.9% |
| Humanity’s Last Exam (no tools) | 45.9% | 10.2% | n/a | 56.9% |
| GDPval-AA v2.1 (Elo) | 1620 | 735 | 1437 | 1840 |
Two readings are possible. The generous one: a small model now beats a rival’s small model by wide margins and is competitive on FrontierCode. The honest one: on Terminal-Bench 4.0, the benchmark closest to real agentic coding, Haiku 5.5 reaches 39.2% while Sonnet 5.5 reaches 70.6%. That is not a gap you paper over with a clever prompt.
Anthropic agrees with the honest reading. Its own positioning is high-volume, narrow work: summaries, compaction, classification, database queries and subagents. For complex agentic coding it points you to Sonnet 5.5 and Opus 5.5. Rare, and welcome, for a vendor to say that about its own launch.
Effort levels come to the small tier#
Haiku 5.5 is the first Haiku-class model with an adjustable effort setting: Low, Med, High, Xhigh and Max. That matters more than it sounds. Until now, small models were a fixed point: fast and shallow. With effort levels you can run the same model id for a routing decision at Low and for a harder extraction at High, without maintaining two prompts for two models.
One caveat from the announcement: the updated tokenizer uses slightly more tokens per task, so compare cost per task, not just per token.
Where it fits in a Claude Code setup#
Claude Code’s changelog already reflects the launch. Version 2.1.293 added claude-haiku-5-5 as the default Haiku model on the Anthropic API. The same release added agentType to the subagentStatusLine payload so scripts can distinguish custom subagent types, which is exactly the plumbing you need once you start mixing models across a swarm.
The pattern I’d use:
- Main loop on Sonnet 5.5 or Opus 5.5. Planning, multi-file edits, debugging. This is where the 70.6% vs 39.2% gap lives, and where mistakes compound.
- Haiku 5.5 for fan-out subagents. Repo search, “which files touch this symbol”, log triage, summarising a long test run into twelve lines.
- Haiku 5.5 for compaction. Summarising a long session is a bounded task, and you do it often. At $0.10 input it stops being a line item.
- Haiku 5.5 as a classifier in your pipeline. Is this PR touching auth? Is this ticket a bug or a feature request? Run it on every event.
This is the spec-driven angle. If your spec decomposes into small, well-defined tasks with clear acceptance criteria, the cheap model can execute many of them. If your spec is vague, no model tier saves you, and the cheap one fails first. Good specs are what let you spend less per unit of work.
Sonnet 5.5 got cheaper too#
Easy to miss: in the same post Anthropic halved Sonnet 5.5’s cache-read price, from $0.20 to $0.10 per million tokens. Agentic loops re-read a lot of context, so cache reads dominate the bill. Anthropic estimates this cuts Sonnet 5.5’s cost on most agentic tasks by about 20%. For many teams that is a bigger saving than switching anything to Haiku.
Sonnet 5.5 pricing from the page: $2 input, $10 output per million tokens, with cache writes at $2.50.
Safety posture#
Anthropic reports far fewer misaligned behaviours and less willingness to cooperate with misuse than Haiku 4.5. On cyber, safeguards are stricter than Haiku 4.5’s but looser than Sonnet 5.5’s, and they still block penetration testing and similar attacker-oriented techniques. Biology safeguards match Sonnet 5, Sonnet 5.5 and Opus 5. Organisations that need broader access can apply to the Cyber Verification Program or the Life Sciences Verification Program.
If your security tooling runs on Haiku, expect refusals on offensive-flavoured prompts and plan to route those elsewhere.
Extras worth knowing#
- Available on the Anthropic API, AWS, Google Cloud and Azure from day one. Model id:
claude-haiku-5-5. - Max 5x subscribers get $100 in monthly API credits, Max 20x get $200, and Team plans up to $500 pooled, rolling out this week.
- The Python and TypeScript SDKs gain beta support for computer use and browser use. The 72.4% OSWorld 2.1 score suggests Haiku is viable for cheap UI automation, though that figure is on an offline subset.
What the customers say#
The quotes are the usual vendor-curated set, with the one real data point from HubSpot: 92.8% averaged over three runs on its internal suite, “the best score we’ve seen on this suite yet”. Box reports 11 points above Haiku 4.5 at about half the latency. Treat them as signals, not evidence, and run your own evals.
Our take#
This is how model tiers should work. Anthropic is not pretending the small model is the big model. It’s pricing Haiku so aggressively that you stop rationing calls, and telling you to keep Sonnet and Opus for the work that needs them. Compare that with tools that hide model routing behind a single “auto” toggle and a subscription meter. In an agentic workflow, being able to choose the model per subagent, from a terminal, with the price printed on the page, is the point.
The practical step this week: audit your subagent and compaction calls, switch the bounded ones to claude-haiku-5-5, and measure cost per completed task before and after.
