Claude Haiku 5.5 is live as of October 7, 2026 — Anthropic's fastest, cheapest, and most capable small model to date. It is designed for high-volume, cost-sensitive work: quick summaries, database queries, classification, document compaction, and subagent tasks that would have been too expensive with previous Haiku generations.
The pricing is the headline: $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens. For prompts over 100k, the price doubles ($0.50 / $2.50). That makes Haiku 5.5 roughly 75% cheaper than Haiku 4.5, which Anthropic attributes to both lower sticker prices and a more efficient tokenizer.
Every other site will repeat that press release. Here is what they skip: the catch is that Haiku 5.5 is still clearly below Sonnet 5.5 and Opus 5.5 on complex agentic coding and sustained reasoning; it shines when tasks are narrow, repetitive, and speed-sensitive. We also note the lack of independent VelsTech benchmarks — everything below comes from Anthropic's reported numbers, with their caveats preserved.
Source note: benchmarks and prices below come from Anthropic's launch announcement and Claude Platform docs (Oct 7, 2026). They are vendor-reported, not independent VelsTech tests. Where Anthropic notes measurement conditions (offline subsets, tool/no-tool splits, pre-release builds), we keep them.
Related: Our Claude Sonnet 5.5 explainer · Our Claude Opus 5.5 explainer · Best GPU for local LLMs · AI API cost calculator.
The short version: Claude Haiku 5.5 in 60 seconds
- What it is: Anthropic's fastest, cheapest Haiku-class model — API ID
claude-haiku-5-5, available on Claude Platform, AWS, Google Cloud, and Microsoft Azure. - Release date: October 7, 2026.
- Price: $0.10/M input, $0.50/M output (≤100k tokens); $0.50/M input, $2.50/M output (over 100k); cache reads $0.01/M (≤100k) / $0.05/M (over 100k); writes $0.125/M (≤100k) / $0.625/M (over 100k).
- Benchmarks: 39.2% Terminal-Bench 4.0, 1620 GDPval-AA Elo, 46.4% Chartography (no tools), 46.4% FrontierCode 1.1, 72.4% OSWorld 2.1 (offline subset), 45.9% Humanity's Last Exam (no tools) / 57.4% (with tools).
- Headline claim: ~75% cheaper than Haiku 4.5; fastest Claude model to date at standard speed; first Haiku with adjustable effort settings.
- Catch: Well below Sonnet 5.5 / Opus 5.5 on agentic coding and complex reasoning; built for high-volume, cost-sensitive, narrowly-scoped work.
- Who should care: Anyone running high-volume summaries, database queries, subagents, or customer-support pipelines — same quality direction as Sonnet 5.5 at a fraction of the cost. Keep Opus for sustained judgment calls; keep Haiku for speed and volume.
What Claude Haiku 5.5 actually is
Anthropic positions Haiku 5.5 as the small-model tier of the Claude 5.5 family (alongside Opus 5.5 and Sonnet 5.5). Unlike earlier Haiku releases, it launches with adjustable effort settings — Low, Medium, High, Max, and Xhigh — letting users trade speed and cost against accuracy in the same way larger Claude models already allow.
The model ID is claude-haiku-5-5. It is available immediately on the Claude Platform, Claude Code, Claude in Chrome, Microsoft 365 integrations, and through Amazon Web Services, Google Cloud, and Microsoft Azure. Zero data retention is available through the Claude Platform.
Anthropic notes that Haiku 5.5 is the first Haiku to come with an updated tokenizer (similar to Sonnet 5.5 and Opus 5.5's), which means it uses slightly more tokens per task than Haiku 4.5. That is part of why the 75% cost-reduction claim accounts for both price cuts and token-efficiency differences.
Specs at a glance
| Feature | Claude Haiku 5.5 | Claude Haiku 4.5 |
|---|---|---|
| Released | Oct 7, 2026 | Oct 1, 2025 (approx.) |
| API model ID | claude-haiku-5-5 | claude-haiku-4-5-20251001 (legacy) |
| Input / output | Text + images → text | Text + images → text |
| Context window | Up to 100k tokens (pricing tier split) | 200k tokens |
| Output speed | Fastest Claude model at standard speed | Fast, but slower per Anthropic |
| Effort settings | Low → Max + Xhigh (first Haiku with this) | Not available |
| Default effort | Medium in apps, High on Platform | Fixed high-effort behaviour |
| Data retention | Zero retention available (Platform) | Zero retention available |
Note the pricing tier split at 100,000 tokens: under that threshold is where the model is designed to operate for the majority of requests. Anthropic reports that ~90% of previous Haiku requests fell under 100k tokens.
Pricing: what it actually costs
| Per 1M tokens (≤100k prompt) | Haiku 5.5 | Haiku 4.5 | Sonnet 5.5 (ref.) | GPT-6 Luna (ref.) |
|---|---|---|---|---|
| Input | $0.10 | $1.00 | $2.00 | $0.10 |
| Output | $0.50 | $5.00 | $10.00 | $0.50 |
| Cache reads | $0.01 | $0.10 | $0.10 (new: $0.10, was $0.20) | $0.01 |
| Cache writes | $0.125 | $1.25 | $2.50 | $0.125 |
| Per 1M tokens (>100k prompt) | Haiku 5.5 | Haiku 4.5 |
|---|---|---|
| Input | $0.50 | $1.00 |
| Output | $2.50 | $5.00 |
| Cache reads | $0.05 | $0.10 |
| Cache writes | $0.625 | $1.25 |
- The cliff: The 100k-token threshold doubles the price. If you routinely send long context windows, budget for the over-100k tier — but Anthropic says 90% of previous Haiku requests stayed under it.
- Concrete example: A support-bot answering 10,000 short queries/day, each using ~2K input + 500 output tokens: ~20M input + 5M output tokens/day = $2 input + $2.50 output = $4.50/day, vs ~$45/day on Haiku 4.5 for the same workload. Use our AI API cost calculator with your real cached-vs-fresh split.
Benchmarks: Haiku 5.5 vs Haiku 4.5 vs Sonnet 5.5 vs GPT-6 Luna
All rows below are Anthropic-reported at stated effort levels, with production safeguards on. The offline-subset note on OSWorld 2.1 and the no-tools / with-tools split on Humanity's Last Exam are Anthropic's own caveats — we keep them.
| Benchmark (what it tests) | Haiku 5.5 | Haiku 4.5 | Sonnet 5.5 (ref.) | GPT-6 Luna (ref.) |
|---|---|---|---|---|
| GDPval-AA v2.1 (44-occupation knowledge work, Elo) | 1620 | 735 | 1844 | 1437 |
| AA-Briefcase v1.1 (long-horizon knowledge work, Elo) | 1578 | 614 | 1811 | 1336 |
| OSWorld 2.1 — partial (computer use, offline subset) | 72.4% | 15.7% | 80.1% | 48.9% |
| Humanity's Last Exam — no tools | 45.9% | 10.2% | 64.5% | — |
| Humanity's Last Exam — with tools | 57.4% | 18.7% | — | — |
| Terminal-Bench 4.0 (agentic terminal coding) | 39.2% | 0.0% | 70.6% | 16.4% |
| FrontierCode 1.1 — Main | 46.4% | — | 52.1% (Xhigh) | 42.4% |
| Chartography — no tools (visual chart reading) | 46.4% | 6.4% | 61.6% | 29.1% |
The gap is clear: Haiku 5.5 roughly doubles Haiku 4.5 on knowledge work (GDPval-AA 1620 vs 735; AA-Briefcase 1578 vs 614) and makes the biggest jump on computer use (OSWorld 2.1: 72.4% vs 15.7%). But it remains well below Sonnet 5.5 and Opus 5.5 on agentic coding (Terminal-Bench 39.2% vs 70.6% / 66.4%) and chart reading (46.4% vs 61.6% / 64.4%).
Anthropic's efficiency curves show that Haiku 5.5 at Medium effort outperforms Haiku 4.5's best scores on OSWorld, GDPval-AA, and Humanity's Last Exam at roughly half the per-task cost — a pattern consistent with the Sonnet 5.5 efficiency story.
What changed from Haiku 4.5
Three real differences: (1) adjustable effort settings — the first Haiku to offer Low/Medium/High/Max/Xhigh tradeoffs; (2) a faster standard-speed inference profile, making it the fastest Claude model at standard settings (though not faster than Opus in Fast Mode); and (3) the updated tokenizer, which slightly increases token usage per task but pairs with a ~90% price reduction for requests under 100k.
Anthropic also reports that Haiku 5.5 pairs well with Opus 5.5 and Sonnet 5.5 as a subagent — for instance, while a larger model builds architecture or handles complex reasoning, a Haiku 5.5 subagent can pull specific line items, compile summaries, or process document batches in parallel.
The India angle
- Price ≈ ₹8.30 / ₹41.50 per million tokens (≤100k) at current USD/INR rates (~₹83/$), before GST. Realistically ~₹9.50 / ₹47.50 after 18% GST.
- Cloud-only: no local download option. Use Claude Platform, AWS, Google Cloud, or Azure from Indian accounts.
- For cost-sensitive pipelines (customer support, document processing, database query generation), the ~75% price cut over Haiku 4.5 makes Haiku 5.5 competitive with open-source alternatives in pure cost-per-query terms for high-volume, low-complexity tasks.
- Compare with local vs cloud AI for when a private rig remains the better deal.
Which to pick
- Pick Haiku 5.5 if: your workload is high-volume, narrowly-scoped, speed-sensitive — summaries, classification, subagent pulls, database queries, document compaction, customer-support triage. You run thousands of calls per hour, not sustained multi-hour reasoning.
- Skip Haiku 5.5 if: your work requires sustained agentic coding (Terminal-Bench-level), complex open-ended reasoning, or long-context sustained judgment. Use Opus 5.5 or Sonnet 5.5 instead.
- Consider instead: open-weight local models (Qwen, Llama 3 family, Mistral) for fully private/offline work — see our local LLM GPU guide and cost calculator to compare real-world spend.
What this means if you run local AI
Nothing to download: Haiku 5.5 is cloud-only and closed-weight, like the rest of the Claude 5.5 family. Your local rig keeps its role — private, offline, predictable cost — but the efficiency frontier it competes against just moved again: faster standard-speed inference, lower per-query cost, and the first Haiku with adjustable effort are patterns worth copying into local agent designs.
If you run a hybrid setup (local model for private work, cloud API for speed), budget for Haiku 5.5 on the high-volume, low-complexity side of the split. See local vs cloud AI for the tradeoff math, and our AI API cost calculator to price a Haiku workload before committing.
Bottom line
Claude Haiku 5.5 is the budget/speed tier upgrade that makes the Claude 5.5 family complete. It roughly doubles Haiku 4.5 on knowledge work and computer use, introduces adjustable effort settings, and cuts cost by ~75% for prompts under 100k tokens. It does not replace Sonnet or Opus for complex agentic coding or sustained judgment — Anthropic is explicit that Opus 5.5 stays clearly stronger on open-ended work. But for everyday high-volume tasks — summaries, subagents, database queries, classification — the default answer just got faster and much cheaper.
Sources
- Anthropic: Introducing Claude Haiku 5.5 (Oct 7, 2026)
- Claude Haiku 5.5 System Card
- Claude Platform Docs: Haiku 5.5 migration guide
- Our Claude Opus 5.5 explainer
- Our Claude Sonnet 5.5 explainer
FAQ
What is Claude Haiku 5.5?
Anthropic's fastest, cheapest small model released Oct 7, 2026 — model ID claude-haiku-5-5, $0.10/$0.50 per million tokens (≤100k prompts), 39.2% Terminal-Bench 4.0, 1620 GDPval-AA Elo, first Haiku with adjustable effort settings (Low → Xhigh), and the first Haiku with an updated tokenizer.
How much does Claude Haiku 5.5 cost?
$0.10/M input, $0.50/M output for prompts ≤100k tokens; $0.50/M input, $2.50/M output for longer prompts. Cache reads $0.01/M (≤100k) or $0.05/M (over 100k); writes $0.125/M (≤100k) or $0.625/M (over 100k). Roughly 75% cheaper than Haiku 4.5.
How does Haiku 5.5 compare to Sonnet 5.5 and Opus 5.5?
Haiku 5.5 trails both on complex agentic coding (39.2% Terminal-Bench vs 70.6% Sonnet / 66.4% Opus) and sustained reasoning. It is the budget/speed tier for high-volume, narrowly-scoped work — summaries, queries, subagents, classification. Use Opus for hard reasoning; keep Sonnet for everyday agentic coding.
Can I run Claude Haiku 5.5 locally?
No. It is cloud-only and closed-weight. Use the Claude API / Platform, AWS, Google Cloud, Microsoft Azure, or Claude.ai. For private/offline work, see our local AI guides.