Claude Haiku 5.5 is live as of October 7, 2026 — Anthropic's fastest, cheapest, and most capable small model to date. It is designed for high-volume, cost-sensitive work: quick summaries, database queries, classification, document compaction, and subagent tasks that would have been too expensive with previous Haiku generations.

The pricing is the headline: $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens. For prompts over 100k, the price doubles ($0.50 / $2.50). That makes Haiku 5.5 roughly 75% cheaper than Haiku 4.5, which Anthropic attributes to both lower sticker prices and a more efficient tokenizer.

Every other site will repeat that press release. Here is what they skip: the catch is that Haiku 5.5 is still clearly below Sonnet 5.5 and Opus 5.5 on complex agentic coding and sustained reasoning; it shines when tasks are narrow, repetitive, and speed-sensitive. We also note the lack of independent VelsTech benchmarks — everything below comes from Anthropic's reported numbers, with their caveats preserved.

Source note: benchmarks and prices below come from Anthropic's launch announcement and Claude Platform docs (Oct 7, 2026). They are vendor-reported, not independent VelsTech tests. Where Anthropic notes measurement conditions (offline subsets, tool/no-tool splits, pre-release builds), we keep them.

Related: Our Claude Sonnet 5.5 explainer · Our Claude Opus 5.5 explainer · Best GPU for local LLMs · AI API cost calculator.

The short version: Claude Haiku 5.5 in 60 seconds

What Claude Haiku 5.5 actually is

Anthropic positions Haiku 5.5 as the small-model tier of the Claude 5.5 family (alongside Opus 5.5 and Sonnet 5.5). Unlike earlier Haiku releases, it launches with adjustable effort settings — Low, Medium, High, Max, and Xhigh — letting users trade speed and cost against accuracy in the same way larger Claude models already allow.

The model ID is claude-haiku-5-5. It is available immediately on the Claude Platform, Claude Code, Claude in Chrome, Microsoft 365 integrations, and through Amazon Web Services, Google Cloud, and Microsoft Azure. Zero data retention is available through the Claude Platform.

Anthropic notes that Haiku 5.5 is the first Haiku to come with an updated tokenizer (similar to Sonnet 5.5 and Opus 5.5's), which means it uses slightly more tokens per task than Haiku 4.5. That is part of why the 75% cost-reduction claim accounts for both price cuts and token-efficiency differences.

Specs at a glance

FeatureClaude Haiku 5.5Claude Haiku 4.5
ReleasedOct 7, 2026Oct 1, 2025 (approx.)
API model IDclaude-haiku-5-5claude-haiku-4-5-20251001 (legacy)
Input / outputText + images → textText + images → text
Context windowUp to 100k tokens (pricing tier split)200k tokens
Output speedFastest Claude model at standard speedFast, but slower per Anthropic
Effort settingsLow → Max + Xhigh (first Haiku with this)Not available
Default effortMedium in apps, High on PlatformFixed high-effort behaviour
Data retentionZero retention available (Platform)Zero retention available

Note the pricing tier split at 100,000 tokens: under that threshold is where the model is designed to operate for the majority of requests. Anthropic reports that ~90% of previous Haiku requests fell under 100k tokens.

Pricing: what it actually costs

Per 1M tokens (≤100k prompt)Haiku 5.5Haiku 4.5Sonnet 5.5 (ref.)GPT-6 Luna (ref.)
Input$0.10$1.00$2.00$0.10
Output$0.50$5.00$10.00$0.50
Cache reads$0.01$0.10$0.10 (new: $0.10, was $0.20)$0.01
Cache writes$0.125$1.25$2.50$0.125
Per 1M tokens (>100k prompt)Haiku 5.5Haiku 4.5
Input$0.50$1.00
Output$2.50$5.00
Cache reads$0.05$0.10
Cache writes$0.625$1.25

Benchmarks: Haiku 5.5 vs Haiku 4.5 vs Sonnet 5.5 vs GPT-6 Luna

All rows below are Anthropic-reported at stated effort levels, with production safeguards on. The offline-subset note on OSWorld 2.1 and the no-tools / with-tools split on Humanity's Last Exam are Anthropic's own caveats — we keep them.

Benchmark (what it tests)Haiku 5.5Haiku 4.5Sonnet 5.5 (ref.)GPT-6 Luna (ref.)
GDPval-AA v2.1 (44-occupation knowledge work, Elo)162073518441437
AA-Briefcase v1.1 (long-horizon knowledge work, Elo)157861418111336
OSWorld 2.1 — partial (computer use, offline subset)72.4%15.7%80.1%48.9%
Humanity's Last Exam — no tools45.9%10.2%64.5%—
Humanity's Last Exam — with tools57.4%18.7%——
Terminal-Bench 4.0 (agentic terminal coding)39.2%0.0%70.6%16.4%
FrontierCode 1.1 — Main46.4%—52.1% (Xhigh)42.4%
Chartography — no tools (visual chart reading)46.4%6.4%61.6%29.1%

The gap is clear: Haiku 5.5 roughly doubles Haiku 4.5 on knowledge work (GDPval-AA 1620 vs 735; AA-Briefcase 1578 vs 614) and makes the biggest jump on computer use (OSWorld 2.1: 72.4% vs 15.7%). But it remains well below Sonnet 5.5 and Opus 5.5 on agentic coding (Terminal-Bench 39.2% vs 70.6% / 66.4%) and chart reading (46.4% vs 61.6% / 64.4%).

Anthropic's efficiency curves show that Haiku 5.5 at Medium effort outperforms Haiku 4.5's best scores on OSWorld, GDPval-AA, and Humanity's Last Exam at roughly half the per-task cost — a pattern consistent with the Sonnet 5.5 efficiency story.

What changed from Haiku 4.5

Three real differences: (1) adjustable effort settings — the first Haiku to offer Low/Medium/High/Max/Xhigh tradeoffs; (2) a faster standard-speed inference profile, making it the fastest Claude model at standard settings (though not faster than Opus in Fast Mode); and (3) the updated tokenizer, which slightly increases token usage per task but pairs with a ~90% price reduction for requests under 100k.

Anthropic also reports that Haiku 5.5 pairs well with Opus 5.5 and Sonnet 5.5 as a subagent — for instance, while a larger model builds architecture or handles complex reasoning, a Haiku 5.5 subagent can pull specific line items, compile summaries, or process document batches in parallel.

The India angle

Which to pick

What this means if you run local AI

Nothing to download: Haiku 5.5 is cloud-only and closed-weight, like the rest of the Claude 5.5 family. Your local rig keeps its role — private, offline, predictable cost — but the efficiency frontier it competes against just moved again: faster standard-speed inference, lower per-query cost, and the first Haiku with adjustable effort are patterns worth copying into local agent designs.

If you run a hybrid setup (local model for private work, cloud API for speed), budget for Haiku 5.5 on the high-volume, low-complexity side of the split. See local vs cloud AI for the tradeoff math, and our AI API cost calculator to price a Haiku workload before committing.

Bottom line

Claude Haiku 5.5 is the budget/speed tier upgrade that makes the Claude 5.5 family complete. It roughly doubles Haiku 4.5 on knowledge work and computer use, introduces adjustable effort settings, and cuts cost by ~75% for prompts under 100k tokens. It does not replace Sonnet or Opus for complex agentic coding or sustained judgment — Anthropic is explicit that Opus 5.5 stays clearly stronger on open-ended work. But for everyday high-volume tasks — summaries, subagents, database queries, classification — the default answer just got faster and much cheaper.

Sources

FAQ

What is Claude Haiku 5.5?

Anthropic's fastest, cheapest small model released Oct 7, 2026 — model ID claude-haiku-5-5, $0.10/$0.50 per million tokens (≤100k prompts), 39.2% Terminal-Bench 4.0, 1620 GDPval-AA Elo, first Haiku with adjustable effort settings (Low → Xhigh), and the first Haiku with an updated tokenizer.

How much does Claude Haiku 5.5 cost?

$0.10/M input, $0.50/M output for prompts ≤100k tokens; $0.50/M input, $2.50/M output for longer prompts. Cache reads $0.01/M (≤100k) or $0.05/M (over 100k); writes $0.125/M (≤100k) or $0.625/M (over 100k). Roughly 75% cheaper than Haiku 4.5.

How does Haiku 5.5 compare to Sonnet 5.5 and Opus 5.5?

Haiku 5.5 trails both on complex agentic coding (39.2% Terminal-Bench vs 70.6% Sonnet / 66.4% Opus) and sustained reasoning. It is the budget/speed tier for high-volume, narrowly-scoped work — summaries, queries, subagents, classification. Use Opus for hard reasoning; keep Sonnet for everyday agentic coding.

Can I run Claude Haiku 5.5 locally?

No. It is cloud-only and closed-weight. Use the Claude API / Platform, AWS, Google Cloud, Microsoft Azure, or Claude.ai. For private/offline work, see our local AI guides.