Claude Sonnet 5.5 is live as of September 28, 2026 — the second model in Anthropic's Claude 5.5 family after Claude Opus 5.5 (Sept 22, 2026). Anthropic calls it a clear upgrade over Claude Sonnet 5 (June 30, 2026): it runs 30%+ faster and costs up to 30% less per task for most work, while jumping from 10.3% to 70.6% on Terminal-Bench 4.0.
The positioning is deliberate: where Opus 5.5 is built for complex work requiring careful judgment, Sonnet 5.5 is the fast, low-cost complement for well-scoped everyday tasks — fixing bugs, iterating on code, and producing polished documents, slides and spreadsheets. It is also, Anthropic says, the first Sonnet with a sharp eye for design.
Source note: benchmarks and prices below come from Anthropic's launch announcement and Claude Platform docs. They are vendor-reported, not an independent VelsTech benchmark. Where Anthropic discloses caveats (effort settings, fallback models, measurement bugs), we keep them.
The short version: Claude Sonnet 5.5 in 60 seconds
- What it is: Sonnet-tier model for everyday agentic coding, bug-fixing, documents, slides and spreadsheets. API model ID
claude-sonnet-5-5, available on all platforms with zero data retention. - Release date: September 28, 2026, on the Claude API/Platform, Claude Code, Amazon Web Services, Google Cloud and Microsoft Azure. Claude Haiku 5.5 follows in the coming weeks.
- Price: $2 per million input tokens / $10 per million output tokens — identical sticker price to Sonnet 5, but typically up to 30% cheaper per finished task because it uses far fewer tokens. Cache reads $0.20/M.
- Benchmarks: 70.6% Terminal-Bench 4.0 (Sonnet 5: 10.3%, Opus 5.5: 66.4%), 1844 GDPval-AA Elo (two points below Opus 5.5, ~400 above Sonnet 5), 80.1% OSWorld 2.1, first Sonnet to beat Pokémon Red from screenshots alone.
- Efficiency: at Low/Medium effort it beats Sonnet 5's best scores for roughly a tenth of the per-task cost on Terminal-Bench, CursorBench and AA-Briefcase.
- Safety: first Sonnet to launch with Opus-class cyber safeguards and anti-distillation classifiers; biology safeguards unchanged from Sonnet 5; routine development work unaffected.
- Who should care: anyone running Sonnet 5 in Claude Code, Copilot-style tools, support bots or document pipelines — same price, faster, better. If you use Opus for hard reasoning, keep it; Sonnet 5.5 complements rather than replaces it.
Claude Sonnet 5.5 specs at a glance
| Feature | Claude Sonnet 5.5 | Claude Sonnet 5 |
|---|---|---|
| Released | Sept 28, 2026 | June 30, 2026 |
| API model ID | claude-sonnet-5-5 | claude-sonnet-5 |
| Input / output | Text + images → text | Text + images → text |
| Output speed | 30%+ faster (fastest Sonnet to date) | Baseline |
| Thinking | Effort levels low → max; thinking-off users must switch to between_tools (see migration guide) | Thinking on/off |
| Default effort | Medium in apps, High on Platform | Medium in apps, High on Platform |
| Data retention | Zero retention available | Zero retention available |
Platform model IDs: Claude API claude-sonnet-5-5, on AWS, Google Cloud
and Microsoft Azure through their usual Claude listings. If you run Sonnet with
thinking off, read the
Sonnet 5.5 migration guide
before switching — the old thinking-off flag is replaced by between_tools.
Sonnet 5.5 pricing: what it actually costs
| Per 1M tokens | Sonnet 5.5 | Sonnet 5 | Opus 5.5 |
|---|---|---|---|
| Input | $2 | $2 | $4 |
| Output | $10 | $10 | $20 |
| Cache reads | $0.20 | $0.20 | $0.20 |
| Cache writes | $2.50 | $2.50 | $5 |
Same sticker, lower bill: Sonnet 5.5 needs far fewer tokens per task, so Anthropic reports up to 30% lower cost per task in its testing. Balyasny Asset Management's numbers make it concrete — ~121K tokens per finance answer vs 497K on Sonnet 5. If you estimate agent spend, use our AI API cost calculator with your real cached-vs-fresh input split.
Benchmarks: Sonnet 5.5 vs Sonnet 5 vs Opus 5.5 vs GPT-6 Sol
All rows below are Anthropic-reported, at stated effort levels, with production safeguards on. Anthropic cautions that benchmark margins are a noisy guide to real-world gaps — and that Opus 5.5 remains clearly stronger at complex, open-ended work needing sustained judgment.
| Benchmark (what it tests) | Sonnet 5.5 | Sonnet 5 | Opus 5.5 | GPT-6 Sol |
|---|---|---|---|---|
| Terminal-Bench 4.0 (agentic terminal coding) | 70.6% | 10.3% | 66.4% | —* |
| FrontierCode v1.1 Main (would-merge code changes) | 46.2% (Max) / 52.1% (Xhigh) | 42.4% | 54.4% | 49.3% |
| CursorBench 4.0 (multi-file Cursor sessions) | 55.5% | 34.1% | 57.8% | —* |
| GDPval-AA v2.1 (44-occupation knowledge work, Elo) | 1844 | 1449 | 1846 | 1487 |
| AA-Briefcase v1.1 (long-horizon knowledge work, Elo) | 1811 | 1359 | 1822 | 1483 |
| Humanity's Last Exam with tools (multidisciplinary reasoning) | 64.5% | 54.9% | 67.7% | — |
| OSWorld 2.1 partial (computer use) | 80.1% | 57.0% | 81.8% | — |
| Chartography, no tools (visual chart reading) | 61.6% | 15.6% | 64.4% | 53.6% |
* OpenAI did not report GPT-6 Sol on Terminal-Bench or CursorBench, so Anthropic reports GPT-5.6 Sol there instead. FrontierCode penalizes out-of-scope edits, which is why Sonnet 5.5 scores lower at Max than Xhigh effort. GDPval/AA-Briefcase ran on a pre-release build with a structured-output bug Anthropic says is fixed and likely understated Sonnet 5.5.
The shape of the table matters more than any cell: Sonnet 5.5 sits within a couple of points of Opus 5.5 on knowledge work, computer use and chart reading, while Sonnet 5 trails by 20–45 points on the same rows. The one-sided jump — Terminal-Bench 10.3% → 70.6% — is the largest single-generation Sonnet gain Anthropic has reported.
Efficiency is the real upgrade
Anthropic's accuracy-vs-cost curves are the point of this release: Sonnet 5.5 at Low or Medium effort beats Sonnet 5's best score for about a tenth of the per-task cost on Terminal-Bench, CursorBench and AA-Briefcase. At High effort on FrontierCode it matches GPT-6 Sol's best for about a fifth of the cost.
- Epic Games: held a higher-tier quality bar on a system-design audit and data-flow review across tens of thousands of lines of gameplay architecture, with less prescriptive prompting.
- Base44 (118 real app builds): Sonnet 5.5 builds scored level with Opus 5 in 3.6 iterations on average vs 7.7, with the fewest failed tool calls of any model compared.
- Balyasny (2,441 finance tasks): ahead of Sonnet 5 at ~121K tokens per answer vs 497K — the best quality-to-cost tradeoff of seven models tested.
- Zendesk (support tickets): fewer wrong decisions, tickets processed 20% faster than current production models.
- Unity (multi-step Editor + coding benchmark): 90% task completion with runtime-checked results, beating similar models.
The mechanism, per early testers: Sonnet 5.5 batches tool calls instead of dribbling them out (CodeRabbit notes the compulsive web-search habit and token hunger of Sonnet 5 are gone), understands codebases faster, and rarely stalls mid-build asking the user questions.
Coding, documents and design
Coding is where the jump shows most: +10 points over Sonnet 5 at the same High-effort FrontierCode setting for about one-fifteenth the per-task cost, and within ~2 points of Opus 5.5 on CursorBench. SpaceXAI calls 55.5% on CursorBench "frontier-level … second only to Opus 5.5", and Unity's Creator game team would let Opus set the architecture and Sonnet implement it.
The surprise beat is design taste: testers report polished UIs with minimal editing, slide decks from templates that need little cleanup (Anthropic's internal test turned earnings materials plus a template into a send-as-is 10-slide operating review), and clearer everyday writing in the Opus 5.5 style. If your workload is "fix this bug, draft this doc, build this slide" rather than multi-hour autonomy, this is the model tier that just got much better value.
Safety and safeguards: first Sonnet with Opus-class cyber defaults
- Alignment: improves on or matches Sonnet 5 across most of Anthropic's ~1,850-scenario behavioral audit; closest to Opus 5.5 at resisting sandbox escape, least likely of any model to probe container limits.
- Cyber: capabilities comparable to Opus 5, so it launches with Opus-style cyber safeguards and fallbacks — the first Sonnet to do so. Routine bug-fixing is unaffected; higher-risk tasks visibly fall back to Sonnet 5. Defenders can apply to the expanded Cyber Verification Program.
- Biology: same safeguards as Sonnet 5; most research, education and clinical work unaffected. Full-spectrum biology access via the Life Sciences Verification Program.
- Distillation: first Sonnet with reasoning-extraction classifiers plus expanded preserved thinking — moving sessions between accounts may behave differently (see platform docs).
Availability: where to use Claude Sonnet 5.5 today
- Claude API / Platform:
claude-sonnet-5-5— Claude Code, Cowork, Claude in Chrome, Microsoft 365 integration, zero data retention. - Clouds: Amazon Web Services, Google Cloud, Microsoft Azure.
- Chat: Claude.ai on web, iOS and Android.
- Coming weeks: Claude Haiku 5.5 for high-volume, cost-sensitive work.
Sonnet 5.5 vs Sonnet 5: should you switch?
For almost every Sonnet 5 workload, yes — same sticker price, 30%+ faster, better
scores at every effort level, and lower per-task cost from fewer tokens. Budget half
a day for the between_tools thinking-off migration and a side-by-side on
your own prompts, and watch for safeguard fallbacks if your agents touch security
tooling. If you use Opus for hard reasoning, don't downgrade: Anthropic is explicit
that Opus 5.5 stays clearly stronger on complex open-ended work.
What it means for local-AI users
Nothing to download: Sonnet 5.5 is cloud-only, like Opus 5.5 and Fable 5.1. Your local rig keeps its role — private, offline, predictable cost — while the efficiency frontier it chases just moved: batched tool calls, fewer tokens per task, and effort levels that trade thinking for speed are patterns worth copying into local agents today. See best GPU for local LLMs and local vs cloud AI for the tradeoff math, and our AI API cost calculator to price a Sonnet 5.5 workload before you commit.
Bottom line
Claude Sonnet 5.5 is the mid-tier upgrade that actually changes buying decisions: near-Opus knowledge work and computer use, a 7x Terminal-Bench jump, 30%+ more speed and up to 30% lower per-task cost — at exactly Sonnet 5's price. Vendor tables are not independent tests, and Opus keeps the crown for sustained judgment calls. But for everyday agentic coding, support automation and document pipelines, the default answer is now Sonnet 5.5. Haiku 5.5 will decide how far down the stack the 5.5 story goes.
Sources
- Anthropic: Introducing Claude Sonnet 5.5 (Sept 28, 2026)
- Claude Sonnet 5.5 System Card
- Claude Platform Docs: Sonnet 5.5 migration guide
- Our Claude Opus 5.5 explainer
- Our Claude Fable 5.1 explainer
FAQ
What is Claude Sonnet 5.5?
Anthropic's Sonnet-tier model released Sept 28, 2026 — the second Claude 5.5 model after Opus 5.5, built for everyday agentic coding, bug-fixing, documents and slides at 30%+ more speed and up to 30% lower per-task cost than Sonnet 5.
When was Claude Sonnet 5.5 released?
September 28, 2026, six days after Claude Opus 5.5 (Sept 22, 2026). Claude Haiku 5.5 follows in the coming weeks.
How much does Claude Sonnet 5.5 cost?
$2 per million input tokens and $10 per million output tokens — the same sticker price as Sonnet 5 — with $0.20 cache reads and $2.50 cache writes. Because it uses far fewer tokens per task, typical work costs up to 30% less than Sonnet 5.
What is the Claude Sonnet 5.5 model ID?
API ID claude-sonnet-5-5 on the Claude Platform, also live on AWS, Google Cloud and Microsoft Azure with zero data retention. Thinking-off users must switch to the new between_tools setting.
Is Claude Sonnet 5.5 better than Sonnet 5 or Opus 5.5?
On Anthropic's reported table it beats Sonnet 5 on every row — most dramatically Terminal-Bench 4.0 (70.6% vs 10.3%) — and sits within ~2 points of Opus 5.5 on knowledge work, computer use and chart reading, while Opus stays clearly stronger on complex open-ended work.
Can I run Claude Sonnet 5.5 locally?
No. It is closed and cloud-only. Use the API, clouds or Claude.ai — or run open-weight models locally for privacy and offline work; see our local-AI guides.