
Claude Sonnet 5.5: Same $2/$10, but the 30% Saving Only Shows Up on Long Jobs
Anthropic shipped Claude Sonnet 5.5 on 28 September 2026 at exactly Sonnet 5's price: $2 per million input tokens, $10 per million output. A day later we pushed 19 graded calls through OpenRouter. On a long HTML file the bill fell 11.1%. On a tiny invoice sum it rose 12.8%.
Same price. Different bill.
What changed when Claude Sonnet 5.5 shipped on 28 September?
Swap the model id to claude-sonnet-5-5 and your per-token rate stays at $2 input and $10 output.
Anthropic's pricing page lists Sonnet 5.5 and Sonnet 5 on identical rows, cache and batch included, so every cost difference between them is a token count.
| Item | Sonnet 5.5 at launch |
|---|---|
| API model id | claude-sonnet-5-5 |
| Input / output, per 1M tokens | $2 / $10 |
| Cache read / cache write | $0.20 / $2.50 (5 min), $4 (1 hour) |
| Batch input / output | $1 / $5 |
| Context / max output | 1M / 128K tokens |
| Default effort | Medium in Claude Code and apps, High on the API |
| Availability | Zero data retention; AWS, Google Cloud, Microsoft Azure |
| GitHub Copilot | Pro, Pro+, Max, Business, Enterprise, at list price |
Sources: Anthropic's Sonnet 5.5 model page, launch post and GitHub's changelog. Background: what Sonnet 5 changed and per-token prices across providers.
How close does Sonnet 5.5 get to Opus 5.5 on Anthropic's table?
Sonnet 5.5 trails Opus 5.5 by under three points on four of six percentage rows, and its one win sits inside the error bars.
Cells re-typed from the hero image:
| Benchmark | Sonnet 5.5 | Sonnet 5 | Opus 5.5 | GPT-6 Sol |
|---|---|---|---|---|
| Terminal-Bench 4.0 | 70.6% | 10.3% | 66.4%¹ | n/a |
| FrontierCode 1.1 Main | 46.2% (Max)², 52.1% (Xhigh) | 42.4% | 54.4% | 49.3% |
| CursorBench 4.0 | 55.5% | 34.1% | 57.8% | n/a |
| GDPval-AA v2.1³ | 1844 | 1449 | 1846 | 1487⁴ |
| AA-Briefcase v1.1³ | 1811 | 1359 | 1822 | 1483⁴ |
| HLE (with tools) | 64.5% | 54.9% | 67.7% | n/a |
| OSWorld 2.1 (partial) | 80.1% | 57.0% | 81.8% | n/a |
| Chartography (no tools) | 61.6% | 15.6% | 64.4% | 53.6%⁴ |
¹ Opus at Xhigh effort. ² Max scored below Xhigh: it ran more code-review subagents, which Cognition found caused timeouts or out-of-scope edits in two cases. ³ Run by Artificial Analysis on a pre-release deployment with a structured-output bug, since fixed. ⁴ A GPT-6 Sol image bug may not be reflected. Anthropic's Terminal-Bench and CursorBench charts plot GPT-5.6 Sol.
FrontierCode (Xhigh) and CursorBench trail by 2.3, OSWorld by 1.7, Chartography by 2.8. HLE is the exception at 3.2. The Elo rows sit 2 and 11 apart.
That Terminal-Bench row feeds the r/ClaudeCode thread "Sonnet 5.5 beats Opus 5.5 at coding and it's half the price??". The system card puts standard errors at ±2.5 and ±2.6 points, so the 4.2-point gap is about 1.2 combined standard errors, and Sonnet ran at max while Opus ran at xhigh. Artificial Analysis measured Sonnet 5.5 at 64% on the same test. The widest Opus lead is off the launch page: SWE-Bench Pro, 81.3 against 89.9.
GPT-6 Sol is just the column Anthropic printed; for how GPT-6 Astra compares with Opus 5.5, see our three-way test.
Does Anthropic's 30% saving survive 19 API calls?
Only on the long job: 11.1% cheaper on an HTML task averaging 2,564 output tokens, 10-13% dearer on two tiny tasks.
On 29 September at 14:25 UTC we sent 19 graded calls through OpenRouter to Sonnet 5.5 and Sonnet 5: 3 tasks, 3 trials, 2 models, plus one high-effort extra. Total bill, $0.183434. All 18 matched calls passed.
# sonnet55_bench.py, 29 Sep 2026, POST https://openrouter.ai/api/v1/chat/completions
body = {
"model": model, # anthropic/claude-sonnet-5.5 or anthropic/claude-sonnet-5
"messages": messages,
"max_tokens": max_tokens, # 12000 on the HTML task
"reasoning": {"effort": effort}, # low: invoice, merge_intervals; medium: HTML
}No LLM graded anything. The invoice (800k cache-read, 200k input, 50k output tokens) had to return exactly 1.06, merge_intervals ran 6 executed asserts, and the HTML needed doctype, canvas, requestAnimationFrame and no external script. Costs are usage.cost summed over 3 trials.
| Task (effort) | Sonnet 5.5 | Sonnet 5 | Delta | Mean output tokens | Mean latency |
|---|---|---|---|---|---|
| Invoice (low) | $0.002754 | $0.002442 | +12.8% | 71 / 61 | 2.63 s / 4.49 s |
| merge_intervals (low) | $0.006054 | $0.005502 | +10.0% | 172 / 154 | 3.06 s / 3.68 s |
| 150-bird HTML (medium) | $0.077508 | $0.087156 | -11.1% | 2,564 / 2,886 | 15.54 s / 21.61 s |
| Matched total | $0.086316 | $0.0951 | -9.2% | n/a | n/a |
Anthropic promises "up to 30% less per task than its predecessor". Matthias Bastian at The Decoder added the right caveat: "Independent testing still needs to confirm these claims."
In our 19-call meter at low and medium effort, one long task, the saving was -11.1% on the HTML and flipped to +12.8% and +10.0% on the tiny tasks. Same list price means every saving is a token count, and on our two smallest tasks Sonnet 5.5 wrote more tokens, not fewer.
Speed held up better. HTML latency fell 28.1%, and throughput rose from 133.5 to 165.0 tokens per second (+23.6%): right direction, a bit under 30%. Sonnet 5 also logged 54-63 reasoning tokens per invoice call; Sonnet 5.5 logged 0 on all 10 calls.
That HTML task is a 150-bird version of Anthropic's 400-starling launch-page demo prompt. The frames below are Anthropic's recording, not our run.


What Anthropic's launch testers reported
"Without changing any of our prompts, Claude Sonnet 5.5 did better than Sonnet 5 on almost all of our offline Slackbot evals, in fewer steps and with about 14% fewer output tokens." Curtis Allen, Principal Engineer, Slack
"The new model managed tens of thousands of lines of code for gameplay system architecture, kept responses snappy, handled multi-hour tasks, and delivered with less prescriptive prompting." Daniel Vogel, Chief Operating Officer, Epic Games
Both are vendor-collected, and both describe long, multi-step work. Our tiny tasks went the other way.
Not measured: agentic work, the benchmarks above, the cyber fallback. 3 trials is a spot meter; to run your own, read usage on every call.
Why does effort level move the bill more than the model?
Artificial Analysis measured $0.59 per task at medium and $7.60 at max, a 13x jump for 15 index points.
At max it counted "~193k Output Tokens per Intelligence Index Task. This is the highest token use we have measured".
| Effort | AA cost per task | AA Intelligence Index | AA rank, 29 Sep |
|---|---|---|---|
| Low | $0.41 | 36 | 63 of 216 |
| Medium | $0.59 | 41 | 44 of 216 |
| High | $1.08 | 47 | 23 of 216 |
| Max | $7.60 | 56 | 3 of 216 |
Rows come from AA's per-effort model pages, retrieved 29 September. AA's launch article said Sonnet 5.5 took second place on the index; the model page showed 56, rank 3 of 216, that day. Max also cost "~50% higher than Sonnet 5's Cost per Task".
Our own high-effort call ignored the setting: merge_intervals at effort=high cost $0.002018 for 172 output tokens, identical to low. Effort bites only on work big enough to think about, which is where cutting LLM API spend starts.
What returns a 400 after you swap the model id?
Any request with thinking disabled fails; change it to between_tools and keep effort at high or below.
Anthropic's migration guide spells out the error:
# Before (Sonnet 5 habit): 400 invalid_request_error on claude-sonnet-5-5
thinking = {"type": "disabled"}
# '"thinking.type.disabled" is not supported for this model. Use
# "thinking.type.between_tools" for the lowest thinking setting, ...'
# After: accepted at low / medium / high effort, 400 at xhigh / max
thinking = {"type": "between_tools"}
output_config = {"effort": "medium"} # fixed for the conversation under between_toolstool_choiceofanyortoolreturns 400. Useautoplusstrict: true(20 strict tools max), orautoalone on Amazon Bedrock.- Thinking blocks from Opus 5, Opus 5.5, Fable or Mythos are dropped silently; accounts created on or after 31 Aug 2026 get a 400 for replaying a block after editing history. The mechanics match Fable 5.1's thinking-block binding.
- Computer use needs
computer_toolset_20260801on the Claude API and Google Cloud. - Refusals name a
stop_detailscategory, includingcyber,frontier_llm("could assist the development of competing AI models") andreasoning_extraction. - Server-side fallback (
fallbacks: "default", beta, Claude API only) retriescyberandfrontier_llmdeclines on Sonnet 5; the system card says 1.2% of Terminal-Bench requests fell back. - The minimum cacheable prompt drops to 512 tokens from 1,024 (prompt caching minimums), and
/claude-api migratein Claude Code applies these changes.
Security teams hit item 4 first. One HN commenter, johnmlussier, said that despite being in the Cyber Verification Program they "can't use Opus 5.5 or Sonnet 5.5 for any authorized bounty work". Anthropic says the expanded program comes "Soon".
The same checklist says "Re-run your effort sweep, and re-baseline cost". Do that before trusting the 30%.
Sonnet 5.5 or Opus 5.5: which one gets the default slot?
Default to Sonnet 5.5 for scoped coding and high-volume calls; keep Opus 5.5 for open-ended judgment and any max-effort work.
Sonnet lists at half of Opus 5.5's $4/$20, yet Artificial Analysis measured Sonnet at max costing $7.60 per task against Opus 5.5 at $5.98, about 27% more.
| Signal | Sonnet 5.5 | Opus 5.5 | Leans |
|---|---|---|---|
| Input / output per 1M | $2 / $10 | $4 / $20 | Sonnet, half the price |
| Cache read / 5-min write | $0.20 / $2.50 | $0.20 / $5 | Tie on reads |
| Cache-heavy mix (below) | $1.06 | $1.96 | Sonnet, 1.85x not 2x |
| AA cost per task, max effort | $7.60 | $5.98 | Opus, 27% cheaper |
| SWE-Bench Pro | 81.3% | 89.9% | Opus by 8.6 points |
Mix: 800k cache-read + 200k input + 50k output tokens
Sonnet 5.5: 0.8 x $0.20 + 0.2 x $2 + 0.05 x $10 = $0.16 + $0.40 + $0.50 = $1.06
Opus 5.5: 0.8 x $0.20 + 0.2 x $4 + 0.05 x $20 = $0.16 + $0.80 + $1.00 = $1.96
Ratio: 1.96 / 1.06 = 1.85xAleksei Aleinikov's "approximately 1.65×" depends on a mix he doesn't state. Plug in yours: the heavier your cache share, the smaller Sonnet's edge.
Anthropic itself says "Opus 5.5 remains clearly stronger at complex, open-ended work requiring sustained judgment." One HN commenter, datadrivenangel, went further: "Opus 5.5 on Low seems smarter, cheaper, and faster than sonnet on medium". Untested by us. Opus pricing and limits live in our Opus 5.5 breakdown.
Our default from 29 September: claude-sonnet-5-5 at medium effort. Max-effort jobs go to Opus 5.5, because at max, Sonnet stops being the cheap model.
Frequently Asked Questions
When was Claude Sonnet 5.5 released?
Anthropic released Claude Sonnet 5.5 on 28 September 2026 as claude-sonnet-5-5, the same day on Amazon Web Services, Google Cloud, Microsoft Azure and GitHub Copilot. Anthropic's docs say it won't be retired before 28 September 2027.
Is Claude Sonnet 5.5 cheaper than Sonnet 5?
Not per token: both list at $2 input and $10 output. In our meter a long HTML task cost 11.1% less on Sonnet 5.5 and a tiny invoice task 12.8% more; at max effort, Artificial Analysis measured about 50% more per task.
Can I use Sonnet 5.5 in Claude Code and GitHub Copilot?
Yes. Claude Code defaults it to Medium effort; the API defaults to High. In GitHub Copilot, Pro through Enterprise, it is "billed at provider list pricing under usage-based billing". To pin it, see how to switch models in Claude Code.
What is Sonnet 5.5's context window and knowledge cutoff?
It takes 1M tokens of context and returns up to 128K output tokens, with a reliable knowledge cutoff of June 2026 per the system card.
Is there a Haiku 5.5?
Not as of 29 September 2026. Anthropic said on 28 September that Claude Haiku 5.5 "will join the Claude 5.5 family in the coming weeks."