
Claude Opus 5.5: $4/$20 Pricing, the Limits, and What We Metered
Claude Opus 5.5 went live on 22 Sep 2026 at $4 per million input tokens and $20 per million output tokens. Cache reads are $0.20. That is 20% under Claude Opus 5 on input and output, and 60% under it on cache reads ($0.50). Platform id: claude-opus-5-5. OpenRouter already listed anthropic/claude-opus-5.5 at 18:57 UTC, same three rates, 1,000,000-token context.
Nine graded calls went out between 18:57 and 18:58 UTC. Eight matched Opus 5 at effort=low. All eight passed. Opus 5.5 billed $0.010564. Opus 5 billed $0.01269.
A ninth call, Opus 5.5 at effort=high, also passed and billed $0.003936. Experiment total: $0.02719. The cover is Anthropic's chart from the launch page, saved as-is.
Claude Opus 5.5 API price versus Opus 5
Cache-heavy traffic is where the new card moves. Uncached calls only fell about 20% on our meter.
Anthropic's launch table, reproduced below and in the price screenshot, is the standard card. Fable 5.1 stays at $10/$50. Its cache read is $0.25, so the Opus 5.5 cache line undercuts Fable by $0.05 per million, not by half. The 2.5x gap is input and output: $10/$4 and $50/$20.
| Line, per 1M tokens | Opus 5.5 | Opus 5 | Fable 5.1 |
|---|---|---|---|
| Input | $4 | $5 | $10 |
| Output | $20 | $25 | $50 |
| Cache read | $0.20 | $0.50 | $0.25 |
| Cache write, 5 min | $5 | $6.25 | $12.50 |
| Cache write, 1 hour | $8 | $10 | $20 |
| OpenRouter batch input / output | $2 / $10 | $2.50 / $12.50 | $5 / $25 |
| Fast mode input / output | $8 / $40 | not on this card | not on this card |

The 5-minute cache write ($5 vs $6.25) is on Anthropic's page and on the Artificial Analysis note.
The 1-hour write ($8) and the batch route ($2/$10, cache read $0.10) are OpenRouter's /api/v1/models payload at 18:57 UTC. The launch table does not print those rows.
Fast mode is on the launch page: up to 2.5x speed, at $8 and $40. That is 2x standard Opus 5.5, and it sits above Opus 5's $5/$25.
Nine graded calls, and the bill they produced
All 8 matched calls passed. Opus 5.5 cost $0.010564 against Opus 5 at $0.01269, a 16.8% cut.
OpenRouter chat completions. No temperature. reasoning.effort was low, plus one coding call at high.
Two tasks, two trials. Graders were exact match and executed tests. Cached tokens: 0. Invoice prompts: 104 tokens.
# 22 Sep 2026, 18:57 UTC, OpenRouter
MODELS = ["anthropic/claude-opus-5.5", "anthropic/claude-opus-5"]
# invoice answer must be 1.96
# 800_000 * 0.20/1e6 + 200_000 * 4/1e6 + 50_000 * 20/1e6
# coding: merge_intervals, hidden tests, function only| Set | Opus 5.5 | Opus 5 |
|---|---|---|
| Invoice, 2 trials | 2/2, $0.002652, 45 and 46 out, 3.089s and 5.100s | 2/2, $0.00332, 46 and 46 out, 6.765s and 3.997s |
| merge_intervals, 2 trials | 2/2, $0.007912, 168 and 168 out, 3.597s and 3.239s | 2/2, $0.00937, 161 and 155 out, 3.435s and 3.582s |
| Matched total | $0.010564, 8/8 | $0.01269, 8/8 |
| effort=high, 1 coding call | pass, $0.003936, 167 out, 0 reasoning tokens | not run |
The invoice pair is the clean 20% case: 20.1% cheaper, almost the same completion count. The function pair is the miss: Opus 5.5 wrote 336 completion tokens against 316, and the bill fell only 15.6%. List price dropped 20%. Token count rose. Net saving shrank. effort=high did not lengthen that function: 167 completion tokens, 0 reasoning tokens, $0.003936, tests passed. Artificial Analysis saw the other regime, about 119,000 output tokens per Intelligence Index task at max against about 73,000 for Opus 5, with cost per task still level. Do not book a 40% discount there. Our 9 calls never entered it.
Latency is a tie on this sample. Coding means: 3.42s and 3.51s. Invoice means: 4.09s and 5.38s. The ranges overlap, 5.100s versus 3.997s.
Anthropic says output is more than 30% faster than Opus 5. These 9 calls do not show it.
Where Anthropic's chart leads, and where it does not
Astra leads two numbered rows: AutomationBench at 41.4%, and Terminal-Bench-Science at 64.6%.
Every other numbered cell on the 22 Sep chart is ahead for Opus 5.5. The cover image is that chart. Cells below are copied from it.
Anthropic says these Opus 5.5 numbers use adaptive thinking at max effort, with one flag: Terminal-Bench 4.0 is Opus 5.5 at xhigh and GPT-6 Astra at high, Astra's figure as reported by OpenAI.
| Benchmark | Opus 5.5 | Fable 5.1 | Opus 5 | GPT-6 Astra | GPT-5.6 Sol |
|---|---|---|---|---|---|
| Terminal-Bench 4.0 | 66.4% | 55.8% | 52.3% | 57.9% | 37.3% |
| FrontierCode v1.1 (Main) | 54.4% | 50.3% | 48.0% | 53.3% | 47.5% |
| CursorBench 4.0 | 57.8% | 51.8% | 46.6% | n/a | 41.7% |
| GDPval-AA v2.1 | 1846 | 1735 | 1708 | 1542 | 1588 |
| AutomationBench | 40.0% | 31.4% | 26.9% | 41.4% | 28.8% |
| Humanity's Last Exam, with tools | 67.7% | 65.6% | 63.6% | 57.2% | n/a |
| Terminal-Bench-Science 0.1 | 58.7% | 52.6% | 29.0% | 64.6% | 22.4% |
| OSWorld 2.0, partial | 81.8% | 80.7% | 74.0% | n/a | n/a |
| Chartography, with tools | 89.0% | 88.4% | 83.4% | n/a | n/a |
Read the footnotes before you route on a one-point gap.
- Safeguards stayed on. Cyber fallback was Opus 4.8. Biology and frontier-LLM fallback was Opus 5. Anthropic says that likely lowers the score.
- Zapier ran AutomationBench with no fallback, so an intervention counted as a miss. 40.0% is early access. Astra's 41.4% is the public board.
- Terminal-Bench 4.0 error is ±2.6 points. 66.4 versus 57.9 clears it. FrontierCode 54.4 versus 53.3 does not.
- Anthropic says the daily-use gap versus Fable 5.1 is narrower than this chart.
- Artificial Analysis, same day, other setup: Terminal-Bench 4.0 at 59.6%, level with Astra at xhigh, and Humanity's Last Exam at 61.4% versus 59.1% for Fable 5.1. Do not average 66.4 and 59.6.
Default effort is a different row. FrontierCode at medium is 54.6%. CursorBench at medium is 52.5%, not the 57.8% above. Anthropic says that medium FrontierCode score beats Astra's best at about a fifth of the cost per task.
Limits the $4 price does not lift
Cyber tasks reroute to Opus 4.8, and thinking stays on for every API account.
The cheaper card did not delete the constraints that made the last two Claude launches operationally awkward. What changed on 22 Sep, and what did not:
- Five-hour limits go up on Pro, Max, Team, and seat-based Enterprise, plus a reset you can save. No new message count. The weekly cap is not described as gone. Older window rules are in the usage-limit piece.
- Cyber, biology, and distillation safeguards match the Fable class. Routine bug-fixing stays. Most cyber tasks re-route to Opus 4.8. Biology needs the Life Sciences Verification Program. The wider cyber program was not open on launch day.
- Thinking cannot be switched off. The launch page points at the model docs for that restriction. Preserved thinking, the anti-distillation control from Fable 5.1, blocks API users from editing prior reasoning. It applies to Opus 5.5 on API accounts created on or after 31 Aug 2026.
- The model is watermarked for the EU AI Act. Zero data retention is still offered, same as prior Opus models.
- Containment-boundary attempts fell about 85% versus Opus 5 and Mythos 5.1, Anthropic says, and the rest were low severity and self-reported. The model often suspects an eval, so that audit predicts production poorly.
- Sonnet 5.5 and Haiku 5.5 are not in the API today. The page says they follow in the coming weeks. No date.
If your workload is offensive security or wet-lab design, the $4 price is not the price of an unrestricted Opus 5.5. The classifier spends some of those calls on an older model, and your trace should show the fallback or you will bill the wrong id.
What early testers told Anthropic
GitHub, Optiver, and Box report fewer tokens than Opus 5. Those lines are quotes Anthropic published.
They are signed early-access notes, not an independent board. @claudeai posted at 16:31 UTC. Same-hour replies praised Pro headroom and debugging. No shared trace.
| Who, on the launch page | What they said | What it does not prove |
|---|---|---|
| Mario Rodriguez, GitHub | Fewest tokens they measured. VS Code terminal tasks in under half the steps of Opus 5. | No public log. |
| Noyan Tokgozoglu, Optiver | Opus 5 quality in about half the turns, time, and output tokens. Cost down 40% to 50%. | Their desk. Our cut was 16.8%, cache at 0. |
| Yashodha Bhavnani, Box | A third of the tokens, 40% less verbose, accuracy held. | Their content eval. |
| Zimu Li, Factory | Medium matched Opus 5 at high, with 20% to 25% fewer output tokens. | Our high pin did not shrink tokens. |
| Fabian Hedin, Lovable | A third to half fewer steps, and fewer tokens. | Their builder loop, not the API card. |
The quotes describe fewer steps and fewer tokens, a larger cut than 20% off the list price. Our function did the opposite: 336 completion tokens versus 316.
Both can be true. Anthropic says a 200,000-line audit finished in under 3 hours versus over 20 for Opus 5, at 2.5x fewer tokens. We did not re-run it.
What a cache-heavy turn does to the bill
A 500,000-token cache reread drops from $0.25 on Opus 5 to $0.10 on Opus 5.5.
Arithmetic on the published rates, not one of the 9 calls. Those calls cached nothing, so they only show the 20% cut minus extra tokens. Plug in your own prefix size. The mix moves the percent.
| Workload | Opus 5 | Opus 5.5 | Cut |
|---|---|---|---|
| 500,000 cache-read tokens | $0.25 | $0.10 | 60% |
| 20,000 fresh input + 8,000 output, no cache | $0.30 | $0.24 | 20% |
| 500,000 cache + 20,000 in + 8,000 out | $0.55 | $0.34 | 38% |
| Same 20,000 in + 8,000 out at fast-mode $8/$40 | n/a | $0.48 | a price increase vs standard 5.5 |
38% on that third row is how a 20% token cut becomes "about 40% less on typical workloads": a cache-heavy agent, default settings, plus fewer tokens. Strip the cache and you are back at 20% until the model talks more. Fast mode runs the other way, $0.48 against $0.24 on the uncached slice. The batch row in the price table ($2/$10) is a catalog price. We sent no batch job.
Which id to pin after 22 September
Pin claude-opus-5-5 for new agent traffic unless the task is cyber or biology.
The July Opus 5 piece priced a Fable-class model at half of Fable's token card. That rate is stale as of 22 Sep.
Opus 5.5 is another 20% off Opus 5's input and output, and 60% off cache reads. Safeguards now track Fable 5.1, not the looser Opus line. Grok 4.7, metered the same day, is a different card.
- Uncached, short API calls like the ones we graded: move them to Opus 5.5. Pass rate tied at 8/8. Bill down 16.8% on this set, and 20.1% when the completion counts match.
- Agent loops that reread a large prefix: move them too, and budget something nearer the 38% worked example than the 16.8% invoice. Confirm
cached_tokensis non-zero on the response or you are still on the 20% path. - Max-effort research evals: keep the old Opus 5 budget until you measure. AA's Intelligence Index cost per task did not fall with the token price.
- Cyber offense, biology R&D, distillation-style probing: do not plan on Opus 5.5 answering. Expect Opus 4.8 or Opus 5, or a refusal. Verification programs are the door Anthropic named.
- Fast mode: pay $8/$40 only when 2.5x speed is worth 2x tokens. It is not the cheap tier.
- Leave Fable 5.1 pinned only where you have measured a quality gap these 9 calls cannot see. Fable is 2.5x on input and output. Cache reads are $0.25 versus $0.20.
Frequently Asked Questions
What is the Claude Opus 5.5 API model id?
On the Claude Platform the id is claude-opus-5-5. On OpenRouter it is anthropic/claude-opus-5.5, batch anthropic/claude-opus-5.5:batch. We called the standard id at 18:57 UTC on 22 Sep 2026. Listed context: 1,000,000 tokens. Fast mode is a price ($8/$40), not a second id on the launch page.
Does Claude Opus 5.5 remove the five-hour usage limit?
No. The five-hour limit rises on Pro, Max, Team, and seat-based Enterprise, and one reset can be saved. No message count was published. The weekly cap is not described as removed. API tokens-per-minute limits are a separate meter. Check Settings, then Usage.
Can thinking be turned off on Opus 5.5?
Anthropic says no. The 22 Sep page says thinking can no longer be switched off. Preserved thinking also blocks accounts created on or after 31 Aug 2026 from editing prior reasoning. Our effort=high call still returned a normal function, with 0 reasoning tokens on the OpenRouter usage object. A high pin is not a visible chain of thought.
Is the 40% cost cut real on every workload?
Not on ours. Eight uncached calls cost 16.8% less. The invoice pair alone cost 20.1% less. A worked mix with 500,000 cache-read tokens plus a short completion lands at 38%. Artificial Analysis says max-effort Intelligence Index tasks cost about the same as Opus 5, because output tokens rose about 1.6x.
When do Claude Sonnet 5.5 and Haiku 5.5 ship?
Anthropic said both follow in the coming weeks, on the 22 Sep 2026 page, with no date. Opus 5.5 was the only 5.5 id we could call that evening. A cheaper non-Opus route is still the previous Sonnet or Haiku card.