
Claude Opus 5 Is Here: Near-Fable-5 Intelligence at Half the Price
Anthropic shipped Claude Opus 5 on July 24, 2026, and the pitch is blunt: close to Fable 5's frontier intelligence, at half the price. The headline proof is Frontier-Bench v0.1, where Opus 5 scores 43.3% and more than doubles Opus 4.8's 21.1%. The catch, and there's always one: it doesn't win everything. GPT-5.6 Sol still edges it on one agentic-coding test, and on health and biology the crown goes to Mythos 5. Here's what actually changed, the full benchmark table, and whether you should re-point your default model today.
Key takeaways:
- Claude Opus 5 launched July 24, 2026 on the Claude API, Claude.ai, Claude Code, and Claude Cowork, with the model ID
claude-opus-5. - It leads most of Anthropic's published benchmarks (Frontier-Bench, GDPval-AA, ARC-AGI-3, OSWorld 2.0, AutomationBench), but trails on a couple of coding, legal, and science tests.
- Pricing is unchanged from Opus 4.8 at $5/$25 per million tokens, with a Fast mode at roughly 2x the price for about 2.5x the speed.
What Shipped Today: Claude Opus 5 in One Minute
Claude Opus 5 is Anthropic's new frontier model, released July 24, 2026, and available the same day on the API, Claude.ai, Claude Code, and Claude Cowork. The model ID is claude-opus-5. Anthropic's framing, per the announcement, is that it "comes close to the frontier intelligence of Claude Fable 5 at half the price." Standard pricing stays at $5 input / $25 output per million tokens, the same as Opus 4.8.
This is a real generational step for agent work, not a point release. If you're coming from Claude Opus 4.8, the API shape and Claude Code integration carry over unchanged, so the migration is a model-ID swap. What's different is the ceiling: on Frontier-Bench v0.1, Anthropic says Opus 5 more than doubles 4.8's score at a lower cost per task, and it posts state-of-the-art numbers on GDPval-AA and ARC-AGI-3.
The positioning matters. Fable 5 is Anthropic's most expensive frontier model. Getting within striking distance of it at Opus prices is the whole story of this release, and it's why the "half the price" line is the one Anthropic leads with.
What's Actually New in Opus 5 vs 4.8?
The biggest 4.8-to-5 shift is judgment: Opus 5 is markedly better at verifying its own work and iterating until a task is actually done, not just plausibly finished. Anthropic also cites stronger software engineering, knowledge work, and scientific research, plus improved visual output. Pricing and the core API stayed put.
Better judgment and self-verification
The clearest behavioral change is that Opus 5 checks itself. Anthropic quotes Zimu Li, a technical staff member, describing the model as one that "verifies branches, checks templates" rather than charging ahead. For anyone who relies on an agent to catch its own mistakes in production, that's the trait that moves the needle. It's also the hardest thing to sell on a benchmark chart, so treat it as a claim to test on your own tasks.
Frontier intelligence at Opus cost
Sualeh Asif, co-founder of Cursor, summed the release up as "near Fable 5 intelligence at Opus speed and cost." That's the practical read: teams that wanted Fable-5-class reasoning but couldn't justify the price now have a middle option. On OSWorld 2.0, Anthropic says Opus 5 surpasses Fable 5 at roughly one-third of the cost, which is the cleanest single example of the value argument.
Alignment and safety posture
Anthropic reports Opus 5 has its lowest misaligned-behavior score to date (2.3) and calls it the model most aligned to its Constitution. On the security side, the cybersecurity classifiers are described as about 85% less restrictive than Fable 5's, with an enterprise Cyber Verification Program for teams that need it. As with any vendor safety claim, useful to know, still worth validating against your own red-team cases.
Opus 5 Benchmarks: Where It Wins (and Where It Loses)
Opus 5 leads most of Anthropic's published benchmarks, and the wins that matter for agent builders are lopsided: Frontier-Bench v0.1 (43.3%), GDPval-AA (1861 Elo), ARC-AGI-3 (30.2%), OSWorld 2.0 (70.6%), and Zapier's AutomationBench (26.0%). The honest exceptions: GPT-5.6 Sol wins DeepSWE v1.1, Fable 5 edges the FrontierCode and legal tests, and Mythos 5 takes health and biology.
| Benchmark | Opus 5 | Fable 5 | Opus 4.8 | GPT-5.6 Sol |
|---|---|---|---|---|
| Agentic terminal coding, Frontier-Bench v0.1 | 43.3% | 33.7% | 21.1% | 34.4% |
| Knowledge work, GDPval-AA v2 (Elo) | 1861 | 1747 | 1593 | 1736 |
| Novel problem-solving, ARC-AGI-3 | 30.2% | — | 1.5% | 7.8% |
| Agentic search, BrowseComp | 90.8% | 87.4% | 84.3% | 90.4% |
| Multidisciplinary reasoning, Humanity's Last Exam (with tools) | 64.7% | 63.9% | 57.9% | — |
| Agentic computer use, OSWorld 2.0 | 70.6% | 66.1% | 55.7% | 62.6% |
| Agentic coding, DeepSWE v1.1 | 68.8% | 69.7% | 59.0% | 72.7% |
| Agentic coding, FrontierCode v1.1 (Main) | 53.4% | 53.5% | 46.5% | 47.5% |
| Business workflows, AutomationBench | 26.0% | 17.4% | 17.0% | 18.1% |
| Legal Agent Benchmark (held-out) | 11.7% | 13.3% | 10.4% | 2.5% |
| Health, HealthBench Professional | 59.8% | 66.0% | 57.4% | 60.5% |
| Biology, BioMysteryBench (hard) | 49.4% | 46.5% | 42.4% | — |

Read the deltas, not just the absolute scores. Frontier-Bench jumped +22.2 points over Opus 4.8 (21.1 to 43.3), the biggest single leap and the reason Anthropic calls this a coding-and-agents release. GDPval-AA gained +268 Elo (1593 to 1861), clearing every model in the table on knowledge work. ARC-AGI-3 is the eye-catcher: 30.2% versus 7.8% for GPT-5.6 Sol and 1.5% for Opus 4.8, which Anthropic frames as roughly 3x the next-best public model on novel problem-solving. On AutomationBench, 26.0% is about 1.5x the field, a direct signal for anyone building business-process agents.
Two honest notes on the losses. First, the leader in the Fable 5 column for health (66.0%) and the biology human-solved split is actually Mythos 5, Anthropic's science-specialized model, not Fable 5, so those aren't like-for-like general-model wins. Second, the coding picture is genuinely split: GPT-5.6 Sol takes DeepSWE v1.1 (72.7% vs 68.8%), and Fable 5 wins FrontierCode by a hair (53.5% vs 53.4%). If your work is pure SWE-bench-style coding, the "best overall" label doesn't automatically pick Opus 5.
What Does Opus 5 Cost? Pricing and Fast Mode
Standard Opus 5 pricing is $5 per million input tokens and $25 per million output tokens, identical to Opus 4.8 per Anthropic's published pricing, so there's no price increase for the upgrade. The "half the price" claim is relative to Fable 5, not to 4.8: you get near-Fable-5 intelligence without paying Fable-5 rates. Fast mode runs at about 2x the base price for roughly 2.5x the default speed, and the API supports fallback routing.
So what does that mean per task? On standard pricing you pay nothing extra versus 4.8 and get a large capability jump, which makes the upgrade close to free for most workloads. Fast mode is the interactive-session lever: you trade a 2x price premium for output that streams about 2.5x faster on the same model, which pays off when you're sitting there waiting and hurts when you're running overnight batches where wall-clock time is irrelevant. If rate limits are your real constraint, our guide to Claude usage limits covers how the tiers interact with cost.
# Claude Code: point your default model at Opus 5
/model claude-opus-5
# API request body (model ID swap):
{
"model": "claude-opus-5",
"messages": [
{ "role": "user", "content": "Refactor this module and verify the tests pass." }
]
}Where Opus 5 Lands for Production Agents
The gains cluster exactly where production agents live: terminal coding (Frontier-Bench), computer use (OSWorld 2.0), agentic search (BrowseComp), and multi-step business workflows (AutomationBench). Those four are the workloads that break most often in real deployments, so a model that jumps on all four at once changes what's shippable.
That's the part worth sitting with. A benchmark like AutomationBench measures whether an agent can complete a real multi-tool workflow end to end, and Opus 5's 26.0% versus roughly 17% for the field is the difference between "demos well" and "survives contact with a messy CRM." The OSWorld 2.0 result (70.6%, beating Fable 5 at a third of the cost) says the same thing for computer-use agents that click through actual software. For teams shipping customer-facing AI agents, those are the numbers that translate into fewer 2 a.m. failures.
It also lowers the cost of a design pattern we lean on: cheap self-verification. When the model is genuinely better at checking its own branches, you can spend fewer tokens on a separate critic pass and still catch the same errors. That's a real budget line for anyone running agents at volume.
How We're Rolling Opus 5 Into Our Stack
We build and run production AI agents on the Claude stack at Techsy, so a frontier release is a same-day evaluation, not a headline we read and move on from. Here's our honest first read, qualitative where we don't yet have clean before/after numbers, specific where we do.
The swap itself is trivial, and that's the point. Our content and automation pipelines pin a default model in config; changing claude-opus-4-8 to claude-opus-5 and restarting a session needed zero other changes, same as every recent Opus bump. If you route across providers and want to A/B Opus 5 against Fable 5 or GPT-5.6 Sol on your own tasks, a proxy lets you switch models without rewriting your app.
What we're watching first is the self-verification claim, because it's the one that would actually change our architecture. Our current agents run an explicit critic step to catch bad edits before they land; if Opus 5 reliably flags its own weak branches, we can thin that step out and save the tokens. We're not publishing a number on that until we've run it across a repeatable task set, the same discipline we'd want from anyone quoting agent metrics. If you want the method, it's the one in our guide to evaluating AI agents in production: same prompts, same repo state, count the wrong turns, not the vibes.
Should You Switch From Opus 4.8? (Switch / Wait / Stay)
Whether to move depends on your workload, not on which model "wins" overall. The agentic and knowledge-work jumps are large and real, so most teams should switch. Pure-coding shops with a GPT or Fable pipeline have a reason to test before committing, and near-ceiling users on cheap tasks can stay put with almost nothing lost.
Switch now if: your work leans on agentic tasks, computer use, business-process automation, or knowledge work. Frontier-Bench (+22.2 over 4.8), AutomationBench (~1.5x the field), and GDPval-AA (+268 Elo) are the biggest real gains, and pricing didn't change, so there's no cost penalty to upgrading.
Wait if: your workflow is pure agentic coding and you're already on GPT-5.6 Sol or Fable 5 for it. GPT-5.6 Sol still leads DeepSWE v1.1 and Fable 5 edges FrontierCode, so run your own coding eval before you re-point. Also wait if you're mid-project and a model swap would muddy an evaluation you're already running.
Stay on 4.8 if: you're cost-sensitive on tasks that were already near-ceiling for you and you don't touch the agentic workloads where Opus 5 pulls ahead. You'd be paying switching cost for a delta you won't feel. For the wider field, see how the models stack up in our Claude Sonnet 5 breakdown and the Fable 5 and Mythos 5 overview.
We pick and wire models like this for client agent builds every week. If you're deciding which model to standardize on for a production agent, get a free consultation and we'll help you match the model to the workload instead of the marketing.
About the Author
Mert Batur Gurbuz is Co-Founder of Techsy.io, where the team ships AI agents, automation systems, and voice/SDR pipelines for B2B clients. He studies at the University of Birmingham and writes about the LLM tooling stack the Techsy team actually runs in production.
Co-Founder, Techsy.io, University of Birmingham · LinkedIn
Frequently Asked Questions
What is Claude Opus 5?
Claude Opus 5 is Anthropic's frontier model, released July 24, 2026, with the model ID claude-opus-5. Anthropic positions it as coming close to Fable 5's intelligence at half the price, and it leads most of its published benchmarks, including Frontier-Bench, GDPval-AA, and ARC-AGI-3. It's available on the Claude API, Claude.ai, Claude Code, and Claude Cowork.
When was Claude Opus 5 released?
Claude Opus 5 was released on July 24, 2026. It went live the same day across the Claude API, Claude.ai, Claude Code, and Claude Cowork, with no staged rollout, so you can re-point your default model to claude-opus-5 immediately.
Is Claude Opus 5 better than Opus 4.8?
For most work, yes. Opus 5 more than doubles Opus 4.8 on Frontier-Bench v0.1 (43.3% vs 21.1%), gains 268 Elo on GDPval-AA, and jumps from 1.5% to 30.2% on ARC-AGI-3. Pricing is unchanged, so there's no cost penalty. The exceptions are a few coding and science tests where GPT-5.6 Sol, Fable 5, or Mythos 5 still lead.
How much does Claude Opus 5 cost?
Standard pricing is $5 per million input tokens and $25 per million output tokens, identical to Opus 4.8. The "half the price" line refers to Fable 5, not 4.8: Opus 5 delivers near-Fable-5 intelligence at Opus rates. Fast mode costs about 2x the base price and runs roughly 2.5x faster on the same model.
Does Claude Opus 5 beat Fable 5?
Not across the board. Opus 5 surpasses Fable 5 on OSWorld 2.0, BrowseComp, GDPval-AA, and Frontier-Bench, and it does so at a fraction of Fable 5's price. But Fable 5 still edges FrontierCode and the legal benchmark, and Mythos 5 (Anthropic's science model) leads health and biology. The pitch is "near Fable 5 at half the price," not "better than Fable 5 at everything."
What does Opus 5 lose at?
GPT-5.6 Sol wins agentic coding on DeepSWE v1.1 (72.7% vs 68.8%), Fable 5 wins FrontierCode v1.1 by 0.1 points and the held-out legal benchmark (13.3% vs 11.7%), and Mythos 5 leads HealthBench Professional and the biology tests. If your workload is dominated by any of those, benchmark it before you switch.
Is Claude Opus 5 good for building AI agents?
It's built for it. The largest gains land on agentic terminal coding (Frontier-Bench), computer use (OSWorld 2.0), agentic search (BrowseComp), and multi-step business workflows (AutomationBench), which are the workloads production agents run. Its stronger self-verification also lets you spend fewer tokens on a separate critic pass.
How do I switch to Opus 5 in Claude Code and the API?
Set your model to claude-opus-5 (use /model claude-opus-5 in Claude Code) and you're done for most teams. In the API, change the model field to claude-opus-5; the request shape is unchanged from Opus 4.8, so no other code changes are needed.
What is Fast mode in Claude Opus 5?
Fast mode is a higher-speed tier that runs Opus 5 at roughly 2.5x the default speed for about 2x the standard price. It keeps you on the full claude-opus-5 model rather than routing you to a smaller one, which makes it useful for interactive sessions where you're waiting on output and not worth it for latency-insensitive batch jobs.