
Claude Fable 5.1 Review: We Metered It Against Fable 5 and Opus 5
Claude Fable 5.1 shipped on 2026-09-01 at the same $10/$50 as Fable 5, cache reads cut to $0.25. We ran 14 matched OpenRouter tasks the next morning at 08:55 UTC. The 45% agentic saving was not on that invoice: all three models passed 14/14, and 5.1 billed $0.14693 against Opus 5 at $0.080725.
Anthropic's generally available frontier ID is claude-fable-5-1, GA 2026-09-01, $10/$50 per million tokens, cache read $0.25. Mythos 5.1 is the same weights with different (looser cyber/bio) safeguards, Glasswing only. Knowledge cutoff Jun 2026. Thinking always on.
Claude Fable 5.1 vs Fable 5 vs Opus 5: the switch table
Stay on Opus 5 for this 14-task set; Fable 5.1 matched accuracy and cost 1.11× Fable 5.
We called anthropic/claude-fable-5.1 the morning after GA. The list card did not move: $10/$50, same as Fable 5. Claude Opus 5 is $5/$25. Cache read is the only list-price change ($0.25 vs $1.00), and it did not fire here.
Opus 5 is 45% cheaper than Fable 5.1 on this 14-task set because the list price is $5/$25 vs $10/$50, not because it used fewer tokens.
| Fable 5.1 | Fable 5 | Opus 5 | |
|---|---|---|---|
| List $ in/out | $10/$50 | $10/$50 | $5/$25 |
| Cache read | $0.25 / MTok | $1.00 / MTok | n/a |
| Default effort | High in Claude Code; Medium on claude.ai / Cowork | n/a | High |
| Vendor headline | Terminal-Bench-Science 52.6% | 24.7% | 29.0% |
| Invoice (our meter) | $0.14693, 14/14 | $0.13285, 14/14 | $0.080725, 14/14 |
| Switch | cache-heavy agents | short uncached calls | bounded single-shot |
Same weights, two doors, and one of them will not answer
Mythos 5.1 is the same weights with looser cyber/bio safeguards, and it is not on OpenRouter.
Anthropic's launch sentence is literal: "Claude Fable 5.1 and Claude Mythos 5.1 are the same model, but with different levels of safeguards." Fable is GA (claude-fable-5-1, Bedrock anthropic.claude-fable-5-1). Mythos is claude-mythos-5-1 behind Project Glasswing / CVP / LSVP, currently a set of US organizations. 1M context, 128k max output, Jun 2026 cutoff vs Fable 5's Jan 2026. Our experiment cost $0.53795. Mythos was not called.
# OpenRouter slugs we actually sent on 2026-09-02
MODELS = [
"anthropic/claude-fable-5.1",
"anthropic/claude-fable-5",
"anthropic/claude-opus-5",
# claude-mythos-5-1: not listed / Glasswing only
]Which Claude Fable 5.1 benches actually moved?
Terminal-Bench-Science 0.1 is the jump: 52.6% on Fable 5.1 versus 24.7% on Fable 5.
Fable 5 sat at 24.7% and Opus 5 at 29.0%. The jump is why agentic coding loops matter more than our 14 short calls.
| Bench | Fable 5.1 | Fable 5 | Opus 5 | GPT-5.6 Sol |
|---|---|---|---|---|
| Terminal-Bench-Science 0.1 | 52.6% | 24.7% | 29.0% | 22.4% |
| Terminal-Bench 4.0 | 55.8% (Mythos 5.1: 60.9%) | 42.0% | 52.3% | 37.3% |
| GDPval-AA v2 | 1853 | 1723 | 1824 | 1711 |
| OSWorld 2.0 partial / strict | 77.9% / 41.7% | 72.9% / 36.1% | 75.4% / 39.6% | n/a |
| HLE no tools / with tools | 60.9% / 65.0% | 57.8% / 63.8% | 56.6% / 63.6% | n/a |
| AutomationBench | 31.4% | 17.1% | 26.9% | 19.6% |
| CursorBench 3.2.0 | 73.4% | 70.5% | 70.0% | 67.2% |
Source: Anthropic announcement, 2026-09-01.
| Source | Finding | Scope |
|---|---|---|
| Artificial Analysis, 2026-09-01 | Intelligence Index 66 at max | Fable 5 max 62; Opus 5 max 63 |
| Artificial Analysis | $3.76 / task vs Fable 5 $3.14 (+20%) | ~1.7× output; ~$5.16 without the cut; xhigh 65 at $2.72 |
| Snorkel, 2026-09-01 | 58% fewer tokens, 36% faster | vs Opus 5 successes only; 18 / 5 / 2 / 2; pass@1 61.5%; build/dep 18% vs 67% |
Snorkel's 58% / 36% is vs Opus 5 on successful runs only, not vs Fable 5. Opus still solved three more tasks (18 both / 5 Opus-only / 2 Fable-only).
- Terminal-Bench-Science SE ±3.5-4.5 pts. Public leaderboard (3 trials/task, Claude Code runner) reports Opus 5 30.0% and Fable 5 21.4%; Anthropic's setup reproduces 29.0% and 24.7%, both within noise.
- OSWorld 2.0 is the authors' August 2026 task release; Fable 5 and Opus 5 were re-run under the same conditions. Not comparable to older OSWorld numbers.
- Production safeguards were on. Interventions scored as zero on OSWorld (both Fables) and on AutomationBench (Fable 5). Other cyber tasks fell back to Opus 4.8; biology to Opus 5. That likely understates raw capability.
Is the 25% Claude Fable 5.1 saving on your invoice?
The cache-read line is 75% cheaper; the whole short follow-up was only 21% cheaper.
Anthropic's launch post estimates ~25% typical and ~45% agentic savings vs Fable 5. Both need cache reads to dominate, and output tokens not to jump. The docs rate card: cache write $12.50 / $20, output $50, cache read $1.00 → $0.25 / MTok. Workload cost models decide if that cut survives cutting the token bill elsewhere.
| Scenario | Label | Fable 5 | Fable 5.1 | Delta |
|---|---|---|---|---|
| Row A: 20-turn / ~1M (9.5M cache reads, writes $12.50, output $50) | list-price identity, not our meter | $72.00 ($9.50 cache) | $64.88 ($2.38 cache) | -9.9% |
| Row B: Intelligence Index task | Artificial Analysis, 2026-09-01 | $3.14 | $3.76 (~$5.16 without the cut) | +20% at max (~1.7× output) |
| Row C: 1,165-token prefix, turn 2 | Techsy OpenRouter, 2026-09-02 | $0.005086 | $0.00402125 | whole turn 21% cheaper |
| Turn-2 line (Invoice, our meter) | Fable 5.1 | Fable 5 |
|---|---|---|
| Cached input tokens | 1,165 | 1,166 |
| Cache-read line | $0.000291 | $0.001166 |
OpenRouter usage.cost | $0.00402125 | $0.005086 |
The cache-read line is 75.0% cheaper ($0.000291 vs $0.001166). Whole-turn savings stop at 21% because output still bills at $50/M. Scale the cache line to 40 rereads × 200,000 tokens: $2.00 on 5.1 vs $8.00 on 5.
AY Automate asked "Is Claude Fable 5.1 more expensive than Fable 5?" on 2026-09-01 and answered "No." List price is unchanged at $10/$50. The invoice is not: Artificial Analysis measured +20% at max, THE DECODER updated the same day with that +20%, and our 14 uncached calls billed 1.11×.
HN user george_max (item 49525611, 2026-09-01) called it "just cache reads" and "15% more," matching AA at max effort. Subscription 5-hour windows do not price cache reads the API way.
One quiz, not a bench: Fable 5's turn-2 cache answer used $2.00 / $2.50 (5.1 prices) on a prompt that said "on this model."
What our OpenRouter meter showed on Claude Fable 5.1
All three models passed 14/14 at effort=low; Opus 5 was 45% cheaper than Fable 5.1.
We benchmarked 14 prompts. Accuracy tied; Opus 5 still won on $5/$25 vs $10/$50.
- When: 2026-09-02, 08:55 UTC, OpenRouter chat-completions. Public rates on the Fable 5.1 model page.
- What:
anthropic/claude-fable-5.1vsanthropic/claude-fable-5vsanthropic/claude-opus-5. - How:
reasoning.effort=lowexcept three Fable 5.1 High coding calls. Notemperature(5.1 rejects non-default sampling). - Graders: JSON-schema, exact numeric match, executed unit tests. Artefacts:
fable_bench.py,benchmark-raw.json,benchmark-aggregate.json. We tested the three slugs in our benchmark: 49 calls, $0.53795.
| Metric (14 calls, effort=low) | Fable 5.1 | Fable 5 | Opus 5 |
|---|---|---|---|
| Tasks passed | 14/14 | 14/14 | 14/14 |
| Median latency | 5.647 s | 5.383 s | 4.797 s |
| Median completion tokens | 94 | 90 | 103 |
| Median reasoning tokens | 0 | 8 | 23 |
| Total completion tokens | 2,547 | 2,271 | 2,843 |
| Total reasoning tokens | 0 | 253 | 380 |
| Total cost | $0.14693 | $0.13285 | $0.080725 |
Coding passed 6/6 each at low.
| Task (2 trials) | Fable 5.1 out | Fable 5 out | Opus 5 out |
|---|---|---|---|
merge_intervals | 175, 168 | 96, 96 | 155, 155 |
semver_satisfies | 414, 411 | 405, 401 | 493, 487 |
lru_ttl_cache | 387, 418 | 395, 339 | 469, 459 |
| Coding median latency | 6.523 s | 5.389 s | 6.316 s |
| Coding total cost | $0.11095 | $0.09878 | $0.06154 |
Treat Anthropic's "more concise in plans and summaries" as a claim about agent traces, not about single-shot function dumps. Fable 5.1 spent ~175 tokens on merge_intervals where Fable 5 spent 96, both passing.
What this does not measure:
- Mythos 5.1 (not callable from this account).
- Multi-hour agents, Terminal-Bench, SWE-bench, computer use.
- Default effort (we pinned
lowexcept round 3). - Cache writes at 5-minute vs 1-hour TTL.
- Refusals, fallback to Opus, tool-loop batching.
Which three requests return 400 on Fable 5.1?
Forced tool_choice of type any or tool returns HTTP 400 on claude-fable-5-1.
What's new in Fable 5.1 names the error: tool_choice: type "tool" and "any" are not supported for this model. Prefill of the assistant turn is also 400. Non-default temperature / top_p / top_k is also 400; we sent no temperature for that reason. Thinking blocks bind to the producing model (thinking-binding-controls-2026-08-01): 5.1 reads earlier blocks, earlier models cannot read 5.1's, and a router fallback drops them. Editing anything before a 5.1 thinking block errors or drops it for accounts created on/after 2026-08-31. Mythos 5.1 skips that prefix check.
# HTTP 400 invalid_request_error on anthropic/claude-fable-5.1
payload = {
"model": "anthropic/claude-fable-5.1",
"messages": [{"role": "user", "content": "Look up the cache rate."}],
"tools": [{"type": "function", "function": {"name": "lookup"}}],
"tool_choice": {"type": "tool", "name": "lookup"}, # type "any" 400s too
}
# error: tool_choice: type "tool" and "any" are not supported for this model.- Set
tool_choicetoautoandstrict: trueon the tool schema, or switch to structured outputs. - Keep history append-only; never edit a prefix in front of a 5.1 thinking block.
- Drop thinking blocks when you fall back to an earlier model.
- Omit
temperature/top_p/top_k, and do not prefill the assistant turn.
Why does Claude Code default Fable 5.1 to High?
Claude Code defaults Fable 5.1 to High; on lru_ttl_cache that 4×'d a task that already passed at Low.
Anthropic's announcement sets High in Claude Code and Medium on claude.ai / Cowork. Five effort levels: low / medium / high / xhigh / max. Thinking is always on; thinking: disabled is 400. Help Center requires Claude Code 2.1.250 or later.
| Task | Low (out / reason / cost) | High (out / reason / cost / latency) |
|---|---|---|
merge_intervals | 172 / 0 / $0.010 | 175 / 0 / $0.01024 / 4.86 s |
semver_satisfies | 413 / 0 / $0.023 | 401 / 0 / $0.02238 / 7.60 s |
lru_ttl_cache | 403 / 0 / $0.022 | 1,713 / 1,150 / $0.08798 / 22.07 s |
merge_intervals and semver_satisfies looked like Low. lru_ttl_cache did not. High spent 4.25× the completion tokens (1,713 vs 403), added 1,150 reasoning tokens, took 22.07 s vs ~7 s, and cost a 4.0× bill ($0.08798 vs $0.022) for tests we already had. Pin effort=low (or Medium on claude.ai) for bounded functions.
Simon Willison (2026-09-01, not our pelican): effort cannot be off; max was 33× the output of low on one prompt.
Should you stay on Fable 5, stay on Opus 5, or move?
Stay on Opus 5 for bounded single-shot work; move cache-heavy agents to Fable 5.1.
Docs say start with Opus 5 and reach for Fable 5.1 when High-effort evals still fall short. Our eval is the 14/14 meter: Opus 5 at $0.0807 vs $0.1469 (45% cheaper). Do not switch Fable 5 to 5.1 for short uncached calls: 14/14 both, 5.1 cost 1.11× more here.
| Rule | Flip if | Invoice (our meter) |
|---|---|---|
| Stay on Opus 5 for bounded single-shot | High-effort evals still miss long agent sessions | $0.080725, 14/14, 45% cheaper than $0.14693 |
| Move cache-heavy agents to Fable 5.1 | the session never rereads a large prefix | 75.0% cheaper cache-read line ($0.000291 vs $0.001166) |
| Do not leave Claude Code on default High for tasks that already pass at Low | the job is a Terminal-Bench-Science-class loop | 4.0× on lru_ttl_cache ($0.08798 vs $0.022) |
| Do not switch Fable 5 → 5.1 for short uncached calls | you need the Science jump or the $0.25 cache read | 1.11× ($0.14693 vs $0.13285), 14/14 both |
Fable 5.1 billed $0.14693 (1.11×) on uncached 14/14; the win is the cache-read line. Fable 5 billed $0.13285 until a 1,165-token prefix rereads. Opus 5 billed $0.080725, 45% cheaper, with more tokens (2,843 vs 2,547).
- Fable 5 vs Opus 4.8 lives on the June launch post.
- Claude Opus 5 owns Opus-line pricing.
- 5.1 is the point-release that post predicted; Fable 6 is still unannounced.
- Export-control Q&A for the older Fable 5 pull.
- Token-efficiency as a ranking spine lives on Grok 4.6.
- Route Opus vs Fable in production.
Bounded single-shot work stays on Opus 5. Cache-heavy agents move to Fable 5.1 for the $0.25 cache-read line, with effort pinned to Low when tests already pass. Short uncached Fable 5 calls stay put.
Frequently Asked Questions
What is the Claude Fable 5.1 API ID?
claude-fable-5-1 on the Anthropic API, anthropic.claude-fable-5-1 on Bedrock, and anthropic/claude-fable-5.1 on OpenRouter. Vertex, Foundry, and Claude Platform keep claude-fable-5-1. Mythos is claude-mythos-5-1, Glasswing-only; it is not an OpenRouter model you can type, and we did not call it on 2026-09-02.
Does the cache-read cut apply to Claude Code Max quotas?
No. The 75% cut is an API cache-read price ($1.00 → $0.25 / MTok). Subscription 5-hour windows do not bill cache reads that way, so the 25% typical saving does not stretch Max quotas. r/ClaudeCode (2026-09-01) reported burning the 5-hour window in 20 minutes after the High default landed.
What Claude Code version does Fable 5.1 need?
2.1.250 or later, per the Help Center on 2026-09-01. Fable 5 was 2.1.170+. Older builds miss the new ID, the High default that 4.0×'d our LRU task, and the effort picker 5.1 expects. Upgrade before you point traffic at it.
Is Fable 6 next?
5.1 is the point-release that post predicted. Fable 6 is still unannounced. Polymarket threads asking for a 5.1 date are stale: GA was 2026-09-01, and the announcement names no Fable 6 calendar. Do not stall a 5 vs 5.1 switch waiting for it.
Is Claude Fable 5.1 a Covered Model?
Yes. 30-day data retention. Not available under zero data retention unless expressly authorized. Enterprise Frontier Safeguards (customer-cloud storage) roll out later this fall; eligible customers can use ZDR until then. A statistical text watermark ships on all platforms (EU AI Act, models after 2026-08-02); the detection API is in private preview.