ai-machine-learning

Claude Fable 5.1 Review: We Metered It Against Fable 5 and Opus 5

Written by Mert Batur
Sep 2, 2026
9 read
Claude Fable 5.1 Review: We Metered It Against Fable 5 and Opus 5

Claude Fable 5.1 Review: We Metered It Against Fable 5 and Opus 5

Claude Fable 5.1 shipped on 2026-09-01 at the same $10/$50 as Fable 5, cache reads cut to $0.25. We ran 14 matched OpenRouter tasks the next morning at 08:55 UTC. The 45% agentic saving was not on that invoice: all three models passed 14/14, and 5.1 billed $0.14693 against Opus 5 at $0.080725.

Anthropic's generally available frontier ID is claude-fable-5-1, GA 2026-09-01, $10/$50 per million tokens, cache read $0.25. Mythos 5.1 is the same weights with different (looser cyber/bio) safeguards, Glasswing only. Knowledge cutoff Jun 2026. Thinking always on.

Claude Fable 5.1 vs Fable 5 vs Opus 5: the switch table

Stay on Opus 5 for this 14-task set; Fable 5.1 matched accuracy and cost 1.11× Fable 5.

We called anthropic/claude-fable-5.1 the morning after GA. The list card did not move: $10/$50, same as Fable 5. Claude Opus 5 is $5/$25. Cache read is the only list-price change ($0.25 vs $1.00), and it did not fire here.

Opus 5 is 45% cheaper than Fable 5.1 on this 14-task set because the list price is $5/$25 vs $10/$50, not because it used fewer tokens.

Fable 5.1Fable 5Opus 5
List $ in/out$10/$50$10/$50$5/$25
Cache read$0.25 / MTok$1.00 / MTokn/a
Default effortHigh in Claude Code; Medium on claude.ai / Coworkn/aHigh
Vendor headlineTerminal-Bench-Science 52.6%24.7%29.0%
Invoice (our meter)$0.14693, 14/14$0.13285, 14/14$0.080725, 14/14
Switchcache-heavy agentsshort uncached callsbounded single-shot

Same weights, two doors, and one of them will not answer

Mythos 5.1 is the same weights with looser cyber/bio safeguards, and it is not on OpenRouter.

Anthropic's launch sentence is literal: "Claude Fable 5.1 and Claude Mythos 5.1 are the same model, but with different levels of safeguards." Fable is GA (claude-fable-5-1, Bedrock anthropic.claude-fable-5-1). Mythos is claude-mythos-5-1 behind Project Glasswing / CVP / LSVP, currently a set of US organizations. 1M context, 128k max output, Jun 2026 cutoff vs Fable 5's Jan 2026. Our experiment cost $0.53795. Mythos was not called.

python
# OpenRouter slugs we actually sent on 2026-09-02
MODELS = [
    "anthropic/claude-fable-5.1",
    "anthropic/claude-fable-5",
    "anthropic/claude-opus-5",
    # claude-mythos-5-1: not listed / Glasswing only
]

Which Claude Fable 5.1 benches actually moved?

Terminal-Bench-Science 0.1 is the jump: 52.6% on Fable 5.1 versus 24.7% on Fable 5.

Fable 5 sat at 24.7% and Opus 5 at 29.0%. The jump is why agentic coding loops matter more than our 14 short calls.

BenchFable 5.1Fable 5Opus 5GPT-5.6 Sol
Terminal-Bench-Science 0.152.6%24.7%29.0%22.4%
Terminal-Bench 4.055.8% (Mythos 5.1: 60.9%)42.0%52.3%37.3%
GDPval-AA v21853172318241711
OSWorld 2.0 partial / strict77.9% / 41.7%72.9% / 36.1%75.4% / 39.6%n/a
HLE no tools / with tools60.9% / 65.0%57.8% / 63.8%56.6% / 63.6%n/a
AutomationBench31.4%17.1%26.9%19.6%
CursorBench 3.2.073.4%70.5%70.0%67.2%

Source: Anthropic announcement, 2026-09-01.

SourceFindingScope
Artificial Analysis, 2026-09-01Intelligence Index 66 at maxFable 5 max 62; Opus 5 max 63
Artificial Analysis$3.76 / task vs Fable 5 $3.14 (+20%)~1.7× output; ~$5.16 without the cut; xhigh 65 at $2.72
Snorkel, 2026-09-0158% fewer tokens, 36% fastervs Opus 5 successes only; 18 / 5 / 2 / 2; pass@1 61.5%; build/dep 18% vs 67%

Snorkel's 58% / 36% is vs Opus 5 on successful runs only, not vs Fable 5. Opus still solved three more tasks (18 both / 5 Opus-only / 2 Fable-only).

  1. Terminal-Bench-Science SE ±3.5-4.5 pts. Public leaderboard (3 trials/task, Claude Code runner) reports Opus 5 30.0% and Fable 5 21.4%; Anthropic's setup reproduces 29.0% and 24.7%, both within noise.
  2. OSWorld 2.0 is the authors' August 2026 task release; Fable 5 and Opus 5 were re-run under the same conditions. Not comparable to older OSWorld numbers.
  3. Production safeguards were on. Interventions scored as zero on OSWorld (both Fables) and on AutomationBench (Fable 5). Other cyber tasks fell back to Opus 4.8; biology to Opus 5. That likely understates raw capability.

Is the 25% Claude Fable 5.1 saving on your invoice?

The cache-read line is 75% cheaper; the whole short follow-up was only 21% cheaper.

Anthropic's launch post estimates ~25% typical and ~45% agentic savings vs Fable 5. Both need cache reads to dominate, and output tokens not to jump. The docs rate card: cache write $12.50 / $20, output $50, cache read $1.00 → $0.25 / MTok. Workload cost models decide if that cut survives cutting the token bill elsewhere.

ScenarioLabelFable 5Fable 5.1Delta
Row A: 20-turn / ~1M (9.5M cache reads, writes $12.50, output $50)list-price identity, not our meter$72.00 ($9.50 cache)$64.88 ($2.38 cache)-9.9%
Row B: Intelligence Index taskArtificial Analysis, 2026-09-01$3.14$3.76 (~$5.16 without the cut)+20% at max (~1.7× output)
Row C: 1,165-token prefix, turn 2Techsy OpenRouter, 2026-09-02$0.005086$0.00402125whole turn 21% cheaper
Turn-2 line (Invoice, our meter)Fable 5.1Fable 5
Cached input tokens1,1651,166
Cache-read line$0.000291$0.001166
OpenRouter usage.cost$0.00402125$0.005086

The cache-read line is 75.0% cheaper ($0.000291 vs $0.001166). Whole-turn savings stop at 21% because output still bills at $50/M. Scale the cache line to 40 rereads × 200,000 tokens: $2.00 on 5.1 vs $8.00 on 5.

AY Automate asked "Is Claude Fable 5.1 more expensive than Fable 5?" on 2026-09-01 and answered "No." List price is unchanged at $10/$50. The invoice is not: Artificial Analysis measured +20% at max, THE DECODER updated the same day with that +20%, and our 14 uncached calls billed 1.11×.

HN user george_max (item 49525611, 2026-09-01) called it "just cache reads" and "15% more," matching AA at max effort. Subscription 5-hour windows do not price cache reads the API way.

One quiz, not a bench: Fable 5's turn-2 cache answer used $2.00 / $2.50 (5.1 prices) on a prompt that said "on this model."

What our OpenRouter meter showed on Claude Fable 5.1

All three models passed 14/14 at effort=low; Opus 5 was 45% cheaper than Fable 5.1.

We benchmarked 14 prompts. Accuracy tied; Opus 5 still won on $5/$25 vs $10/$50.

  1. When: 2026-09-02, 08:55 UTC, OpenRouter chat-completions. Public rates on the Fable 5.1 model page.
  2. What: anthropic/claude-fable-5.1 vs anthropic/claude-fable-5 vs anthropic/claude-opus-5.
  3. How: reasoning.effort=low except three Fable 5.1 High coding calls. No temperature (5.1 rejects non-default sampling).
  4. Graders: JSON-schema, exact numeric match, executed unit tests. Artefacts: fable_bench.py, benchmark-raw.json, benchmark-aggregate.json. We tested the three slugs in our benchmark: 49 calls, $0.53795.
Metric (14 calls, effort=low)Fable 5.1Fable 5Opus 5
Tasks passed14/1414/1414/14
Median latency5.647 s5.383 s4.797 s
Median completion tokens9490103
Median reasoning tokens0823
Total completion tokens2,5472,2712,843
Total reasoning tokens0253380
Total cost$0.14693$0.13285$0.080725

Coding passed 6/6 each at low.

Task (2 trials)Fable 5.1 outFable 5 outOpus 5 out
merge_intervals175, 16896, 96155, 155
semver_satisfies414, 411405, 401493, 487
lru_ttl_cache387, 418395, 339469, 459
Coding median latency6.523 s5.389 s6.316 s
Coding total cost$0.11095$0.09878$0.06154

Treat Anthropic's "more concise in plans and summaries" as a claim about agent traces, not about single-shot function dumps. Fable 5.1 spent ~175 tokens on merge_intervals where Fable 5 spent 96, both passing.

What this does not measure:

  • Mythos 5.1 (not callable from this account).
  • Multi-hour agents, Terminal-Bench, SWE-bench, computer use.
  • Default effort (we pinned low except round 3).
  • Cache writes at 5-minute vs 1-hour TTL.
  • Refusals, fallback to Opus, tool-loop batching.

Which three requests return 400 on Fable 5.1?

Forced tool_choice of type any or tool returns HTTP 400 on claude-fable-5-1.

What's new in Fable 5.1 names the error: tool_choice: type "tool" and "any" are not supported for this model. Prefill of the assistant turn is also 400. Non-default temperature / top_p / top_k is also 400; we sent no temperature for that reason. Thinking blocks bind to the producing model (thinking-binding-controls-2026-08-01): 5.1 reads earlier blocks, earlier models cannot read 5.1's, and a router fallback drops them. Editing anything before a 5.1 thinking block errors or drops it for accounts created on/after 2026-08-31. Mythos 5.1 skips that prefix check.

python
# HTTP 400 invalid_request_error on anthropic/claude-fable-5.1
payload = {
    "model": "anthropic/claude-fable-5.1",
    "messages": [{"role": "user", "content": "Look up the cache rate."}],
    "tools": [{"type": "function", "function": {"name": "lookup"}}],
    "tool_choice": {"type": "tool", "name": "lookup"},  # type "any" 400s too
}
# error: tool_choice: type "tool" and "any" are not supported for this model.
  1. Set tool_choice to auto and strict: true on the tool schema, or switch to structured outputs.
  2. Keep history append-only; never edit a prefix in front of a 5.1 thinking block.
  3. Drop thinking blocks when you fall back to an earlier model.
  4. Omit temperature / top_p / top_k, and do not prefill the assistant turn.

Why does Claude Code default Fable 5.1 to High?

Claude Code defaults Fable 5.1 to High; on lru_ttl_cache that 4×'d a task that already passed at Low.

Anthropic's announcement sets High in Claude Code and Medium on claude.ai / Cowork. Five effort levels: low / medium / high / xhigh / max. Thinking is always on; thinking: disabled is 400. Help Center requires Claude Code 2.1.250 or later.

TaskLow (out / reason / cost)High (out / reason / cost / latency)
merge_intervals172 / 0 / $0.010175 / 0 / $0.01024 / 4.86 s
semver_satisfies413 / 0 / $0.023401 / 0 / $0.02238 / 7.60 s
lru_ttl_cache403 / 0 / $0.0221,713 / 1,150 / $0.08798 / 22.07 s

merge_intervals and semver_satisfies looked like Low. lru_ttl_cache did not. High spent 4.25× the completion tokens (1,713 vs 403), added 1,150 reasoning tokens, took 22.07 s vs ~7 s, and cost a 4.0× bill ($0.08798 vs $0.022) for tests we already had. Pin effort=low (or Medium on claude.ai) for bounded functions.

Simon Willison (2026-09-01, not our pelican): effort cannot be off; max was 33× the output of low on one prompt.

Should you stay on Fable 5, stay on Opus 5, or move?

Stay on Opus 5 for bounded single-shot work; move cache-heavy agents to Fable 5.1.

Docs say start with Opus 5 and reach for Fable 5.1 when High-effort evals still fall short. Our eval is the 14/14 meter: Opus 5 at $0.0807 vs $0.1469 (45% cheaper). Do not switch Fable 5 to 5.1 for short uncached calls: 14/14 both, 5.1 cost 1.11× more here.

RuleFlip ifInvoice (our meter)
Stay on Opus 5 for bounded single-shotHigh-effort evals still miss long agent sessions$0.080725, 14/14, 45% cheaper than $0.14693
Move cache-heavy agents to Fable 5.1the session never rereads a large prefix75.0% cheaper cache-read line ($0.000291 vs $0.001166)
Do not leave Claude Code on default High for tasks that already pass at Lowthe job is a Terminal-Bench-Science-class loop4.0× on lru_ttl_cache ($0.08798 vs $0.022)
Do not switch Fable 5 → 5.1 for short uncached callsyou need the Science jump or the $0.25 cache read1.11× ($0.14693 vs $0.13285), 14/14 both

Fable 5.1 billed $0.14693 (1.11×) on uncached 14/14; the win is the cache-read line. Fable 5 billed $0.13285 until a 1,165-token prefix rereads. Opus 5 billed $0.080725, 45% cheaper, with more tokens (2,843 vs 2,547).

Bounded single-shot work stays on Opus 5. Cache-heavy agents move to Fable 5.1 for the $0.25 cache-read line, with effort pinned to Low when tests already pass. Short uncached Fable 5 calls stay put.

Frequently Asked Questions

What is the Claude Fable 5.1 API ID?

claude-fable-5-1 on the Anthropic API, anthropic.claude-fable-5-1 on Bedrock, and anthropic/claude-fable-5.1 on OpenRouter. Vertex, Foundry, and Claude Platform keep claude-fable-5-1. Mythos is claude-mythos-5-1, Glasswing-only; it is not an OpenRouter model you can type, and we did not call it on 2026-09-02.

Does the cache-read cut apply to Claude Code Max quotas?

No. The 75% cut is an API cache-read price ($1.00 → $0.25 / MTok). Subscription 5-hour windows do not bill cache reads that way, so the 25% typical saving does not stretch Max quotas. r/ClaudeCode (2026-09-01) reported burning the 5-hour window in 20 minutes after the High default landed.

What Claude Code version does Fable 5.1 need?

2.1.250 or later, per the Help Center on 2026-09-01. Fable 5 was 2.1.170+. Older builds miss the new ID, the High default that 4.0×'d our LRU task, and the effort picker 5.1 expects. Upgrade before you point traffic at it.

Is Fable 6 next?

5.1 is the point-release that post predicted. Fable 6 is still unannounced. Polymarket threads asking for a 5.1 date are stale: GA was 2026-09-01, and the announcement names no Fable 6 calendar. Do not stall a 5 vs 5.1 switch waiting for it.

Is Claude Fable 5.1 a Covered Model?

Yes. 30-day data retention. Not available under zero data retention unless expressly authorized. Enterprise Frontier Safeguards (customer-cloud storage) roll out later this fall; eligible customers can use ZDR until then. A statistical text watermark ships on all platforms (EU AI Act, models after 2026-08-02); the detection API is in private preview.

Tags

claude-fable-5-1claude-mythos-5-1prompt-cachingclaude-opus-5openrouter

Share this article

Start Your Project

Ready to build something extraordinary?

Let's turn your vision into reality. Our team is ready to help you create software that makes a difference.