Anthropic now sells a $4 model that beats its own $10 flagship on eight of the nine benchmarks it publishes. That is a remarkable result — and a fragile one, because four of those margins sit inside the disclosed error bars, and Anthropic says so out loud. Here is the whole comparison, including the places the frontier tier still wins.

TL;DR

Anthropic now sells a $4 model that beats its own $10 flagship on most of the benchmarks it publishes. Claude Opus 5.5 wins eight of the nine rows in Anthropic's comparison table against Claude Fable 5.1 — but the margins are thin (0.6 points on Chartography, 1.1 on FrontierCode) and Anthropic itself says "the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest."

  • Price gap: Opus 5.5 is 2.5× cheaper per token on both input ($4 vs $10) and output ($20 vs $50), and even wins on cache reads ($0.20 vs $0.25).
  • Cost per task: derived from Anthropic's own cost claims, Opus 5.5 runs a FrontierCode task for roughly a fifth of what a $10/$50 model costs.
  • Where Fable 5.1 still wins: the hardest reasoning, the Mythos 5.1 sibling for vetted security and biology work, and one Browserbase agent benchmark (82% vs 74%).
  • The interesting architectural fact: Fable 5.1 can read Opus 5.5 thinking blocks, so you can start on the cheap model and escalate mid-conversation without losing reasoning.
8 / 9
published rows won by
Opus 5.5
2.5×
cheaper per token
input and output
0.6 pt
narrowest margin
Chartography
$0.20
cache read
vs Fable's $0.25

Two tiers, one lab, twenty-one days apart

Claude Fable 5.1 shipped on 1 September 2026. It is a Mythos-class model — the tier Anthropic introduced above Opus in June — with a restricted sibling, Claude Mythos 5.1, reserved for vetted access through the Cyber and Life Sciences verification programmes. Fable 5.1 held headline pricing at $10 / $50 per million tokens, the same as Fable 5, but cut cache reads 75% (from $1.00 to $0.25) and claimed token-billed workloads would run about 25% cheaper than Fable 5, or up to 45% cheaper on highly agentic work.

Twenty-one days later, Anthropic shipped Claude Opus 5.5 at $4 / $20. That is the first time the Opus tier has undercut the Fable tier on cache reads as well as on headline price — and cache reads are, in Anthropic's own words, "the majority of agentic and coding work costs."

SpecificationClaude Opus 5.5Claude Fable 5.1Verdict
Released22 September 20261 September 2026
TierOpus classMythos class (above Opus)Fable
Model IDclaude-opus-5-5claude-fable-5-1
Input price$4 / MTok$10 / MTokOpus 5.5
Output price$20 / MTok$50 / MTokOpus 5.5
Cache read$0.20 / MTok$0.25 / MTokOpus 5.5
Batch discount50% ($2 / $10)not documentedOpus 5.5
Fast mode$8 / $40, up to 2.5× speednot offeredOpus 5.5
Context window1M tokens1M tokenstie
Max output128K (300K on Batches, beta)128Ktie
ThinkingAdaptive, always onAdaptive, always ontie
Default effortmediumhigh
Forced tool userejected (400)rejected (400)tie
Knowledge cutoffJune 2026June 2026tie
Restricted siblingnoneClaude Mythos 5.1 (Cyber / Life Sciences programmes)Fable
Watermarking / ZDRYes / yesYes / —tie
Claude Fable 5.1 $50 Claude Opus 5.5 $20 Cache read /M Fable 5.1 $0.25 vs Opus 5.5 $0.20 Input /M Fable 5.1 $10 vs Opus 5.5 $4
The tier inversion, in one picture. The cheaper model is the newer one, and it also wins the cache-read line — which matters more than the headline for long agent loops, because Anthropic describes cache reads as the majority of agentic and coding cost.

Head-to-head: every published row

BenchmarkOpus 5.5Fable 5.1DeltaWinner
Terminal-Bench 4.0agentic coding · SE ±2.6 (Opus 5.5)66.4%55.8%+10.6Opus 5.5
FrontierCode v1.1 (Main)agentic coding54.4%50.3%+4.1Opus 5.5
CursorBench 4.0coding-agent tasks57.8%51.8%+6.0Opus 5.5
GDPval-AA v2.1knowledge work · Elo1,8461,735+111Opus 5.5
AutomationBenchbusiness workflows · Zapier40.0%31.4%+8.6Opus 5.5
Humanity's Last Examwith tools67.7%65.6%+2.1Opus 5.5
Terminal-Bench-Science 0.1agentic scientific research · SE ±3.5–558.7%52.6%+6.1Opus 5.5
OSWorld 2.0computer use · partial credit81.8%80.7%+1.1Opus 5.5
Chartographyvisual chart recognition · with tools89.0%88.4%+0.6Opus 5.5 (tie)

Opus 5.5 takes every row. That is a genuinely unusual result: a cheaper model in a lower tier beating the flagship across an entire published table. Three facts stop it from being a rout:

  • Four of the nine margins are inside the noise. Chartography (+0.6), OSWorld (+1.1), HLE (+2.1) and FrontierCode (+4.1) are all within or near the disclosed standard errors, and the OSWorld and Terminal-Bench figures come with explicit caveats about partial credit and harness choice.
  • Only Terminal-Bench 4.0 is a large, clean gap. +10.6 points with a standard error of ±2.6 is real. Everything else is a rounding argument dressed as a win.
  • The vendor says so. "At these levels of capability we've found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest."
Opus 5.5 Fable 5.1 Terminal-Bench 4.0 66.4 55.8 FrontierCode v1.1 54.4 50.3 CursorBench 4.0 57.8 51.8 HLE (with tools) 67.7 65.6
Paired bars, four of the nine rows. Opus 5.5 leads everywhere, and the two bars converge as you move from agentic coding toward vision and reasoning-heavy evaluations. The capability gap and the price gap point in opposite directions.

What the price difference actually buys

The sticker gap is 2.5×. The real gap depends on how much of your bill is cache reads, output tokens, and retries. Anthropic gives us enough claims to triangulate.

The FrontierCode derivation

Anthropic states that on FrontierCode, Opus 5.5 "beats GPT-6 Astra at roughly 20% of the cost per task." GPT-6 Astra and Claude Fable 5.1 have identical list pricing — $10 input and $50 output per million tokens. If a $4/$20 model completes the same task for a fifth of what a $10/$50 model costs, then Opus 5.5 is running frontier-class code work at roughly one fifth the cost per task of Fable 5.1.

That is a derived figure, not a published one — it leans on Astra's and Fable 5.1's token usage being in the same region, which no lab has published. But the direction is corroborated from the other side: Artificial Analysis measured Fable 5.1 at max effort at $7.63 per Intelligence Index task versus $3.26 for GPT-6 Astra, and described Astra as matching Fable 5.1's score "at ~40% of the cost per task." Two independent estimates both put a $10/$50 model at roughly 2–2.5× the cost of the alternatives for the same work — with Opus 5.5 claiming to go considerably further.

The HAProxy test — the only same-task comparison

Anthropic ran the cleanest available head-to-head: ask Opus 5.5 and Fable 5.1 to translate HAProxy — widely used load-balancing software — from C into Rust, and grade both rewrites against HAProxy's own regression suite.

ResultClaude Opus 5.5Claude Fable 5.1Delta
Wall-clock time9.5 hours12 hours−21%
CostOpus 5.5 51% cheaper
Regression suitenearly all passednearly all passedtie

This is the single most useful data point in the whole release, because it is the one comparison where both models did the same job under the same grader. The outcome: equivalent correctness, faster wall-clock, and half the bill. If your work looks like a large mechanical translation or migration project, the tier question is already answered.

The research report test

Anthropic asked Opus 5.5, Fable 5.1 and Opus 5 to produce a quarterly earnings report from a copy of the web where the earnings release was hard to find, with an automated grader verifying every figure and quote. 16 of 18 Opus 5.5 reports cleared the quality bar. Fable 5.1 cleared zero. A single invented number failed a report. This is not a speed or cost result; it is a reliability result, and it points the same direction as the price table.

Where Fable 5.1 still earns its premium

A review that only quoted the vendor table would be malpractice. These are the places the frontier tier still wins.

Fable 5.1's real advantages

  • The restricted ceiling. Claude Mythos 5.1 — the same weights family, without the production safeguards — is available only through the Cyber and Life Sciences Verification Programmes. For vetted security researchers and biotech R&D, Opus 5.5's safeguards are routing you to Opus 4.8 and Opus 5. That is a functional gap, not a benchmark gap.
  • Mythos-class reasoning headroom. Independent indices keep Fable 5.1 at the top of the table. On Artificial Analysis v4.3 it shares first place with GPT-6 Astra at max effort; on the earlier v4.2 scale it led outright at 57 against 54 for Opus 5.
  • Browser-agent work at the hardest end. Browserbase, an early tester, reported that on their hardest browser-agent benchmark Fable 5.1 completed 82% of tasks in about 10 minutes each, against 74% for Opus 5 — using fewer tokens than either.
  • Long-session readability. Jane Street reported Fable 5.1 "remains readable over long, multi-step tasks" where prior models became hard to follow.

Where the gap has genuinely closed

  • Cache economics. $0.20 vs $0.25 per million cache reads. The cheap tier now wins the single largest cost line in agentic work.
  • Knowledge-work reliability. 16/18 graded reports clearing QC against 0/18 is the most lopsided same-task result in the release.
  • Business workflows. AutomationBench 40.0% vs 31.4% — a 27% relative improvement, the largest relative margin in the table.
  • Scientific research. 58.7% vs 52.6% on Terminal-Bench-Science 0.1, and roughly double Fable 5's own result (24.7%) a month earlier.
  • Efficiency features Fable 5.1 does not have: batch pricing at 50% off and an $8/$40 fast mode.

The uncomfortable question for Fable 5.1 buyers. Cognition said it would move its Opus 5 traffic in Devin to Fable 5.1 on launch day. That logic — a Mythos-class model at a lower cost per task than the Opus tier — is exactly the logic Anthropic has now inverted three weeks later. If your Fable 5.1 usage is dominated by code review, refactors, migrations and research synthesis rather than by genuinely frontier reasoning, you are paying a 2.5× premium for margins your workload cannot measure.

The architectural detail nobody is talking about

Opus 5.5 and Fable 5.1 share a thinking-block format lineage that most releases do not. According to Anthropic's documentation:

  • Fable 5.1 and Mythos 5.1 can read Opus 5.5 thinking blocks on the Claude API.
  • Opus 5.5 can read blocks from Opus 5 and earlier Opus/Sonnet/Haiku models — but not from Fable or Mythos models.
  • When a request carries a block the target model cannot read, the API drops it, the request still succeeds, and the dropped blocks are not billed.

Translated into a routing pattern: you can start a conversation on Opus 5.5 and escalate to Fable 5.1 mid-task without losing the model's reasoning. The reverse handoff loses it. For teams who want a cheap-by-default, expensive-on-failure pipeline, that is a one-way valve pointing in exactly the useful direction.

# cheap by default, escalate on failure - reasoning survives the hop
conversation = [{"role": "user", "content": task}]

resp = client.messages.create(
    model="claude-opus-5-5",        # $4 / $20 per MTok
    thinking={"type": "adaptive"},  # always on - cannot be disabled
    output_config={"effort": "medium"},
    messages=conversation,
)

if not passes_checkpoint(resp):          # escalate without losing reasoning
    conversation += [{"role": "assistant", "content": resp.content}]
    resp = client.messages.create(       # Fable 5.1 READS Opus 5.5 thinking blocks
        model="claude-fable-5-1",        # $10 / $50 per MTok
        thinking={"type": "adaptive"},
        messages=conversation,
    )

Two caveats. First, Opus 5.5's default effort is medium while Fable 5.1's is high — if you hand a task over, set effort explicitly or the escalated leg runs at a different depth than the first. Second, for API accounts created on or after 31 August 2026, both models enforce a prefix check: if the system prompt, tools or an earlier message changed since a thinking block was produced, replaying that block returns a 400. Keep conversations append-only, and steer with mid-conversation system messages instead of editing history.

Verdict: who should pay the 2.5×

The short version

Default to Opus 5.5. Pay for Fable 5.1 only when a task has already failed twice, or when you are in a vetted security or biology programme that needs Mythos access. The published table goes 8–1 to the cheaper model, the only same-task comparison (HAProxy) goes to the cheaper model on both speed and cost at equal correctness, and the largest same-task reliability test (graded research reports) goes 16–0 to the cheaper model. Anthropic's own words concede the benchmark gap is smaller than the scores suggest.

The remaining reason to pay the premium is a ceiling argument, and ceiling arguments are the hardest to falsify — which is exactly why you should test rather than assume. Build a set of 20 hard tasks you have already solved by hand, run both, and count failures. If Fable 5.1 does not win more than about a third of them, the premium is not buying you anything you can measure.

If this is your workloadRoute toReasoning
Code review, refactors, migrations, auditsOpus 5.5HAProxy: equal correctness, 21% faster, 51% cheaper. 200K-line audit: 3h vs Opus 5's 20h.
Research reports, financial analysis, decksOpus 5.516/18 graded reports passed QC vs 0/18 for Fable 5.1. 1,846 vs 1,735 Elo.
High-volume agentic loopsOpus 5.5$0.20 cache reads, batch at 50% off, and an $8/$40 fast mode Fable 5.1 does not offer.
Business process automationOpus 5.5AutomationBench 40.0% vs 31.4% — the largest relative margin in the table.
Scientific pipelines and toolingOpus 5.5Terminal-Bench-Science 0.1: 58.7% vs 52.6%.
Hardest browser-agent workflowsFable 5.1Browserbase: 82% vs 74% completion on their hardest benchmark, at fewer tokens.
Genuinely frontier reasoning, cost no objectFable 5.1Still the Mythos-class tier; leads the independent Intelligence Index at max effort.
Authorised offensive security, biotech R&DFable 5.1 / Mythos 5.1Opus 5.5's safeguards route these to Opus 4.8 and Opus 5; Mythos access is the only path.
Latency-critical interactive editingOpus 5.5Generates output more than 30% faster than Opus 5, and offers fast mode at $8/$40.

FAQ

Is Opus 5.5 actually better than Fable 5.1, or just cheaper?

On Anthropic's published table it wins all nine rows, but only one margin (Terminal-Bench 4.0, +10.6 points) is large and clean. Four others are within the disclosed error bars. Independent indices still place Fable 5.1 at the top of the frontier at maximum effort. The honest summary is: equivalent on most published measures, clearly cheaper, and the vendor says the real-world gap is narrow.

Why is the cheaper model newer?

Anthropic's framing is that Opus 5.5 "performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5." The company is deliberately shipping efficiency — fewer tokens, cheaper cache reads, faster output — rather than chasing the absolute capability ceiling, which stays with the Mythos-class tier. It is a portfolio decision, not an accident.

Can I mix them in one conversation?

One way, yes. On the Claude API, Fable 5.1 and Mythos 5.1 read Opus 5.5 thinking blocks, so an escalation from Opus 5.5 to Fable 5.1 preserves reasoning. Opus 5.5 cannot read Fable or Mythos blocks — the API silently drops them, the request succeeds, and you are not billed for the dropped blocks.

Does Fable 5.1 have any feature Opus 5.5 lacks?

Yes, three: the restricted Mythos 5.1 sibling for vetted security and life-sciences access; a higher default effort (high vs medium), which matters if you never set effort explicitly; and a general capability ceiling that independent evaluators still rank first at maximum effort. Opus 5.5 has features in the other direction — batch pricing, fast mode, and cheaper cache reads.

What about Sonnet 5.5 — will it make this question moot?

Anthropic says Sonnet 5.5 arrives "in the coming weeks, with many of the same improvements to performance, efficiency, and safety." Sonnet 5 currently sits at $2 / $10. If Sonnet 5.5 inherits the 5.5-family efficiency work, the interesting comparison stops being Opus 5.5 vs Fable 5.1 and becomes Opus 5.5 vs Sonnet 5.5 for everyday coding.

Run the same task on both tiers

Pick three hard tasks from last month, run them on Opus 5.5 and Fable 5.1, and count the failures. If the flagship does not win more than a third, you just saved 2.5× on every call this quarter.

Open a new chat on CodingFleet →

Sources & further reading

Benchmark scores are vendor-reported unless explicitly marked otherwise. Cross-vendor numbers come from different harnesses, effort settings and tool configurations and are directional, not strictly comparable. Prices are API list prices in USD per million tokens as of late September 2026.