On June 9, 2026, Anthropic launched Claude Fable 5 — its first Mythos-class model available to the public. At $10/$50 per million tokens (input/output), it was the undisputed king of the Anthropic lineup. Six weeks later, on July 24, Anthropic launched Claude Opus 5 at half the price ($5/$25). And then something unexpected happened: the cheaper model started winning.
Opus 5 leads Fable 5 on 7 of 12 shared benchmarks, including a 9.6-point blowout on Frontier-Bench v0.1 (43.3% vs 33.7%) and a 4.5-point lead on OSWorld 2.0. Fable 5 holds onto SWE-bench Pro by just 0.8 points (80.0% vs 79.2%) and CursorBench by 0.3 points. The rest is either Opus 5 or a tie.
This isn't just a comparison — it's a cannibalization. Anthropic's own $25 model makes its $50 flagship nearly impossible to justify for most coding work. Here's the full breakdown.
| Spec | Claude Opus 5 | Claude Fable 5 |
|---|---|---|
| Released | July 24, 2026 | June 9, 2026 |
| API Model ID | claude-opus-5 | claude-fable-5 |
| Model Tier | Opus (flagship workhorse) | Mythos-class (premium flagship) |
| Input Price | $5.00 / 1M tokens | $10.00 / 1M tokens |
| Output Price | $25.00 / 1M tokens | $50.00 / 1M tokens |
| Cached Input | $0.50 / 1M tokens | $1.00 / 1M tokens |
| Context Window | 1,000,000 tokens | 1,000,000 tokens |
| Max Output | 128,000 tokens | 128,000 tokens |
| Knowledge Cutoff | May 2026 | January 2026 |
| Reasoning | Adaptive (always on), effort: low→max | Adaptive (always on), effort: low→max |
| Fast Mode | Yes — 2.5× speed at 2× price ($10/$50) | Not offered |
| Data Retention | No mandatory retention | 30-day mandatory retention |
| Safety Fallback | Routes to Opus 4.8 (~85% fewer classifier hits) | Routes to Opus 4.8 on cyber/bio/chemistry |
| BenchAlign Score | 85.88 (#1 overall) | 82.76 (#3 overall) |
Benchmark Scoreboard: Opus 5 Wins 7–3 (2 Ties)
Across 12 shared benchmarks, Opus 5 leads on 7, Fable 5 leads on 3, and 2 are ties. The most striking result: Opus 5 wins the benchmarks that matter most for real-world coding.
| Benchmark | What It Measures | Claude Opus 5 | Claude Fable 5 | Δ | Winner |
|---|---|---|---|---|---|
| SWE-bench Pro | Real GitHub issues resolved | 79.2% | 80.0% | −0.8 | Fable 5 |
| SWE-bench Verified | Curated software tasks | 96.0% | 95.0% | +1.0 | Opus 5 |
| Frontier-Bench v0.1 | Agentic coding | 43.3% | 33.7% | +9.6 | Opus 5 |
| Terminal-Bench 2.1 | CLI agent tasks | 89.1% | 88.0% | +1.1 | Opus 5 |
| DeepSWE v1.1 | Long-horizon engineering | 68.8% | 69.7% | −0.9 | Fable 5 |
| GDPval-AA v2 | Knowledge work (Elo) | 1,861 | 1,747 | +114 | Opus 5 |
| OSWorld 2.0 | Computer use | 70.6% | 66.1% | +4.5 | Opus 5 |
| CursorBench 3.2 | In-editor coding | 70.1% | 70.4% | −0.3 | Fable 5 |
| AutomationBench | Business automation | 26.0% | 17.4% | +8.6 | Opus 5 |
| HLE (w/ tools) | Expert reasoning | 64.8% | 64.5% | +0.3 | Tie |
| FrontierCode Diamond | Hardest coding subset | 29.3% | 29.3% | 0.0 | Tie |
| MCP Atlas | Tool orchestration | 85.8% | 83.3% | +2.5 | Opus 5 |
The Four Benchmarks That Tell the Whole Story
1. Frontier-Bench v0.1 — The 9.6-Point Stunner
Frontier-Bench is Anthropic's own agentic coding benchmark — real software engineering tasks requiring multi-file changes, debugging, and feature building. It's the eval that most directly maps to what developers actually pay for.
Opus 5 scores 43.3%. Fable 5 scores 33.7%. That's a 9.6-point gap — and Opus 5 achieves it at roughly half the cost per task. This is the result that reframes the entire comparison. The pre-launch consensus was that Opus 5 would land "comparable to Fable 5, but won't surpass it." It didn't just surpass — it opened a wide lead on the benchmark that matters most for coding.
On CursorBench 3.2, the in-editor coding benchmark measured independently by Cursor, Opus 5 lands within 0.3 points of Fable 5's peak (70.1% vs 70.4%) — at half the cost per task. For teams running AI in the editor, that's near-parity for half the bill.
2. SWE-bench Pro — Fable 5's Last Stand
SWE-bench Pro is the hardest real-world coding benchmark: 1,865 real GitHub issues from actively maintained repositories. Fable 5 scores 80.0%. Opus 5 scores 79.2%. That 0.8-point gap is Fable 5's strongest remaining argument — and it's razor-thin.
For context: Opus 4.8 scored 69.2%. Opus 5 closed 10.8 of the 11.1-point gap to Fable 5 in a single generation. The remaining 0.8 points cost 2× the price. For most teams, that math doesn't work.
3. OSWorld 2.0 — Computer Use Dominance
OSWorld 2.0 measures computer use: navigating GUIs, clicking buttons, filling forms, reading screens. Opus 5 scores 70.6% to Fable 5's 66.1% — a 4.5-point lead — while spending roughly half the budget (~$25 vs ~$47 per task).
If you're building computer-use agents, Opus 5 is the clear pick. It's better and cheaper.
4. GDPval-AA v2 — Knowledge Work
GDPval-AA v2 is human-graded economic knowledge work expressed as an Elo rating. Opus 5 scores 1,861 to Fable 5's 1,747 — a 114-point Elo gap. For teams using AI as a day-to-day knowledge assistant rather than a pure coding engine, Opus 5 is the stronger model.
Where Fable 5 Still Makes Sense
Fable 5 isn't obsolete. It still wins on:
- SWE-bench Pro — by 0.8 points. If you're doing the hardest real-world bug fixes and every fraction of a percent matters, Fable 5 is marginally better.
- DeepSWE v1.1 — by 0.9 points. Long-horizon engineering tasks favor Fable 5's deeper reasoning.
- CursorBench 3.2 — by 0.3 points. In-editor coding at max effort, Fable 5 holds a negligible edge.
- Cybersecurity & biology research — Fable 5 (as Mythos 5's public face) remains stronger on offensive cyber and long-running autonomous biology tasks. Opus 5 was intentionally not trained on cyber tasks.
But these wins come with a catch: Fable 5's safety classifiers route ~5% of requests to Opus 4.8 — a model that scores 69.2% on SWE-bench Pro. When Fable 5 falls back, you're not getting Mythos-class performance. Opus 5's classifiers intervene ~85% less often, and when they do, the fallback is still Opus 4.8 — but you're paying Opus 5 prices, not Fable 5 prices.
Pricing: 2× the Cost for 0.8 Points
| Pricing Tier | Claude Opus 5 | Claude Fable 5 |
|---|---|---|
| Input (per 1M tokens) | $5.00 | $10.00 |
| Output (per 1M tokens) | $25.00 | $50.00 |
| Cached input (per 1M tokens) | $0.50 | $1.00 |
| Batch API (input/output) | $2.50 / $12.50 | $5.00 / $25.00 |
| Fast mode | $10 / $50 (2.5× speed) | N/A |
At 100M output tokens per month, Opus 5 saves $2,500/month versus Fable 5. That's $30,000/year — the cost of a junior developer. For that savings, you give up 0.8 points on SWE-bench Pro and gain leads on 7 other benchmarks.
Even at batch pricing, Opus 5 ($12.50/1M output) is half of Fable 5 ($25.00/1M output). There is no pricing scenario where Fable 5 wins on value.
The Data Retention Difference
This is the factor most comparisons miss. Fable 5 carries a mandatory 30-day data retention requirement — every prompt and output is stored on Anthropic's infrastructure for 30 days for safety classifier analysis. This applies even to organizations that previously had zero-data-retention agreements.
Opus 5 has no mandatory retention. For enterprises with compliance requirements, healthcare companies with HIPAA obligations, or anyone handling proprietary code — this alone may disqualify Fable 5 regardless of benchmark scores.
Speed & Latency
| Speed Metric | Claude Opus 5 | Claude Fable 5 |
|---|---|---|
| Output speed (max effort) | ~59.8 tok/s | ~63.4 tok/s |
| Time to first token (max effort) | ~66s | ~109s |
| Fast mode speed | ~150 tok/s (2.5×, 2× price) | N/A |
Fable 5 generates tokens slightly faster once it starts (63.4 vs 59.8 tok/s), but Opus 5 starts answering 43 seconds sooner (66s vs 109s TTFT). For interactive use, Opus 5 feels significantly more responsive. And Opus 5's Fast Mode at ~150 tok/s is an option Fable 5 doesn't offer at all.
The Verdict
| Choose Claude Opus 5 if... | Choose Claude Fable 5 if... |
|---|---|
| You want better benchmarks at half the price | You need the absolute highest SWE-bench Pro score (by 0.8 pts) |
| You're building computer-use or automation agents | You're doing offensive cybersecurity research |
| You need zero data retention for compliance | You're doing long-running autonomous biology research |
| You want faster time-to-first-token (66s vs 109s) | You need marginally faster token generation (63.4 vs 59.8 tok/s) |
| You want Fast Mode option (2.5× speed) | You're already on Fable 5 and the switching cost exceeds the savings |
| You want fewer safety classifier interventions | You need the best DeepSWE score (by 0.9 pts) |
Bottom Line
Claude Opus 5 makes Claude Fable 5 nearly impossible to justify for most coding work. It leads on 7 of 12 benchmarks, costs exactly half as much, starts answering 43 seconds faster, offers a Fast Mode, and doesn't force 30-day data retention. The benchmarks Fable 5 wins — SWE-bench Pro by 0.8 points, DeepSWE by 0.9, CursorBench by 0.3 — are margins so thin they disappear in real-world variance.
Fable 5 remains the pick for offensive cybersecurity research and autonomous biology work — domains where Opus 5 was intentionally not trained. And if you're already deep in a Fable 5 pipeline with optimized prompts and caching, the switching cost may outweigh the savings.
But for everyone else? Opus 5 is the best model Anthropic has ever shipped. It's not just better value — it's better, period. And at $25 per million output tokens, it's the clearest recommendation in the history of the Claude lineup.
Sources: Anthropic Opus 5 system card (July 24, 2026), Anthropic Fable 5 & Mythos 5 system card (June 9, 2026), BenchLM.ai, Artificial Analysis, ClaudeFast, Codersera, Vellum. Prices verified against official API documentation as of July 25, 2026.