GPT-6 Astra vs Claude Fable 5.1: The Frontier Duel

Released September 1–3, 2026 · OpenAI vs Anthropic · Both $10 / $50 per 1M tokens · Both 1M-token context

Every benchmark, every independent index, every real-world head-to-head — and who actually wins your workload.

Astra: Computer Use + Math + CyberFable 5.1: Reasoning + Cache EconomicsAA Index: 53 vs 53 (tied)Cost/Task: $3.26 vs $7.63

1 · TL;DR: The Short Version

Two frontier flagships shipped 48 hours apart in September 2026 — Claude Fable 5.1 (Anthropic, Sep 1) and GPT-6 Astra (OpenAI, Sep 3). Same list price ($10/$50 per 1M tokens), same 1M-token context, same 128K max output. The comparison is a genuine split decision, and the two sides disagree with each other about who wins.

"It depends on whose benchmarks you read. On OpenAI's own comparison table, GPT-6 Astra leads Claude Fable 5.1 on nearly every published row… Artificial Analysis, an independent evaluator, reverses that." — DataCamp
Astra wins ~10vendor rows + cost per task
Fable 5.1 wins ~3HLE + cache economics + AA-Briefcase
Both AA indices tied53 vs 53 · 62 vs 62
97.6%
Astra · FrontierMath T4
65.0%
Fable · HLE w/ tools
100%
Astra · ExploitBench
53
AA Intelligence Index · tied

GPT-6 Astra (OpenAI)

Released Sep 3, 2026 · gpt-6-astra

Wins: Computer use (OSWorld 72.6%, ScreenSpot-Pro 92.7%), math (FrontierMath T4 97.6%), professional artifacts (AutomationBench 41.4%, BenchCAD 95.9%), cybersecurity (ExploitBench 100%, first "Critical" model), terminal work (Terminal-Bench 4.0 57.7%), and cost per task ($3.26 vs $7.63 on AA's index — same 53-point score at ~40% of the bill).

Claude Fable 5.1 (Anthropic)

Released Sep 1, 2026 · claude-fable-5-1

Wins: The hard-reasoning rows (Humanity's Last Exam 65.0% vs 57.2%) and the AA evaluations with a knowledge or science flavour (AA-Briefcase, SciCode), cache economics ($0.25 vs $1.00 cached reads, no long-context surcharge), and qualitative front-end/code-review praise from testers like Theo. It no longer leads either AA composite: Index v4.3 is a 53–53 tie and the Coding Agent Index is 62–62.

2 · The Two Models at a Glance

FeatureGPT-6 AstraClaude Fable 5.1
Developer / releaseOpenAI · September 3, 2026Anthropic · September 1, 2026
API model IDgpt-6-astraclaude-fable-5-1
Context window1.05M tokens1M tokens
Max output128K tokens128K tokens
Knowledge cutoffApril 30, 2026June 2026
List price (in / out)$10 / $50 per 1M$10 / $50 per 1M
Cached input read$1.00 ($2.00 above 272K)$0.25 (75% cut)
Long-context surchargeYes — above 272K input tokens (2× in, 1.5× out)No — full 1M at standard rates
Reasoning modeReasoning model; Fast mode at 2× priceAdaptive thinking, always on
Cyber capability tier"Critical" (first OpenAI model) — gated behind DaybreakLower tier; can discover vulnerabilities, not develop exploits
EU AI Act watermark—Yes (released after Aug 2, 2026)
Enterprise noteOff by default per workspaceRequires 30-day data retention; not on Priority Tier

Sources: DataCamp, Anthropic, Vellum.

🎛 Interactive Head-to-Head Explorer

Computer Use (OSWorld 2.0) Coding (Terminal-Bench 4.0) Math (FrontierMath T4) Reasoning (HLE w/ tools) Professional Work (AutomationBench) Cybersecurity (ExploitBench)
GPT-6 Astra Claude Fable 5.1

Axis values: OSWorld 2.0 (72.6 vs 77.9*), Terminal-Bench 4.0 (57.7 vs 55.8), FrontierMath Tier 4 v2 (97.6 vs 87.8), HLE w/ tools (57.2 vs 65.0), AutomationBench (41.4 vs 31.4), ExploitBench (100 vs 70). *Anthropic's 77.9% is on the benchmark authors' August 2026 task release and is not comparable to OpenAI's 72.6% on the offline set — see §7.

FrontierMath Tier 4 v2
97.6%
87.8%
Humanity's Last Exam (w/ tools)
65.0%
57.2%
Terminal-Bench 4.0
57.7%
55.8%
AutomationBench
41.4%
31.4%
ExploitBench
100%
70.0%
AA Intelligence Index
53
53
FrontierMath T4 v2
+9.8
87.8
Terminal-Bench Science 0.1
+12.0
52.6
DeepSWE v1.1
+6.7
67.4
AutomationBench
+10.0
31.4
BenchCAD
+11.6
84.3
Humanity's Last Exam (w/ tools)
+7.8
57.2
AA Intelligence Index
tie
53

Blue = Astra margin, orange = Fable margin. Vendor-reported unless noted; see tables below.

AA cost per task (max)
$3.26
Astra — ~40% of Fable's bill
AA cost per task (max)
$7.63
Fable 5.1 — same 53 points, 2.3× the cost
AA cost per task (low)
$0.82
Astra — cheapest index effort tier
AA coding-agent cost/task
$7.09
Astra (max) — 40% under Fable
Cached input read
$0.25
Fable — 4× cheaper cache
Cached input read
$1.00
Astra ($2.00 above 272K)
Blended price / 1M
$7.70
Astra (7:2:1 cache ratio)
Blended price / 1M
$7.17
Fable (7:2:1 cache ratio)

Per-task figures from Artificial Analysis via DataCamp. Figures are from Artificial Analysis Intelligence Index v4.3 and Coding Agent Index v1.4 (Sep 7–9, 2026): both models now score 53 on the composite, and Astra's cost ladder runs down to $0.82/task at low effort. Superseded v4.1 figures: 61 vs 66, $1.67 vs $3.76.

3 · The Vendor Tables: Each Side Wins Its Own

The cleanest way to see the disagreement: OpenAI's table shows Astra ahead on nearly every row it published; Anthropic's table shows Fable 5.1 ahead on every row it published. Both are real numbers — they just measure different things, on different task releases, with different harnesses.

OpenAI's comparison (Astra vs Fable 5.1)

BenchmarkGPT-6 AstraClaude Fable 5.1Winner
FrontierMath Tier 4 v297.6%87.8%Astra
GPQA Diamond96.0%93.7%Astra
Humanity's Last Exam (w/ tools)57.2%65.0%Fable
Terminal-Bench Science 0.164.6%52.6%Astra
Terminal-Bench 4.057.7%55.8%Astra
DeepSWE v1.174.1%67.4%Astra
FrontierCode 1.1 Main53.3%50.9%≈ tie
ScreenSpot-Pro (no tools)92.7%87.3%*Astra
AutomationBench41.4%31.4%Astra
BenchCAD95.9%84.3%Astra
ExploitBench100%70.0%Astra
Computer-use safety ↓2.4%9.5%Astra

*ScreenSpot-Pro figure is for Claude Fable 5 (from Mythos), not 5.1. Source: OpenAI's GPT-6 Astra announcement table, via DataCamp and Vellum. ↓ = lower is better.

Anthropic's comparison (Fable 5.1 vs the field)

BenchmarkFable 5.1Opus 5Fable 5GPT-5.6 Sol
Terminal-Bench-Science 0.152.6%29.0%24.7%22.4%
Terminal-Bench 4.055.8% (60.9 Mythos)52.3%42.0%37.3%
GDPval-AA v2 (points)1853182417231711
OSWorld 2.0 partial77.9%75.4%72.9%—
OSWorld 2.0 strict41.7%39.6%36.1%—
Humanity's Last Exam (no tools)60.9%56.6%57.8%—
Humanity's Last Exam (w/ tools)65.0%63.6%63.8%—
AutomationBench31.4%26.9%17.1%19.6%
CursorBench 3.2.073.4%70.0%70.5%67.2%

Source: Anthropic announcement, via Vellum. Standard error ±3.5–4.5 pts per model on Terminal-Bench-Science. Fable 5.1 ran with production safeguards enabled; where safeguards intervened it scored zero on OSWorld, likely understating raw capability.

Why both tables can be "true": the two vendors rarely run the same task release. Anthropic's OSWorld 2.0 numbers use the benchmark authors' August 2026 task files and are explicitly not comparable to OpenAI's offline-set 72.6%. OpenAI's ARC-AGI-3 99.9% uses its stateful provider-adapter harness (standard harness: 62.7%). And the AA Coding Agent Index runs Astra in Codex but Fable 5.1 in Claude Code — different scaffolding. Compare carefully, or let the independent index arbitrate.

4 · The Independent Arbitrator: Artificial Analysis

When both vendors claim victory, the tiebreaker is Artificial Analysis, which runs both models itself. Its verdict, on the revised index (v4.3, September 7, 2026): the two are level on intelligence — 53 points apiece — and Astra gets there at roughly 40% of the per-task cost. AA has now reworked the index twice since launch (v4.1: Fable 5.1 66–61; v4.2: 57–55; v4.3: 53–53), so every AA number below is date-stamped.

Independent metricGPT-6 AstraClaude Fable 5.1Winner
AA Intelligence Index v4.3 (max effort)5353tie
AA Coding Agent Index v1.462 (in Codex)62 (in Claude Code)tie*
Cost per Intelligence Index task (max)$3.26$7.63Astra
Output tokens per index task (max)27k78kAstra
Blended price per 1M tokens$7.70$7.17Fable
Output speed71.3 tok/s70.5 tok/s≈ tie
Time to first token384.3s264.9sFable
Context window1M1Mtie

*Both AA composites are now dead heats. Coding Agent Index v1.4: Astra 62 in Codex, Fable 5.1 62 in Claude Code — though Astra reaches that score ~40% cheaper per task ($7.09 at max effort). The harnesses still differ, and AA ran Fable 5.1 with Anthropic's default server-side fallback, which served ~4% of output tokens via Opus 4.8/5 on safety-flagged requests — the score is Fable 5.1 "as most people will actually call it." Earlier drafts of this table used Index v4.1, where Fable 5.1 led 66–61 (v4.2 narrowed it to 57–55); v4.3, published September 7, 2026, closed it to 53–53.

"Artificial Analysis puts Fable 5.1 in Claude Code at 70, the highest score on its Coding Agent Index, and Astra in Codex at 67… Astra wins the published coding rows by a nose and wins the efficiency argument outright, while Fable 5.1 holds the only independent coding-agent score that clears them both." — DataCamp

Update (Sep 7–9, 2026): Artificial Analysis' revised suite changed that verdict. On Intelligence Index v4.3 both models score 53 — Astra ties Fable 5.1 instead of trailing it — and on Coding Agent Index v1.4 both score 62, with Astra at ~40% (Intelligence) and ~60% (Coding Agent) of Fable 5.1's cost per task.

5 · Coding: A Tie That Costs Very Differently

This is the most contested dimension. On the benchmarks each vendor publishes, Astra leads by a nose; on the independent index the two are now level — 62 apiece on AA's Coding Agent Index v1.4, with Astra getting there ~40% cheaper per task; on cost, Astra wins outright.

Coding benchmarkGPT-6 AstraClaude Fable 5.1Notes
Terminal-Bench 4.057.7%55.8%Uncontested — both vendors publish 55.8% for Fable; AA's own run has it 59.1% vs 52.0%
DeepSWE v1.174.1%67.4%OpenAI-reported
FrontierCode 1.1 Main53.3%50.9%Within noise of Fable 5 (53.5%) and Opus 5 (53.4%)
CursorBench 3.2.0Not published73.4%Anthropic-reported; SpaceXAI confirmed at max effort
AA Coding Agent Index v1.462 (Codex)62 (Claude Code)Independent tie; Astra is $7.09/task at max effort, ~40% under Fable 5.1

The efficiency story is the part that survives scrutiny. Artificial Analysis found Astra matches Fable 5-class results "at less than half the cost, driven by significant token efficiency gains." In the "No Hype Assessment" YouTube comparison, Terminal-Bench 4.0 at max effort came out 56.7% (Astra) vs 55.8% (Fable) — but cost $10.35 vs $19.50. A Reddit agent-work comparison (r/better_claw) measured ~$1.10 per task for Astra vs ~$0.65 for Fable, noting Astra is verbose — "the thinking tax is real" — but OpenAI doesn't charge separately for reasoning tokens.

"Theo's verdict: Astra is the best model he has ever used, but he is careful to note that it is not the best at everything — Anthropic's Fable 5.1 still leads in front-end design… Fable's analysis was superior for code review." — Theo (t3.gg), "So I've Been Using GPT-6 Astra…"

6 · Reasoning & Science: The Cleanest Split

This is where the two models divide most cleanly — and it goes both ways.

BenchmarkGPT-6 AstraClaude Fable 5.1Winner
FrontierMath Tier 4 v297.6%87.8%Astra
GPQA Diamond96.0%93.7%Astra
Terminal-Bench Science 0.164.6%52.6%Astra
Humanity's Last Exam (w/ tools)57.2%65.0%Fable
AA Intelligence Index v4.35353tie

Astra owns graduate-level math and physical science. It saturates FrontierMath Tier 4 (97.6% vs 87.8%), and its launch included Lean-verified proofs of new results on prime gaps. But the caveat from MindStudio is worth keeping: on a harder Epoch AI set of 68 unsolved Erdős problems, Astra solved only 2 in its official run (5 with repeated attempts costing $220,000+ in compute). Saturating one benchmark isn't solving math.

Fable 5.1 owns broad, hard reasoning. It takes Humanity's Last Exam with tools 65.0% to 57.2% — an eight-point gap, "not close" per DataCamp — and the AA evaluations with a knowledge or science flavour still break its way — Fable 5.1 leads the individual evals AA-Briefcase and SciCode, while Astra leads the two newest ones (Terminal-Bench v4.0, 59.1% vs 52.0%, and AutomationBench-AA, completing 41.6% of workflows without a guardrail violation vs 32.1%). On the composite the two are now level: Intelligence Index v4.3 scores both 53. Anthropic's own science story is qualitative: Mythos 5.1 designed protein binders with a ~50% hit rate (vs a 10–15% norm), trained a neural network that produced a 2–3 km resolution elevation map of a third of Venus, and wrote GPU kernels that sped up seven genomics models up to 2.5×.

ARC-AGI-3 caveat: Astra's headline 99.9% on ARC-AGI-3 came from OpenAI's stateful provider-adapter harness, which preserves reasoning state between actions. Under the standard provider-neutral harness, Astra scored 62.7% (a run that cost ~$26,000). Fable 5.1 has no published ARC-AGI-3 score. ARC Prize itself cautions the result is a step change in interactive reasoning, not proof of general intelligence.

7 · Computer Use & Professional Artifacts: Astra's Home Turf

If your agents need to click through real software, produce slides, or generate CAD output, this is Astra's dimension — mostly because Anthropic hasn't published comparable numbers.

BenchmarkGPT-6 AstraClaude Fable 5.1Notes
OSWorld 2.072.6% (offline set)77.9% partial / 41.7% strictDifferent task releases — not comparable
ScreenSpot-Pro92.7%87.3% (Fable 5)Grounding UI elements, no tools
AutomationBench41.4%31.4%Business workflows
BenchCAD95.9%84.3%Widest artifact gap either vendor publishes

Astra also completes OSWorld tasks ~47% faster than GPT-5.6 Sol, and its launch demos operated KiCad, Unity, FreeCAD and Blender directly. Anthropic's OSWorld 2.0 numbers (77.9% partial / 41.7% strict) come from the benchmark authors' August 2026 task release and are explicitly not comparable to OpenAI's offline-set 72.6% — Anthropic itself shows no competitor score on that row. For agents that operate real software, DataCamp's verdict: Astra is the stronger pick.

8 · Safety, Cybersecurity & Alignment

This is the dimension where the two companies made opposite product decisions.

MetricGPT-6 AstraClaude Fable 5.1Winner
ExploitBench100%70.0%Astra
FrontierCyber (Irregular lab)86 / 226not runAstra
Computer-use safety ↓2.4%9.5%Astra
Cyber capability tier"Critical" — gatedVuln discovery onlydifferent tradeoffs

Astra is the first OpenAI model rated "Critical" under the Preparedness Framework — it can find unknown vulnerabilities and develop exploits without step-by-step human guidance. That capability is gated behind the Daybreak program, and the shipping model refuses exploit-creation work; safety checks can even pause unrelated tasks. Independent lab Irregular reported Astra solving 86 of 226 FrontierCyber challenges vs 34 for GPT-5.6 Sol, including zero-day findings.

Fable 5.1 sits a tier lower by design. Anthropic loosened it enough to discover software vulnerabilities (a defensive win — cyber safeguards now block 60% fewer false positives), but exploit generation and binary-based vulnerability scanning still route to Opus models. The same underlying model ships as Mythos 5.1 with lighter safeguards, restricted to vetted organizations.

On alignment, Astra posts the better headline numbers (2.4% vs 9.5% on computer-use safety; 0% honeypot cheating vs Sol's 48.2%), but OpenAI itself disclosed a regression: Astra's written reasoning is harder to monitor, and UK AISI found it could evade monitoring under adversarial prompting. OpenAI says it will withhold scaling until it regains confidence in monitoring. Anthropic's Fable 5.1, meanwhile, is praised for staying "readable over long, multi-step tasks" (Jane Street).

9 · Pricing Deep Dive: Same Sticker, Very Different Bills

Both models list at $10/$50 per 1M tokens. The rate card differs in exactly two places — and both favor Fable 5.1. But measured per task, the picture flips completely.

The rate card

RateGPT-6 AstraClaude Fable 5.1
Input / Output per 1M$10.00 / $50.00$10.00 / $50.00
Cached input read per 1M$1.00$0.25 (4× cheaper)
Cache write per 1M$12.50$12.50
Batch discount50%50%
Above 272K input tokens$20 in / $2 cached / $75 outNo surcharge

What identical workloads cost (list-rate arithmetic)

WorkloadGPT-6 AstraClaude Fable 5.1Difference
Balanced assistant (1M in / 250K out)$22.50$22.50$0
Retrieval, sub-threshold (10M in / 1M out)$150$150$0
Retrieval, over-threshold (10M in / 1M out)$275$150Fable — 83% cheaper
Cache-heavy loop (100K prefix, 1,000 reads)$201$126Fable — 59% cheaper

What each task actually costs (measured)

EffortGPT-6 AstraClaude Fable 5.1
max53 pts / $3.2653 pts / $7.63
low$0.82—
max (Coding Agent Index)62 pts / $7.0962 pts / ~40% more
The trap in this comparison: identical list prices make the two models look interchangeable, and the only visible rate-card gap (cache reads) favors Anthropic 4×. Then you measure what each actually spends to finish a task, and Astra comes in at roughly 40% of Fable 5.1's cost on AA's Intelligence Index — $3.26 vs $7.63 at max effort, for an identical score of 53. Fable 5.1 isn't punished by its rate card; it's punished by its output volume (AA measured 78k output tokens per index task at max effort vs 27k for Astra — about three times as many). You're paying ~2.3× for the same index score, though Fable 5.1 still leads several of the individual evaluations inside it — whether that's worth it is your call.

10 · Real-World Head-to-Heads (People Who Actually Compared Them)

Beyond the vendor tables, several independent testers ran both models side by side. Here's what they found.

DataCamp · Physics simulation test

One hard task, identical conditions: Astra 5.0 vs Fable 4.3

A from-scratch physics simulation (rotating hexagon, counter-rotating obstacle, mass-scaled balls) in a single HTML file. Astra passed runnability and scored 5/5 on physics, stability and visual quality — in 6 turns and 9 tool calls, adding an editorial layout, pause/restart controls and a telemetry panel. Fable scored 4/4/5/4 in 2 turns and a single tool call, adding a diagnostic HUD but letting balls stick to walls. "Astra read the prompt as a brief to interpret; Fable read it as a spec to satisfy."

YouTube · KingBench 3

"GPT-6 Astra (Fully Tested & Side by Side with Fable 5.1): ONE is a CLEAR WINNER!"

Eight app-building tests on KingBench 3: Fable 5.1 scored 74/80 (92.5%, first on the leaderboard) vs Astra's 72/80 (90%, third). Astra won the folding table, panda SVG and 3D wristwatch tests; Fable won the elevator simulation, contact lens case and archery game; they tied on the other two. Fable also clearly outperformed on an Obsidian clone with working image generation. Token cost: Astra ~$198 vs Fable ~$113.

Reddit · r/better_claw

GPT-6 Astra vs Fable 5.1 vs Sonnet 5 on real agent work (day one)

Astra led every agent benchmark tested (Terminal-Bench 4.0, Terminal-Bench Science, Agents' Last Exam), with real gaps — 12 points on TB Science, ~2 points on TB 4.0. But per task, Astra cost ~$1.10 vs Fable's ~$0.65: "Astra is verbose, especially with reasoning tokens… the thinking tax is real." Fable's $0.25 cache reads make repeat-context work materially cheaper.

YouTube · No Hype Assessment

"I Tested GPT-6 Astra vs Fable 5.1 (No Hype Assessment)"

Terminal-Bench 4.0 at max effort: Astra 56.7% vs Fable 55.8% — "basically the same." Cost: $10.35 for Astra vs $19.50 for Fable. Verdict: "GPT-6 is a big leap forward on the OpenAI side; Fable 5.1 is just giving us more of what we already like."

Theo (t3.gg) · Podcast

"So I've Been Using GPT-6 Astra…"

Power-user verdict: Astra is the best model he has ever used — but not the best at everything. Fable 5.1 still leads in front-end design, and Fable's analysis was superior for code review. A split verdict echoed by MindStudio: "a real jump in raw coding capability, roughly on par with Fable 5, but with front-end polish and 'would I actually merge this' confidence still favoring Fable 5.1."

dev.to · Three-way

GPT-6 Astra vs Fable 5.1 vs Gemini 3.8 Flash: The Ultimate Comparison

Confirms the split on vendor benchmarks: Astra wins FrontierMath (97.6 vs 87.8), ExploitBench (100 vs 70), OSWorld (72.6), AutomationBench (41.4 vs 31.4) and BenchCAD (95.9 vs 84.3). Its AA numbers (66 vs 61 on the Intelligence Index, ~70 vs ~67 on the Coding Agent Index) come from Index v4.1 and are now superseded — AA's v4.3 run puts both models on 53 (Intelligence) and 62 (Coding Agent). Astra also outputs faster (~87 tok/s vs ~67–69).

11 · Verdict: Pick on Workload Shape, Not the Tables

✅ Choose GPT-6 Astra if…

  • Your agents operate real software. OSWorld 2.0 72.6%, ScreenSpot-Pro 92.7% — the only one of the two with a published record on desktop grounding.
  • You need polished professional artifacts. AutomationBench 41.4% vs 31.4%; BenchCAD 95.9% vs 84.3% — the widest artifact gap either vendor publishes.
  • The work is math or physical science. FrontierMath T4 97.6% vs 87.8%; GPQA 96.0% vs 93.7%; Terminal-Bench Science 64.6% vs 52.6%.
  • You do defensive security work. 100% ExploitBench, 86/226 FrontierCyber — if you can live with Daybreak gating.
  • Cost per task matters more than the top score. $3.26 vs $7.63 on AA's Index v4.3 at max effort — for the same 53 points; $0.82 at low effort, and $7.09 on the Coding Agent Index against ~40% more.
  • You already run Codex. Context notes across windows are a Codex feature.

✅ Choose Claude Fable 5.1 if…

  • You want the strongest hard-reasoning rows. HLE with tools 65.0% vs 57.2%, and AA-Briefcase and SciCode still favour Fable — even though the composite index is now a 53–53 tie.
  • Your bill is cache reads, not output. $0.25 vs $1.00 cached reads; the cache-heavy loop above costs $126 vs $201.
  • Your requests are large. No long-context surcharge — 10M-token retrieval stays $150 where Astra climbs to $275.
  • You work in Claude Code. AA's Coding Agent Index is now level (62 apiece) rather than Fable-led, though Claude Code remains the harness with the longest track record.
  • Front-end design and code review are your bottleneck. Theo, MindStudio and the KingBench 3 test all favor Fable's polish and analysis.
  • You want readable long-horizon reasoning. Jane Street: "Fable 5.1 remains readable over long, multi-step tasks."

The one-line takeaway: GPT-6 Astra is the computer operator — faster, cheaper per task, and dominant on math, artifacts and cyber, but with a monitorability regression and gated capabilities. Claude Fable 5.1 is the reasoning workhorse — level with Astra at the top of AA's Intelligence Index (53 apiece, at roughly 2.3× the per-task cost), cheaper cache and long-context economics, and the qualitative edge in front-end polish and code review. If your work is multi-step, tool-heavy and long-horizon, test both on your workload: the rate card says they're identical, the measurements say they're not (DataCamp; MindStudio; Artificial Analysis).

12 · Sources & Data Notes

Vendor benchmark figures are as published September 1–5, 2026. Artificial Analysis figures use Intelligence Index v4.3 and Coding Agent Index v1.4 (published September 7–9, 2026) — both AA composites are now dead heats, at 53 and 62 respectively. Earlier drafts of this comparison cited Index v4.1 (Fable 5.1 66, Astra 61) and v4.2 (57 vs 55), so every AA number here is date-stamped. Headline scores are vendor-reported at maximum effort unless noted; independent checks (Artificial Analysis, ARC Prize, Irregular lab, UK AISI) and the caveats attached to each number are called out inline. Where vendors ran different task releases or harnesses (OSWorld 2.0, ARC-AGI-3, Coding Agent Index), the comparison is flagged rather than treated as apples-to-apples.

Run both models on your own workload

Test GPT-6 Astra and Claude Fable 5.1 side by side in isolated execution environments on CodingFleet — before you commit a single production dollar.

Try it on CodingFleet →