Tutorials, deep dives and product notes — built for developers.
Across all six effort levels, GPT-6 Sol scores within a point of GPT-5.6 Sol and costs about half as much per task. Same-harness data on both, the full effort ladder, where coding improved, where knowledge work regressed, and the November price cliff.
OpenAI shipped GPT-6 Sol 90 minutes after Claude Opus 5.5, at half the token price. Artificial Analysis ran both on the same tests: Sol is cheaper to any index score up to about 44, then runs out of headroom. Every benchmark, both effort ladders and the real bill.
Opus 5.5 wins eight of the nine benchmarks Anthropic publishes against Fable 5.1 at 2.5x lower token price - but four of those margins sit inside the disclosed error bars. Includes the HAProxy same-task test, the 16-0 research-report result and the one-way thinking-block handoff.
Anthropic's Claude Opus 5.5 matches Fable 5.1 on most work at $4/$20 per million tokens: 66.4% on Terminal-Bench 4.0, 1,846 GDPval-AA Elo, and 40% lower cost per task than Opus 5. Full benchmark table with footnotes, real tester results, the four breaking API changes, and where GPT-6 Astra still wins.
Terminal-Bench 4.0 public board remains led by GPT-6 Astra (58.2%); Grok 4.7 is listed at 37.6% ±3.5 ($3.7k, 5.5B tokens). Separate, unranked runs: Anthropic Opus 5.5 66.4%, Cognition SWE-2 27.3%, and Artificial Analysis Codex runs GPT-6 Sol 44% / Luna 13%. Harness caveats included. Updated Sep 25, 2026.
Gemini 3.8 Flash review: DeepSWE v1.1 (73.7%), Terminal-Bench 2.1 (89.4%), full benchmark comparison vs Claude Opus 5 & GPT-5.6 Sol, real Antigravity builds with live links, and 3.8 Flash Cyber breakdown.
Comprehensive comparison of GLM 5.3 vs Claude Opus 5 with interactive radar capability chart, direct score bars, token pricing, and benchmark analysis. Updated August 19, 2026.
Interactive Terminal-Bench 3.0 leaderboard with Claude Opus 5 at 42.7%, GPT-5.6 Sol at 34.6%, GLM 5.3 at 28.3%, and all tracked frontier agent models ranked. Updated August 2026.
Claude Opus 5 vs Kimi K3: Opus 5 leads the independent BenchLM aggregate 85.88 to 79.98 and posts a 43.3–43.5% Frontier-Bench score that K3 has never even been tested on.
Interactive FrontierBench v0.1 leaderboard with Claude Opus 5 leading at 42.7%, GLM-5.3 at 28.3%, Gemini 3.7 Flash at 14.9%, and 12 models ranked by professional computer-work task completion. Renamed Terminal-Bench 3.0. Updated August 21, 2026.
Claude Opus 5 vs Claude Fable 5: Opus 5 beats Fable 5 on 7 of 12 benchmarks including Frontier-Bench (+9.6) and OSWorld 2.0 — at half the price ($25 vs $50/1M output). Fable 5 edges SWE-bench Pro by just 0.8 pts. Full comparison with radar charts, pricing, data retention, and verdict.
Claude Opus 5 vs GPT-5.6 Sol: Opus 5 leads 9 of 12 benchmarks including SWE-bench Pro (+14.6 pts) and ARC-AGI-3 (3.9× better). Sol counters with Terminal-Bench 2.1 (91.9% Ultra) and DeepSWE. Opus 5 costs 17% less on output ($25 vs $30/1M). Full comparison with radar charts, pricing, and verdict.