Tutorials, deep dives and product notes — built for developers.
Terminal-Bench 4.0 checked Oct 6, 2026: Claude Opus 5.5 (64.8%) and Sonnet 5.5 (61.8%) now top the public tbench.ai board, with GPT-6 Astra and GPT-6.1 Sol tied at 58.2%. Vendor runs and Artificial Analysis results are reported separately, with harness, tokens and cost per row.
A data-backed comparison of Gemini 3.7 Flash and Claude Sonnet 5 across benchmarks, pricing, speed, context, and modalities — the speed king vs. the desktop-agent specialist.
Complete benchmark comparison of Gemini 3.6 Flash vs Claude Sonnet 5. Every score sourced from official model cards and system cards. Charts, radar plots, pricing analysis, and a clear verdict on which model to choose for your workload.
FrontierCode 1.1 Main: Opus 5.5 leads at 54.6%; new Cognition entries GPT-6.1 Sol (50.2%) and Claude Haiku 5.5 (46.4%) include effort, rollout spend, API rates and token details. Updated October 9, 2026.
DeepSWE v1.1: Datacurve's standardized leaderboard is kept separate from provider runs. Adds Google-reported Gemini 4 Argon at 77.9% and Ornith-1.5-397B at 56.0%; per-task costs are marked unavailable. Updated October 9, 2026.
GPT-5.6 Terra vs Claude Sonnet 5: nearly identical standard output pricing, 1M+ context, and a close SWE-bench Pro result—yet sharply different terminal scores, tooling, availability, and intro pricing.
Hy3 (295B MoE, Apache 2.0, $0.80/1M) vs Claude Sonnet 5 (proprietary, $10/1M). Sonnet leads every shared benchmark (+0.5 to +8.7 pts). But Hy3 ties on BrowseComp (84.2 vs 84.7), leads MCP Atlas (79.1%), costs 12.5x less. Open-weight agent vs proprietary coder — 5 charts, 10-point verdict.
Claude Fable 5 (80.3% SWE-bench Pro, $50/1M) vs Claude Sonnet 5 (63.2%, $15/1M). Fable 5 leads all 8 shared benchmarks by +8.2 pts avg — but Sonnet 5 delivers 79% of the capability at 30% of the price. Full comparison with 4 custom charts, pricing deep-dive, tokenizer analysis, and a 10-point verdict matrix.
Claude Sonnet 5 vs DeepSeek V4 Pro: Sonnet leads every coding benchmark (+7.8 Pro, +9.2 HLE tools). DeepSeek is #1 global on LiveCodeBench (93.5%), MIT open-weight, and 3.5× cheaper per task ($0.12 vs $0.42 at Sonnet's now-permanent $2/$10). Is 7.8 more Pro points worth 3.5× the cost?
Claude Sonnet 5 vs Qwen 3.7 Max: Sonnet leads coding (+2.6 Pro, +4.8 Verified). Qwen dominates math (92.4% GPQA), runs 35-hour autonomous agents, and is 2.7x cheaper ($3.75 vs $15 output). The coder vs the marathon runner — full comparison.
Claude Sonnet 5 vs Gemini 3.5 Flash: Speed vs Depth. Sonnet leads every coding benchmark (+8.1 Pro, +4.2 TB). Gemini leads MCP Atlas (83.6%), is 4x faster (289 tok/s), 2x cheaper. Coding specialist vs tool orchestration speed king — pick your weapon.
Claude Sonnet 5 vs GLM 5.2: near-ties on every benchmark (±0.6-2.7 pts). GLM 2.3× cheaper on output, MIT open-weight, self-hostable. Sonnet has OSWorld, BrowseComp, Anthropic safety ecosystem. Proprietary premium vs open-weight value — Sonnet's $2/$10 is now permanent.