Tutorials, deep dives and product notes — built for developers.
Claude Sonnet 5 ($2/$10, June 30 — now permanent) beats GPT-5.5 ($5/$30, April 23) on every directly comparable benchmark: +4.6 SWE-bench Pro, +2.2 Terminal-Bench 2.1, +5.2 HLE with tools. At 60% cheaper input and 67% cheaper output. Full benchmark comparison.
Claude Sonnet 5 vs Sonnet 4.6: every benchmark, every gain. +13.4 Terminal-Bench 2.1, +10.6 HLE tools, +5.1 SWE-bench Pro, +223 GDPval (beats Opus 4.8). Same $3/$15 list price. Tokenizer caveat explained. Full comparison with bar charts, radar, and gains chart — all sourced from Anthropic's Sonnet 5 System Card.
Claude Sonnet 5 (63.2% Pro, $15/1M) vs Opus 4.8 (69.2%, $25/1M). Sonnet 5 beats Opus on knowledge work (GDPval 1618 vs 1615), ties on HLE with tools (57.4% vs 57.9%), and delivers 93% of Opus capability at 60% of the price. Full benchmark comparison from Anthropic's Sonnet 5 System Card.
Terminal-Bench 2.1 leaderboard with DeepSeek V4.1 Flash atop the public board and newly tracked Ornith-1.5-397B at 86.1% on Terminus-2 as a separate vendor run. Vals AI and Cognition runs remain distinct. Updated October 9, 2026.
SWE-bench Pro: Anthropic-reported Opus 5.5 leads at 89.9%; newly tracked vendor results include Ornith-1.5-397B (65.1%) and Qwen3.8-Omni-Flash (63.3%), kept distinct from Scale standardized results. Updated October 9, 2026.