Tutorials, deep dives and product notes — built for developers.
Terminal-Bench 4.0 checked Oct 6, 2026: Claude Opus 5.5 (64.8%) and Sonnet 5.5 (61.8%) now top the public tbench.ai board, with GPT-6 Astra and GPT-6.1 Sol tied at 58.2%. Vendor runs and Artificial Analysis results are reported separately, with harness, tokens and cost per row.
Interactive Terminal-Bench 3.0 leaderboard with Claude Opus 5 at 42.7%, GPT-5.6 Sol at 34.6%, GLM 5.3 at 28.3%, and all tracked frontier agent models ranked. Updated August 2026.
Claude Fable 5 leads every benchmark (80.3% Pro, 88.0% Terminal-Bench, ~87% Multi). Now the undisputed #1 for Go coding across all workflows. Updated June 9, 2026.
Claude Fable 5 leads every benchmark (80.3% Pro, 88.0% Terminal-Bench, ~87% Multi). Now the undisputed #1 for all Rust workflows. Updated June 9, 2026.
Terminal-Bench 2.1: DeepSeek V4.1 Flash remains the public-board leader at 90.6%. Adds separate Vals AI snapshot runs: Opus 5.5 87.6% and Sonnet 5.5 83.1%, plus Cognition's unranked SWE-2 provider score.