Tutorials, deep dives and product notes — built for developers.
Interactive Terminal-Bench 3.0 leaderboard with Claude Opus 5 at 42.7%, GPT-5.6 Sol at 34.6%, GLM 5.3 at 28.3%, and all tracked frontier agent models ranked. Updated August 2026.
Terminal-Bench 2.1: DeepSeek V4.1 Flash remains the public-board leader at 90.6%. Adds separate Vals AI snapshot runs: Opus 5.5 87.6% and Sonnet 5.5 83.1%, plus Cognition's unranked SWE-2 provider score.