Tutorials, deep dives and product notes — built for developers.
Terminal-Bench 4.0 public board remains led by GPT-6 Astra (58.2%); Grok 4.7 is listed at 37.6% ±3.5 ($3.7k, 5.5B tokens). Separate, unranked runs: Anthropic Opus 5.5 66.4%, Cognition SWE-2 27.3%, and Artificial Analysis Codex runs GPT-6 Sol 44% / Luna 13%. Harness caveats included. Updated Sep 25, 2026.
Cognition’s FrontierCode v1.1 Main leaderboard refreshed Sep 25: Claude Opus 5.5 leads at 54.6% (medium, $0.80/run). New entries include SWE-2 (50.0%), GPT-6 Sol (49.3%), Grok 4.7 (47.6%), and GPT-6 Luna (42.4%), with rollout costs and API rates distinguished.
DeepSWE v1.1 updated Sep 25: the Datacurve leaderboard remains led by Muse Spark 1.3 (75.4%). Adds separately labeled September runs: SWE-2 73.0%, Grok 4.7 73.0%, GPT-6 Sol 68.8%, and GPT-6 Luna 66.6%; harnesses and missing task costs are disclosed.
Terminal-Bench 2.1 refreshed Sep 25: DeepSeek V4.1 Flash leads the public tbench.ai board at 90.6%. Adds Cognition-reported SWE-2 at 92.8% as a separate, unranked Devin CLI run—not a public-board result. 50+ CLI agents tracked.