Tutorials, deep dives and product notes — built for developers.
Interactive FrontierCode v1.1 Main leaderboard with Claude Fable 5 at 53.5%, Claude Opus 5 at 53.4%, and 32 models ranked by production-code pull request quality. Updated August 7, 2026.
Interactive DeepSWE v1.1 leaderboard with Claude Opus 5 at 74.0%, GPT-5.6 Sol at 72.7%, and 20 models ranked by long-horizon software engineering ability. Updated August 7, 2026.
Interactive Terminal-Bench 2.1 leaderboard updated with Claude Mythos 5 at 88.0%. 45+ models ranked by CLI coding ability. Updated August 6, 2026.
Interactive SWE-bench Pro leaderboard updated with Claude Mythos 5 at 80.3% and Sakana Fugu-Ultra at 73.7%. 40+ models ranked by real coding ability. Updated August 6, 2026.
What SWE-bench Pro actually measures, how it works (1,865 tasks, 41 repos, 123 languages), why OpenAI abandoned SWE-bench Verified, the DeepSWE audit that found 32% verifier errors, and how to use coding benchmarks correctly. The definitive explainer.