FrontierCode v1.1 Main Leaderboard
Production-code quality benchmark by Cognition. 100 private tasks scored by maintainer-authored rubrics — correctness, tests, scope, style, and maintainability. Models ranked by published Main scores.
Last updated: July 24, 2026 · 🆕 Claude Opus 5 (53.4%) added · Source: Cognition FrontierCode · SWE-bench Pro → · Terminal-Bench →
FrontierCode v1.1 Main Score Comparison
Published Main scores on a consistent 0–100 scale. Models without a score are omitted.
🆕 Claude Opus 5 (53.4%): Anthropic's latest flagship scores 53.4% Main and 63.6% Extended on FrontierCode v1.1 per its system card (Jul 24, 2026). Nearly ties Fable 5 at the top.
🆕 Claude Fable 5 (53.5%): Leads the board by 0.1 pts over Opus 5. Anthropic system card reports 53.5% Main at xhigh effort with claude-code harness. Diamond: 29.3%.
🆕 GPT-5.6 Sol (47.5%): OpenAI's flagship scores 47.5% Main on the official Cognition leaderboard. Extended: 60.6%. GPT-5.6 Terra and Luna have no published Main scores (only Extended: 55.8% and 55.1%). Kimi K3 also has no published FrontierCode Main score.
Harness note: Scores combine different agent harnesses (claude-code, codex, SWE-agent). Results from different harnesses are directional rather than directly comparable.
Test these models on production code
20+ LLMs on CodingFleet. Run them on your own PRs. Benchmarks are a compass, not a map.
🚀 Try on CodingFleet →