FrontierCode v1.1 Main Leaderboard

Production-code quality benchmark by Cognition. 100 private tasks scored by maintainer-authored rubrics — correctness, tests, scope, style, and maintainability. Models ranked by published Main scores.

Last updated: July 24, 2026 · 🆕 Claude Opus 5 (53.4%) added · Source: Cognition FrontierCode · SWE-bench Pro → · Terminal-Bench →

#
Model
Prov
Main
Harness
License
$/1M Out
Source
Src

FrontierCode v1.1 Main Score Comparison

Published Main scores on a consistent 0–100 scale. Models without a score are omitted.

About FrontierCode v1.1: Cognition's production-code benchmark. Tests whether AI-generated pull requests are mergeable — scored for correctness, tests, scope, style, and maintainability through maintainer-authored rubrics. 100 private Main tasks (150 in Extended). Each row combines a model with an agent harness at its best-performing effort.
🆕 Claude Opus 5 (53.4%): Anthropic's latest flagship scores 53.4% Main and 63.6% Extended on FrontierCode v1.1 per its system card (Jul 24, 2026). Nearly ties Fable 5 at the top.
🆕 Claude Fable 5 (53.5%): Leads the board by 0.1 pts over Opus 5. Anthropic system card reports 53.5% Main at xhigh effort with claude-code harness. Diamond: 29.3%.
🆕 GPT-5.6 Sol (47.5%): OpenAI's flagship scores 47.5% Main on the official Cognition leaderboard. Extended: 60.6%. GPT-5.6 Terra and Luna have no published Main scores (only Extended: 55.8% and 55.1%). Kimi K3 also has no published FrontierCode Main score.
Harness note: Scores combine different agent harnesses (claude-code, codex, SWE-agent). Results from different harnesses are directional rather than directly comparable.
📊 See also: SWE-bench Pro Leaderboard — real-world bug fixing  |  Terminal-Bench 2.1 Leaderboard — CLI coding  |  DeepSWE v1.1 Leaderboard — long-horizon engineering

Test these models on production code

20+ LLMs on CodingFleet. Run them on your own PRs. Benchmarks are a compass, not a map.

🚀 Try on CodingFleet →