FrontierCode v1.1 Main Leaderboard

Production-code quality benchmark by Cognition. 100 private tasks scored by maintainer-authored rubrics — correctness, tests, scope, style, and maintainability. Models ranked by published Main scores.

Last updated: August 14, 2026 · 🆕 Grok 4.6 (48.0%) and Gemini 3.7 Flash (43.6%) added · Source: Cognition FrontierCode · SWE-bench Pro → · Terminal-Bench →

#
Model
Prov
Main
Harness
License
$/1M Out
Source
Src

FrontierCode v1.1 Main Score Comparison

Published Main scores on a consistent 0–100 scale. Models without a score are omitted.

About FrontierCode v1.1: Cognition's production-code benchmark. Tests whether AI-generated pull requests are mergeable — scored for correctness, tests, scope, style, and maintainability through maintainer-authored rubrics. 100 private Main tasks (150 in Extended). Each row combines a model with an agent harness at its best-performing effort.
🆕 Claude Opus 5 (53.4%): Anthropic's latest flagship scores 53.4% Main and 63.6% Extended on FrontierCode v1.1 per its system card (Jul 24, 2026). Nearly ties Fable 5 at the top.
🆕 Claude Fable 5 (53.5%): Leads the board by 0.1 pts over Opus 5. Anthropic system card reports 53.5% Main at xhigh effort with claude-code harness. Diamond: 29.3%.
🆕 GPT-5.6 Sol (47.5%): OpenAI's flagship scores 47.5% Main on the official Cognition leaderboard. Extended: 60.6%.
🆕 Grok 4.6 (48.0%): Ranks #3 on the official Cognition leaderboard — ahead of GPT-5.6 Sol — at high effort ($2.88/run). Gemini 3.7 Flash (43.6%) is the strongest Flash-tier result yet, +9.2 pts over 3.6 Flash, at medium effort ($3.65/run). Both landed the week of Aug 12–13, 2026.
🆕 New entries (Aug 2026): Kimi K3 (44.2%), Grok 4.5 (42.4%), GPT-5.6 Terra (41.3%), GPT-5.6 Luna (39.8%), Gemini 3.6 Flash (34.4%), DeepSeek V4 Flash 0731 (18.8%), DeepSeek V4 Pro (17.6%), MiniMax M3 (14.7%), Inkling 0.99 (14.0%), Mistral 3.5 Medium (8.0%), and Qwen 3.7 Plus (10.2%) all now have published Main scores on the official Cognition FrontierCode leaderboard.
Harness note: Scores combine different agent harnesses (claude-code, codex, SWE-agent). Results from different harnesses are directional rather than directly comparable.
📊 See also: SWE-bench Pro Leaderboard — real-world bug fixing  |  Terminal-Bench 2.1 Leaderboard — CLI coding  |  DeepSWE v1.1 Leaderboard — long-horizon engineering

Test these models on production code

20+ LLMs on CodingFleet. Run them on your own PRs. Benchmarks are a compass, not a map.

🚀 Try on CodingFleet →