DeepSWE v1.1 Leaderboard

Long-horizon software engineering. 113 original tasks across 91 repositories and 5 languages — written from scratch, not adapted from existing commits. Models ranked by published pass@1 scores on mini-swe-agent.

Last updated: September 11, 2026 · 🆕 DeepSeek V4.1 Flash opens at 74.2% — #2 on the board, ahead of Claude Opus 5 and GPT-6 Astra at $0.60/1M output · Sources: OpenAI · DeepSWE by Datacurve · SWE-bench Pro → · Terminal-Bench →

#
Model
Prov
Score
±CI
Cost
License
Effort
Src

DeepSWE v1.1 Score Comparison

Published pass@1 scores on a consistent 0–100 scale. All models run on mini-swe-agent.

About DeepSWE v1.1: Measures frontier coding agents on original, long-horizon software engineering tasks. Tasks are written from scratch (not adapted from existing commits), making the benchmark contamination-free. v1.1 updates execution and grading by scoring committed code in a clean, isolated environment. All models run on mini-swe-agent for consistency. 113 tasks, 91 repos, 5 languages.
🆕 Muse Spark 1.3 (75.4%): Meta's flagship released Sep 2, 2026. Self-reported 75.4% on DeepSWE v1.1 via Muse Code harness at max effort, leading overall coding benchmarks at $1.25/$4.25 per 1M tokens ($0.55/task).
🆕 GPT-6 Astra (74.1%): OpenAI's next-gen flagship released Sep 2026. Official vendor report scores 74.1% on DeepSWE v1.1 (+1.4 pts over GPT-5.6 Sol) with ~57% lower cost per task than Sol thanks to token efficiency despite $10/$50 rates.
🆕 Claude Fable 5.1 (67.4%): Anthropic's refreshed Mythos-class release scored 67.4% average over 5 trials on DeepSWE v1.1. 75% cheaper cache reads ($0.25/M) sharply reduce agent loop costs.
🆕 GLM-5.3 (69.0%): Z.ai's post-training flagship now scores 69.0% ±3% on the official DeepSWE board (refreshed Aug 20, 2026) — up from 66.9% at launch, tied with Kimi K3 for #4 and +25 pts over GLM-5.2 (44.0%). $1.40/$4.40 per 1M tokens. Open weights releasing in ~2 weeks.
🆕 Claude Opus 5 (74.0%): Anthropic's latest flagship, released July 24, 2026. Leads the DeepSWE v1.1 leaderboard at 74.0% ±4%, ahead of GPT-5.6 Sol (73.0%) and Claude Fable 5 (70.0%). Avg cost: $11.84/task.
🆕 New entries (Aug 2026): Grok 4.6 (67.0% ±2% at $5.50/task), Gemini 3.7 Flash (65.0% ±2% at $2.18/task), DeepSeek V4 Pro 0813 (63.0% ±6% at $0.24/task — updated Aug 20), and Muse Spark 1.2 (55.0% ±2%) joined the official leaderboard. Gemini 3.6 Flash (47.0%) and Gemini 3.5 Flash (36.0%) were re-run and updated. 🆕 Qwen3.8-27B (42.2%) is vendor-reported from Qwen's model card — the strongest open-weight result in its size class.
🆕 New entries (Aug 26, 2026): GLM-5.3-Flash (63.4%) and Qwen3.8-Flash-Next (58.7%) are vendor-reported from their launch model cards — both undercut the flagship tier on price (GLM-5.3-Flash $0.15/$0.03, Qwen3.8-Flash-Next $0.16/$0.47 per 1M tokens) while beating several flagship scores (GLM-5.3-Flash ahead of Opus 4.8's 59.0%; Qwen3.8-Flash-Next ahead of Qwen3.8 Max's 57.0%).
Scores: All scores from official vendor releases and the DeepSWE leaderboard (deepswe.datacurve.ai). Effort setting in brackets. CI = 95% confidence interval.
📊 See also: SWE-bench Pro Leaderboard — real-world bug fixing  |  Terminal-Bench 2.1 Leaderboard — CLI coding  |  MCP Atlas Leaderboard — tool orchestration

Test these models on real engineering tasks

20+ LLMs on CodingFleet. Run them on your own codebases. Benchmarks are a compass, not a map.

🚀 Try on CodingFleet →