DeepSWE v1.1 Leaderboard
Long-horizon software engineering. 113 original tasks across 91 repositories and 5 languages — written from scratch, not adapted from existing commits. Models ranked by published pass@1 scores on mini-swe-agent.
Last updated: August 14, 2026 · 🆕 Grok 4.6 (67.0%), Gemini 3.7 Flash (65.0%), DeepSeek V4 Pro 0813 (63.0%), Muse Spark 1.2 (55.0%) added · Source: DeepSWE by Datacurve · SWE-bench Pro → · Terminal-Bench →
DeepSWE v1.1 Score Comparison
Published pass@1 scores on a consistent 0–100 scale. All models run on mini-swe-agent.
🆕 Claude Opus 5 (74.0%): Anthropic's latest flagship, released July 24, 2026. Leads the DeepSWE v1.1 leaderboard at 74.0% ±4%, ahead of GPT-5.6 Sol (72.7%) and Claude Fable 5 (69.7%). Avg cost: $11.84/task.
🆕 Qwen3.8 Max (57.0%): Alibaba's 2.4T flagship (GA Aug 3, 2026) debuts at 57.0% ±3% — the highest-scoring new entry, at just $3.73/task. 95k out tok, 111 agent steps.
🆕 DeepSeek V4 Flash (53.0%): Open-weight 284B MoE at $0.10/task — the cheapest model on the board by far. 53.0% ±4%, 108k out tok, 153 agent steps. MIT license. Released Jul 31, 2026.
🆕 New entries (Aug 13, 2026): Grok 4.6 (67.0% ±2% at $5.50/task), Gemini 3.7 Flash (65.0% ±2% at $2.18/task), DeepSeek V4 Pro 0813 (63.0% ±6% at $0.06/task — cheapest on the board), and Muse Spark 1.2 (55.0% ±2%) joined the official leaderboard. Gemini 3.6 Flash (47.0%) and Gemini 3.5 Flash (36.0%) were re-run and updated.
Scores: All scores from the official DeepSWE leaderboard (deepswe.datacurve.ai), updated August 13, 2026. Effort setting in brackets. CI = 95% confidence interval.
Test these models on real engineering tasks
20+ LLMs on CodingFleet. Run them on your own codebases. Benchmarks are a compass, not a map.
🚀 Try on CodingFleet →