DeepSWE v1.1 Leaderboard
Long-horizon software engineering. 113 original tasks across 91 repositories and 5 languages — written from scratch, not adapted from existing commits. Models ranked by published pass@1 scores on mini-swe-agent.
Last updated: August 28, 2026 · 🆕 GLM-5.3-Flash (63.4%) & Qwen3.8-Flash-Next (58.7%) added · Sources: Z.ai · DeepSWE by Datacurve · SWE-bench Pro → · Terminal-Bench →
DeepSWE v1.1 Score Comparison
Published pass@1 scores on a consistent 0–100 scale. All models run on mini-swe-agent.
🆕 GLM-5.3 (69.0%): Z.ai's post-training flagship now scores 69.0% ±3% on the official DeepSWE board (refreshed Aug 20, 2026) — up from 66.9% at launch, tied with Kimi K3 for #4 and +25 pts over GLM-5.2 (44.0%). $1.40/$4.40 per 1M tokens. Open weights releasing in ~2 weeks.
🆕 Claude Opus 5 (74.0%): Anthropic's latest flagship, released July 24, 2026. Leads the DeepSWE v1.1 leaderboard at 74.0% ±4%, ahead of GPT-5.6 Sol (73.0%) and Claude Fable 5 (70.0%). Avg cost: $11.84/task.
🆕 New entries (Aug 2026): Grok 4.6 (67.0% ±2% at $5.50/task), Gemini 3.7 Flash (65.0% ±2% at $2.18/task), DeepSeek V4 Pro 0813 (63.0% ±6% at $0.24/task — updated Aug 20), and Muse Spark 1.2 (55.0% ±2%) joined the official leaderboard. Gemini 3.6 Flash (47.0%) and Gemini 3.5 Flash (36.0%) were re-run and updated. 🆕 Qwen3.8-27B (42.2%) is vendor-reported from Qwen's model card — the strongest open-weight result in its size class.
🆕 New entries (Aug 26, 2026): GLM-5.3-Flash (63.4%) and Qwen3.8-Flash-Next (58.7%) are vendor-reported from their launch model cards — both undercut the flagship tier on price (GLM-5.3-Flash $0.15/$0.03, Qwen3.8-Flash-Next $0.16/$0.47 per 1M tokens) while beating several flagship scores (GLM-5.3-Flash ahead of Opus 4.8's 59.0%; Qwen3.8-Flash-Next ahead of Qwen3.8 Max's 57.0%).
Scores: All scores from official vendor releases and the DeepSWE leaderboard (deepswe.datacurve.ai). Effort setting in brackets. CI = 95% confidence interval.
Test these models on real engineering tasks
20+ LLMs on CodingFleet. Run them on your own codebases. Benchmarks are a compass, not a map.
🚀 Try on CodingFleet →