DeepSWE v1.1 Leaderboard

Long-horizon software engineering. 113 original tasks across 91 repositories and 5 languages — written from scratch, not adapted from existing commits. Models ranked by published pass@1 scores on mini-swe-agent.

Last updated: August 7, 2026 · 🆕 Claude Opus 5 (74.0%) leads the board · Source: DeepSWE by Datacurve · SWE-bench Pro → · Terminal-Bench →

#
Model
Prov
Score
±CI
Cost
License
Effort
Src

DeepSWE v1.1 Score Comparison

Published pass@1 scores on a consistent 0–100 scale. All models run on mini-swe-agent.

About DeepSWE v1.1: Measures frontier coding agents on original, long-horizon software engineering tasks. Tasks are written from scratch (not adapted from existing commits), making the benchmark contamination-free. v1.1 updates execution and grading by scoring committed code in a clean, isolated environment. All models run on mini-swe-agent for consistency. 113 tasks, 91 repos, 5 languages.
🆕 Claude Opus 5 (74.0%): Anthropic's latest flagship, released July 24, 2026. Leads the DeepSWE v1.1 leaderboard at 74.0% ±4%, ahead of GPT-5.6 Sol (72.7%) and Claude Fable 5 (69.7%). Avg cost: $11.84/task.
🆕 Qwen3.8 Max (57.0%): Alibaba's 2.4T flagship (GA Aug 3, 2026) debuts at 57.0% ±3% — the highest-scoring new entry, at just $3.73/task. 95k out tok, 111 agent steps.
🆕 DeepSeek V4 Flash (53.0%): Open-weight 284B MoE at $0.10/task — the cheapest model on the board by far. 53.0% ±4%, 108k out tok, 153 agent steps. MIT license. Released Jul 31, 2026.
Scores: All scores from the official DeepSWE leaderboard (deepswe.datacurve.ai), updated August 6, 2026. Effort setting in brackets. CI = 95% confidence interval.
📊 See also: SWE-bench Pro Leaderboard — real-world bug fixing  |  Terminal-Bench 2.1 Leaderboard — CLI coding  |  MCP Atlas Leaderboard — tool orchestration

Test these models on real engineering tasks

20+ LLMs on CodingFleet. Run them on your own codebases. Benchmarks are a compass, not a map.

🚀 Try on CodingFleet →