DeepSWE v1.1 Leaderboard

Long-horizon software engineering. 113 original tasks across 91 repositories and 5 languages — written from scratch, not adapted from existing commits. Models ranked by published pass@1 scores on mini-swe-agent.

Last updated: July 25, 2026 · 🆕 Claude Opus 5 (74.0%) leads the board · Source: DeepSWE by Datacurve · SWE-bench Pro → · Terminal-Bench →

#
Model
Prov
Score
±CI
Cost
License
Effort
Src

DeepSWE v1.1 Score Comparison

Published pass@1 scores on a consistent 0–100 scale. All models run on mini-swe-agent.

About DeepSWE v1.1: Measures frontier coding agents on original, long-horizon software engineering tasks. Tasks are written from scratch (not adapted from existing commits), making the benchmark contamination-free. v1.1 updates execution and grading by scoring committed code in a clean, isolated environment. All models run on mini-swe-agent for consistency. 113 tasks, 91 repos, 5 languages.
🆕 Claude Opus 5 (74.0%): Anthropic's latest flagship, released July 24, 2026. Leads the DeepSWE v1.1 leaderboard at 74.0% ±4%, ahead of GPT-5.6 Sol (72.7%) and Claude Fable 5 (69.7%). Avg cost: $11.84/task.
🆕 GPT-5.6 Sol (72.7%): OpenAI's flagship holds second place with the best cost-efficiency at $8.39/task and only 61 agent steps — the most efficient among top models.
🆕 Kimi K3 (68.5%): Moonshot's 2.8T MoE model impresses at $4.65/task — the best value among frontier models on this benchmark.
Scores: All scores from the official DeepSWE leaderboard (deepswe.datacurve.ai), updated July 25, 2026. Effort setting in brackets. CI = 95% confidence interval.
📊 See also: SWE-bench Pro Leaderboard — real-world bug fixing  |  Terminal-Bench 2.1 Leaderboard — CLI coding  |  MCP Atlas Leaderboard — tool orchestration

Test these models on real engineering tasks

20+ LLMs on CodingFleet. Run them on your own codebases. Benchmarks are a compass, not a map.

🚀 Try on CodingFleet →