DeepSWE v1.1 Leaderboard
Long-horizon software engineering. 113 original tasks across 91 repositories and 5 languages — written from scratch, not adapted from existing commits. Models ranked by published pass@1 scores on mini-swe-agent.
Last updated: July 25, 2026 · 🆕 Claude Opus 5 (74.0%) leads the board · Source: DeepSWE by Datacurve · SWE-bench Pro → · Terminal-Bench →
DeepSWE v1.1 Score Comparison
Published pass@1 scores on a consistent 0–100 scale. All models run on mini-swe-agent.
🆕 Claude Opus 5 (74.0%): Anthropic's latest flagship, released July 24, 2026. Leads the DeepSWE v1.1 leaderboard at 74.0% ±4%, ahead of GPT-5.6 Sol (72.7%) and Claude Fable 5 (69.7%). Avg cost: $11.84/task.
🆕 GPT-5.6 Sol (72.7%): OpenAI's flagship holds second place with the best cost-efficiency at $8.39/task and only 61 agent steps — the most efficient among top models.
🆕 Kimi K3 (68.5%): Moonshot's 2.8T MoE model impresses at $4.65/task — the best value among frontier models on this benchmark.
Scores: All scores from the official DeepSWE leaderboard (deepswe.datacurve.ai), updated July 25, 2026. Effort setting in brackets. CI = 95% confidence interval.
Test these models on real engineering tasks
20+ LLMs on CodingFleet. Run them on your own codebases. Benchmarks are a compass, not a map.
🚀 Try on CodingFleet →