SWE-bench Pro Leaderboard

The hardest coding benchmark for AI. Real GitHub issues, multi-file diffs, production repositories — not memorized answers. Models ranked by published scores.

Last updated: August 21, 2026 · 🆕 Ling 3.0 Flash (56.6%), Qwen3.8-27B (61.7%) & Muse Glimmer (51.2%) added · Compare pricing → · Terminal-Bench →

#
Model
Prov
Pro
Verif
Size
License
$/1M Out
Released
Src

SWE-bench Pro Score Comparison

Published Pro scores on a consistent 0–100 scale. Models without a score are omitted.

About SWE-bench Pro: Tests whether an AI model can resolve real GitHub issues end-to-end. Unlike SWE-bench Verified (contaminated — see OpenAI's Feb 2026 withdrawal), Pro uses actively maintained repositories with no public ground-truth leakage.
🆕 GLM-5.3 & GLM-5.2: GLM-5.2 scored 62.1% on SWE-bench Pro. For GLM-5.3, Z.ai emphasized DeepSWE v1.1 (66.9%), SWE-Marathon v1.1 (42.5%), and Terminal-Bench 3.0 (28.3%), and has not published a separate Pro run.
🆕 Qwen3.8 Max: Qwen reports 67.7% on SWE-bench Pro — the strongest new entry since Opus 5, above GPT-5.6 Sol (64.6%), at $2/$6 per 1M tokens.
🆕 Qwen3.8-27B (61.7%): Alibaba's open-weights 27.8B dense (Apache 2.0, Aug 14, 2026) posts the best SWE-bench Pro score in its class — above Qwen3.7-Plus (57.6%) and Opus 4.6 Max (53.4%) — at $0.45/$3.20 per 1M tokens. Muse Glimmer (51.2%) is Meta's first open-weight Muse release (30B, Apache 2.0, Aug 10, 2026), Meta-reported.
Scores: Vendor-reported unless otherwise noted. “—” means not published. Cross-vendor harnesses can differ, so treat small gaps as directional.

Test these models on real code

20+ LLMs on CodingFleet. Side-by-side testing on your own repos. Benchmarks are a compass, not a map.

🚀 Try on CodingFleet →