Terminal-Bench 2.1 Leaderboard

Command-line agentic coding. Real shell tasks — install packages, debug builds, configure servers, manage git repos. Models tracked with published scores from official sources.

Last updated: September 11, 2026 · 🆕 DeepSeek V4.1 Flash takes #1 at 90.6% (mini-SWE 90.3%, Claude Code 88.0%) at $0.15/$0.60 per 1M with MIT weights · Sources: Z.ai · Moonshot AI · tbench.ai (TB 4.0) · SWE-bench Pro →

🆕 Terminal-Bench 4.0 is here (Sep 1–2, 2026): the benchmark moved to a much harder major version. Official tbench.ai runs — Claude Fable 5.1 57.9% ± 3.8 (#1) · Opus 5 51.8% ± 3.4 · Fable 5 44.5% · GLM-5.3 41.8% · GPT-5.6 Sol 37.3% · Gemini 3.8 Flash 19.1% ± 3.4 (mini-SWE-agent, Sep 2). Anthropic's own TB 4.0 numbers: Mythos 5.1 60.9% · Fable 5.1 55.8% · Opus 5 52.3% · Fable 5 42.0%. TB 4.0 scores are NOT comparable with the TB 2.1 numbers above.
#
Model
Prov
Score
Ver
Size
License
$/1M Out
Harness
Src

Terminal-Bench Score Comparison

Bar length uses the full 0–100 scale. Hover a bar for its source harness.

About Terminal-Bench: Evaluates AI agents on real command-line tasks — package management, build systems, git, server config, file manipulation, shell scripting. 🟢 2.1 = official tbench.ai or Anthropic-verified. 🟡 2.0 = vendor-reported from model cards. Scores NOT directly comparable across versions — 2.1 is harder.
🆕 GLM-5.3 scores 88.2% (2.1) & 28.3% (3.0): Z.ai's post-training flagship released Aug 14, 2026. Climbs to 88.2% on TB 2.1 and makes a massive leap on TB 3.0 (from 4.6% to 28.3%). $1.40/$4.40 per 1M tokens; 1M context. Open weights releasing in ~2 weeks.
🆕 Kimi K3 scores 88.3% (2.1) with KimiCode at max reasoning — only 0.5 points behind GPT-5.6 Sol's 88.8% single-agent result. Moonshot's official launch table; $3/$15 per 1M tokens.
🆕 New entries (Aug 2026): Grok 4.6 scores 88.4% (2.1) per Artificial Analysis. DeepSeek V4 Pro 0813 (87.9%), Qwen3.8 Max (86.6%) and Gemini 3.7 Flash (85.8%) are vendor-reported from their GA launch tables; Muse Spark 1.2 (82.9%) uses Meta's Muse Code harness.
🆕 New entries (Aug 21, 2026): Qwen3.8-27B (73.0%) — Alibaba's open-weights 27.8B dense (Apache 2.0), the strongest TB 2.1 score in its class — and Meta's open-weight Muse Glimmer (51.7%, per Qwen's cross-model comparison table) join the board.
🆕 New entries (Aug 26, 2026): GLM-5.3-Flash (84.3%) — Z.ai's 320B/18B MIT-licensed efficiency model — lands within 0.7 pts of Claude Opus 4.8 (85.0) at flash-tier pricing ($0.15/$0.03 per 1M). Qwen3.8-Flash-Next (125B/6B) has no published TB 2.1 score yet but reports 62.5% SWE-bench Pro and 58.7% DeepSWE 1.1.
Harness warning: Results combine vendor and leaderboard harnesses. Compare directionally unless the same evaluator and scaffold were used.
📊 See also: SWE-bench Pro Leaderboard — real-world bug fixing  |  AI Pricing Calculator — compare costs
🆕 Successor benchmark: FrontierBench v0.1 Leaderboard — the next-generation professional computer-work benchmark from the same team behind Terminal-Bench. Broader scope: 74 tasks across 7 domains (software engineering, ML, security, data science, and more).

Test these models on real terminal tasks

20+ LLMs on CodingFleet. Run them on your own CLI workflows. Benchmarks tell you what won — your code tells you what works.

🚀 Try on CodingFleet →