Terminal-Bench 2.1 Leaderboard
Command-line agentic coding. Real shell tasks — install packages, debug builds, configure servers, manage git repos. Models tracked with published scores from official sources.
Last updated: July 10, 2026 · 🆕 GPT-5.6 Sol leads at 88.8% (91.9% Ultra) · Sources: Anthropic · tbench.ai · DeepSeek · SWE-bench Pro →
🆕 GPT-5.6 Sol leads at 88.8% (2.1) — OpenAI also reports 91.9% with its four-agent Ultra configuration. Terra: 87.4%; Luna: 84.7%. These are OpenAI-reported results.
🆕 Claude Sonnet 5 debuts at 80.4% (2.1) — 97% of Opus 4.8 at 60% of the price. +13.4 pts over Sonnet 4.6. Anthropic System Card. $3/$15 per 1M.
GPT-5.5 remains a useful historical reference at 83.4% (Codex CLI, 2.1). GPT-5.3 Codex (77.3%, 2.0) still beats GPT-5.4 on terminal — Codex specialization holds.
Test these models on real terminal tasks
20+ LLMs on CodingFleet. Run them on your own CLI workflows. Benchmarks tell you what won — your code tells you what works.
🚀 Try on CodingFleet →