Terminal-Bench 4.0 Leaderboard
The definitive test of autonomous agent capability in real terminal environments. 66 professional computer-work tasks across software engineering, machine learning, systems, and operations evaluated over 5 trials per task with an 8-hour agent timeout.
| # | Model & Provider | Agent Harness | Effort | Resolution Rate | Eval Run Cost | Tokens | API Price (Out) | Source / Verify |
|---|
What Changed in Terminal-Bench 4.0?
Terminal-Bench (curated by Stanford, the Laude Institute, Harbor Framework, and open-source contributors) evaluates whether autonomous agents can complete complex terminal tasks end-to-end. Unlike SWE-bench which isolates git patches for individual Python issues, Terminal-Bench tasks require operating shells, installing toolchains, compiling multi-language dependencies, executing benchmarks, and resolving system faults.
1. Task Set Recalibration (66 Tasks)
Version 4.0 removes 8 tasks that became saturated, refusal-prone, or leaked public solutions. 20 tasks received revised environments and tighter validation verifiers, avoiding false-positive passes.
2. 8-Hour Compute Budget & Multi-Trial Rigor
Each score reflects 5 full independent trials per task with an 8-hour agent timeout. All results report run-to-run variance (± margin of error) to ensure reliable separation among frontier models.
3. The Efficiency Divide: Token Economics
While GPT-6 Astra and Claude Fable 5.1 are neck-and-neck at 58.2% vs 57.9%, Astra needed only 1.5B tokens ($3.3k eval cost) compared to 2.7B tokens ($6.2k) for Fable 5.1 and 6.5B tokens ($6.0k) for Opus 5.
4. Open-Weight Breakthrough: GLM-5.3
Z.AI's GLM-5.3 achieved 41.8% resolution rate inside Claude Code harness—surpassing GPT-5.6 Sol (37.3%) and Opus 4.8 (23.6%) while running at an accessible $4.40 per 1M output tokens.
Official Verification Sources & Links
- Harbor Hub Dataset & Leaderboard: hub.harborframework.com/datasets/terminal-bench
- Terminal-Bench Official Portal: tbench.ai/news/terminal-bench-4-0
- Snorkel AI Benchmark Analysis: snorkel.ai/leaderboard/terminal-bench-4-0
- BenchLM Agentic Snapshot: benchlm.ai/benchmarks/terminal-bench-4
- Artificial Analysis Terminal-Bench v4.0: artificialanalysis.ai/evaluations/terminalbench-v4-0
- GitHub Repository: github.com/harbor-framework/terminal-bench
Test Frontier CLI & Terminal Agents on Real Workloads
Compare GPT-6 Astra, Claude Fable 5.1, GLM-5.3 and Gemini in isolated sandbox environments on CodingFleet.
⚡ Launch CodingFleet Sandbox →