Terminal-Bench 4.0 Leaderboard

The definitive test of autonomous agent capability in real terminal environments. 66 professional computer-work tasks across software engineering, machine learning, systems, and operations evaluated over 5 trials per task with an 8-hour agent timeout.

Updated September 2026 66 Curated Tasks Harbor Hub Verified Codex & Claude Code & Grok Build
58.2%
Top Resolution Rate
GPT-6 Astra (Codex max)
41.8%
Top Open-Weight
GLM-5.3 (Z.AI)
$0.3k
Lowest Run Cost
GPT-5.6 Luna ($6/1M API)
18
Tracked Configurations
Across 10 Frontier Models
# Model & Provider Agent Harness Effort Resolution Rate Eval Run Cost Tokens API Price (Out) Source / Verify
💡 Click any row to expand granular benchmark metadata & token breakdown

What Changed in Terminal-Bench 4.0?

Terminal-Bench (curated by Stanford, the Laude Institute, Harbor Framework, and open-source contributors) evaluates whether autonomous agents can complete complex terminal tasks end-to-end. Unlike SWE-bench which isolates git patches for individual Python issues, Terminal-Bench tasks require operating shells, installing toolchains, compiling multi-language dependencies, executing benchmarks, and resolving system faults.

1. Task Set Recalibration (66 Tasks)

Version 4.0 removes 8 tasks that became saturated, refusal-prone, or leaked public solutions. 20 tasks received revised environments and tighter validation verifiers, avoiding false-positive passes.

2. 8-Hour Compute Budget & Multi-Trial Rigor

Each score reflects 5 full independent trials per task with an 8-hour agent timeout. All results report run-to-run variance (± margin of error) to ensure reliable separation among frontier models.

3. The Efficiency Divide: Token Economics

While GPT-6 Astra and Claude Fable 5.1 are neck-and-neck at 58.2% vs 57.9%, Astra needed only 1.5B tokens ($3.3k eval cost) compared to 2.7B tokens ($6.2k) for Fable 5.1 and 6.5B tokens ($6.0k) for Opus 5.

4. Open-Weight Breakthrough: GLM-5.3

Z.AI's GLM-5.3 achieved 41.8% resolution rate inside Claude Code harness—surpassing GPT-5.6 Sol (37.3%) and Opus 4.8 (23.6%) while running at an accessible $4.40 per 1M output tokens.

Official Verification Sources & Links

Test Frontier CLI & Terminal Agents on Real Workloads

Compare GPT-6 Astra, Claude Fable 5.1, GLM-5.3 and Gemini in isolated sandbox environments on CodingFleet.

⚡ Launch CodingFleet Sandbox →