Terminal-Bench 4.0 Leaderboard

The definitive test of autonomous agent capability in real terminal environments. 66 professional computer-work tasks across software engineering, machine learning, systems, and operations evaluated over 5 trials per task with an 8-hour agent timeout.

Updated September 25, 2026 66 Curated Tasks Harbor Hub Verified Codex & Claude Code & Grok Build
58.2%
Top Public-Board Rate
GPT-6 Astra (Codex max)
41.8%
Top Open-Weight
GLM-5.3 (Z.AI)
$0.3k
Lowest Run Cost
GPT-5.6 Luna ($6/1M API)
24
Tracked Configurations
Across 20 Distinct Models
Source distinction (Sep 25): The public tbench.ai leaderboard still leads with GPT-6 Astra at 58.2%; it now lists Grok 4.7 at 37.6% ±3.5 ($3.7k run cost, 5.5B tokens). The table also includes four separately labeled runs outside the tbench.ai public board: Anthropic reports Opus 5.5 at 66.4% ±2.6; Artificial Analysis reports GPT-6 Sol at 44% and Luna at 13% (Codex); Cognition reports SWE-2 at 27.3%. These supplemental runs are unranked and should not be read as like-for-like public-board placements.
# Model & Provider Agent Harness Effort Resolution Rate Eval Run Cost Tokens API Price (Out) Source / Verify
💡 Click any row to expand granular benchmark metadata & token breakdown

What Changed in Terminal-Bench 4.0?

Terminal-Bench (curated by Stanford, the Laude Institute, Harbor Framework, and open-source contributors) evaluates whether autonomous agents can complete complex terminal tasks end-to-end. Unlike SWE-bench which isolates git patches for individual Python issues, Terminal-Bench tasks require operating shells, installing toolchains, compiling multi-language dependencies, executing benchmarks, and resolving system faults.

Reading the expanded table: The 58.2% top-stat remains the highest current public tbench.ai submission. Grok 4.7 is a public-board row (37.6% ±3.5, rank #6). Opus 5.5 (66.4%), SWE-2 (27.3%), and Artificial Analysis’s GPT-6 Sol (44%) / Luna (13%) runs are supplemental, unranked, and not controlled head-to-heads.

1. Task Set Recalibration (66 Tasks)

Version 4.0 removes 8 tasks that became saturated, refusal-prone, or leaked public solutions. 20 tasks received revised environments and tighter validation verifiers, avoiding false-positive passes.

2. 8-Hour Compute Budget & Multi-Trial Rigor

Each score reflects 5 full independent trials per task with an 8-hour agent timeout. All results report run-to-run variance (± margin of error) to ensure reliable separation among frontier models.

3. The Efficiency Divide: Token Economics

While GPT-6 Astra and Claude Fable 5.1 are neck-and-neck at 58.2% vs 57.9%, Astra needed only 1.5B tokens ($3.3k eval cost) compared to 2.7B tokens ($6.2k) for Fable 5.1 and 6.5B tokens ($6.0k) for Opus 5.

4. Open-Weight Breakthrough: GLM-5.3

Z.AI's GLM-5.3 achieved 41.8% resolution rate inside Claude Code harness—surpassing GPT-5.6 Sol (37.3%) and Opus 4.8 (23.6%) while running at an accessible $4.40 per 1M output tokens.

Official Verification Sources & Links

Test Frontier CLI & Terminal Agents on Real Workloads

Compare GPT-6 Astra, Claude Fable 5.1, GLM-5.3 and Gemini in isolated sandbox environments on CodingFleet.

⚡ Launch CodingFleet Sandbox →