FrontierBench v0.1 Leaderboard

Professional computer-work benchmark from the team behind Terminal-Bench and Harbor. 74 tasks across 7 domains — software engineering, ML, security, data science, and more. Models ranked by task completion rate.

Last updated: September 3, 2026 · 🆕 Qwen3.8-Max-0902 (29.0%) added — now the top non-Anthropic result · Source: FrontierBench · GitHub · TB 3.0 → · TB 2.1 →

#
Model
Prov
Score
Harness
License
$/1M Out
Src

FrontierBench v0.1 Score Comparison

Published task completion rates. All models run on their respective agent harnesses.

About FrontierBench v0.1: A continuously maintained professional computer-work benchmark from the team behind Terminal-Bench and Harbor. Built by Stanford and the Laude Institute. 74 tasks across 7 domains: software engineering, machine learning, security, data science, system administration, web development, and file operations. Tasks are designed to test frontier autonomous knowledge work — long-horizon, multi-step professional tasks that require real computer skills.
Successor to Terminal-Bench: FrontierBench is the next-generation benchmark from the same team, designed to measure and evolve with the frontier of agent work. It covers a broader range of professional computer tasks beyond just terminal environments.
Harness note: Scores combine different agent harnesses (codex, claude-code, mini-SWE-agent, Grok Build). Results from different harnesses are directional rather than directly comparable. Claude Opus 5's 42.7% figure is from the official FrontierBench public leaderboard run (mini-SWE-agent harness, refreshed Aug 13, 2026; an earlier capture showed 43.5% ± 1.7%); Anthropic's own launch materials cited a close but distinct self-reported figure (43.3% at max effort, 44.4% at xhigh), within the published margin of error. Grok 4.6 (26.5% ± 1.5%, Grok Build harness) joined on Aug 12, 2026. GLM-5.3 (28.3% ± 2.0%, mini-SWE-agent, max effort) posted the strongest open-weight result on launch (Aug 14, 2026); Gemini 3.7 Flash (14.9% ± 1.8%, Gemini CLI, high effort) is the top Flash-tier entry (LLM-Stats). 🆕 Renamed: in mid-August 2026 the benchmark rebranded to Terminal-Bench 3.0 — frontierbench.ai now serves the TB 3.0 board with identical tasks (74 across 7 domains). The benchmark is currently in v0.1 and will be updated quarterly.
📊 See also: Terminal-Bench 2.1 Leaderboard — CLI agentic coding  |  SWE-bench Pro Leaderboard — real-world bug fixing  |  DeepSWE v1.1 Leaderboard — long-horizon engineering

Test these models on professional computer work

20+ LLMs on CodingFleet. Run them on your own workflows. Benchmarks are a compass, not a map.

🚀 Try on CodingFleet →