FrontierBench v0.1 Leaderboard

Professional computer-work benchmark from the team behind Terminal-Bench and Harbor. 74 tasks across 7 domains — software engineering, ML, security, data science, and more. Models ranked by task completion rate.

Launched: July 2026 · 🆕 GPT-5.6 Sol leads at 34.4% · Source: FrontierBench · GitHub · Terminal-Bench 2.1 →

#
Model
Prov
Score
Harness
License
$/1M Out
Src

FrontierBench v0.1 Score Comparison

Published task completion rates. All models run on their respective agent harnesses.

About FrontierBench v0.1: A continuously maintained professional computer-work benchmark from the team behind Terminal-Bench and Harbor. Built by Stanford and the Laude Institute. 74 tasks across 7 domains: software engineering, machine learning, security, data science, system administration, web development, and file operations. Tasks are designed to test frontier autonomous knowledge work — long-horizon, multi-step professional tasks that require real computer skills.
Successor to Terminal-Bench: FrontierBench is the next-generation benchmark from the same team, designed to measure and evolve with the frontier of agent work. It covers a broader range of professional computer tasks beyond just terminal environments.
Harness note: Scores combine different agent harnesses (codex, claude-code). Results from different harnesses are directional rather than directly comparable. The benchmark is currently in v0.1 and will be updated quarterly.
📊 See also: Terminal-Bench 2.1 Leaderboard — CLI agentic coding  |  SWE-bench Pro Leaderboard — real-world bug fixing  |  DeepSWE v1.1 Leaderboard — long-horizon engineering

Test these models on professional computer work

20+ LLMs on CodingFleet. Run them on your own workflows. Benchmarks are a compass, not a map.

🚀 Try on CodingFleet →