FrontierBench v0.1 Leaderboard
Professional computer-work benchmark from the team behind Terminal-Bench and Harbor. 74 tasks across 7 domains — software engineering, ML, security, data science, and more. Models ranked by task completion rate.
Launched: July 2026 · 🆕 Claude Opus 5 leads at 43.5% · Source: FrontierBench · GitHub · Terminal-Bench 2.1 →
FrontierBench v0.1 Score Comparison
Published task completion rates. All models run on their respective agent harnesses.
Successor to Terminal-Bench: FrontierBench is the next-generation benchmark from the same team, designed to measure and evolve with the frontier of agent work. It covers a broader range of professional computer tasks beyond just terminal environments.
Harness note: Scores combine different agent harnesses (codex, claude-code, mini-SWE-agent). Results from different harnesses are directional rather than directly comparable. Claude Opus 5's 43.5% figure is from the official FrontierBench v0.1 public leaderboard run (mini-SWE-agent harness); Anthropic's own launch materials cited a close but distinct self-reported figure (43.3% at max effort, 44.4% at xhigh), within the published margin of error. The benchmark is currently in v0.1 and will be updated quarterly.
Test these models on professional computer work
20+ LLMs on CodingFleet. Run them on your own workflows. Benchmarks are a compass, not a map.
🚀 Try on CodingFleet →