Terminal-Bench 2.1
How it works
The agent is dropped into a real terminal and asked to complete end-to-end tasks — installing things, wiring up tools, fixing a broken environment — graded on whether the task actually got done. Version 2.1, scored per (agent, model) with an effort setting.
How we ranked it
composite 71 / 100Every benchmark is scored on four weighted criteria — how directly it measures a real coding agent, how much of the stack it covers, how real its tasks are, and how open and reproducible it is. Those blend into the composite that ranks it #3 of 19.
Scores a real (harness × model) pair — the full agent — not just the bare model.
Data-driven: this benchmark exposes 17 harness × model pairs.
Executable, real-world coding tasks with hard pass/fail — not human preference or aggregate scores.
Open data with per-run receipts you can audit.
Our review
Strongly agent-native and highly realistic: it's a real shell doing real work, pass/fail. Reputable and well-run. Younger and smaller than SWE-bench, with a narrower task set, but it captures the 'operate a computer' side of agentic coding that issue-resolution benchmarks miss.
Top on Terminal-Bench 2.1
12 models · AccuracySource: www.tbench.ai/leaderboard/terminal-bench/2.1 ↗ · aggregated 2026-07-19