ixio
← Top Benchmark
benchmark · agent × model

Terminal-Bench 2.1

Accuracy·17 harness × model pairs·homepage ↗
Rank · benchmarks#3 / 19
Score71

How it works

The agent is dropped into a real terminal and asked to complete end-to-end tasks — installing things, wiring up tools, fixing a broken environment — graded on whether the task actually got done. Version 2.1, scored per (agent, model) with an effort setting.

Metric
Accuracy
Pairs
17
Type
Agent × model

How we ranked it

composite 71 / 100

Every benchmark is scored on four weighted criteria — how directly it measures a real coding agent, how much of the stack it covers, how real its tasks are, and how open and reproducible it is. Those blend into the composite that ranks it #3 of 19.

Agent-nativeweight 30%
100

Scores a real (harness × model) pair — the full agent — not just the bare model.

Stack coverageweight 30%
12

Data-driven: this benchmark exposes 17 harness × model pairs.

Task realismweight 25%
95

Executable, real-world coding tasks with hard pass/fail — not human preference or aggregate scores.

Open & reproducibleweight 15%
88

Open data with per-run receipts you can audit.

Our review

Strongly agent-native and highly realistic: it's a real shell doing real work, pass/fail. Reputable and well-run. Younger and smaller than SWE-bench, with a narrower task set, but it captures the 'operate a computer' side of agentic coding that issue-resolution benchmarks miss.

Top on Terminal-Bench 2.1

12 models · Accuracy

Source: www.tbench.ai/leaderboard/terminal-bench/2.1 · aggregated 2026-07-19