Which benchmark should you trust?
Every leaderboard here pulls from public benchmarks — but they're not equal. We rank them by one thing: how directly they measure the coding-agent stack — a real harness driving a real model on real tasks. Four criteria, weighted:
Scores a real (harness × model) pair — the full agent — not just the bare model.
How many distinct agent × model combinations it actually runs.
Executable, real-world coding tasks with hard pass/fail — not human preference or aggregate scores.
Open data with per-run receipts you can audit.
Benchmark leaderboard
19 benchmarks · score 0–100Tap a benchmark for its method, our review, and how we scored it.
CodingAgentBench tops it — the broadest open matrix of coding agents × open-weight models on executable tasks, with public per-run receipts. Model-level benchmarks never drive a harness, so they rank lower — but an executable, traced contest where models write real code holds up far better than human-preference Elo or aggregate indexes, which sink furthest.
Infrastructure benchmarks
source: ComputeSDK ↗These measure the layer under the agent — the sandboxes it executes in, the storage it persists to, the browsers it drives — so they rank providers, not models, and sit outside the scoring above. Run daily in ComputeSDK's open CI; we read the raw results straight from the repo.
Sandbox cold-start: median time-to-interactive (sequential / burst / staggered) + $/hr
Object-storage throughput and latency on real 1–16 MB transfers