ixio
ranking the rankings · benchmarks

Which benchmark should you trust?

19benchmarks/2026-07-21updated

Every leaderboard here pulls from public benchmarks — but they're not equal. We rank them by one thing: how directly they measure the coding-agent stack — a real harness driving a real model on real tasks. Four criteria, weighted:

Agent-native30%

Scores a real (harness × model) pair — the full agent — not just the bare model.

Stack coverage30%

How many distinct agent × model combinations it actually runs.

Task realism25%

Executable, real-world coding tasks with hard pass/fail — not human preference or aggregate scores.

Open & reproducible15%

Open data with per-run receipts you can audit.

Benchmark leaderboard

19 benchmarks · score 0–100

Tap a benchmark for its method, our review, and how we scored it.

CodingAgentBench tops it — the broadest open matrix of coding agents × open-weight models on executable tasks, with public per-run receipts. Model-level benchmarks never drive a harness, so they rank lower — but an executable, traced contest where models write real code holds up far better than human-preference Elo or aggregate indexes, which sink furthest.

Infrastructure benchmarks

source: ComputeSDK

These measure the layer under the agent — the sandboxes it executes in, the storage it persists to, the browsers it drives — so they rank providers, not models, and sit outside the scoring above. Run daily in ComputeSDK's open CI; we read the raw results straight from the repo.

ComputeSDK Sandboxesinfra

Sandbox cold-start: median time-to-interactive (sequential / burst / staggered) + $/hr

20 providers · dailyboard →src ↗
ComputeSDK Storageinfra

Object-storage throughput and latency on real 1–16 MB transfers

6 providers · dailyboard →src ↗
ComputeSDK Browsersinfra

Remote-browser session round-trip (create · connect · navigate) + actions/s

6 providers · dailyboard →src ↗