ixio
← Top Benchmark
benchmark · agent × model

CodingAgentBench

Pass rate·140 harness × model pairs · 3,465 runs·homepage ↗
Rank · benchmarks#1 / 19
Score98

How it works

Runs open coding agents — the CLIs and TUIs people actually use — across a matrix of open-weight models on a suite of coding tasks, and reports pass rate for each (harness × model) pair. Because it varies both the harness and the model, it's one of the few boards that measures the full agent rather than a bare LLM.

Metric
Pass rate
Pairs
140
Runs
3,465
Type
Agent × model

How we ranked it

composite 98 / 100

Every benchmark is scored on four weighted criteria — how directly it measures a real coding agent, how much of the stack it covers, how real its tasks are, and how open and reproducible it is. Those blend into the composite that ranks it #1 of 19.

Agent-nativeweight 30%
100

Scores a real (harness × model) pair — the full agent — not just the bare model.

Stack coverageweight 30%
100

Data-driven: this benchmark exposes 140 harness × model pairs · 3,465 runs.

Task realismweight 25%
95

Executable, real-world coding tasks with hard pass/fail — not human preference or aggregate scores.

Open & reproducibleweight 15%
95

Open data with per-run receipts you can audit.

Our review

Our #1 benchmark. It's the most agent-native board we track and by far the broadest matrix of agent × model combinations, all on executable tasks with public results — exactly the thing this site exists to measure. The main limitation is scope: it leans toward open-weight models, so the frontier closed models are under-represented here and have to be triangulated from other boards.

Top on CodingAgentBench

14 models · Pass rate

Source: codingagentbench.com/leaderboard · aggregated 2026-07-19