CodingAgentBench
How it works
Runs open coding agents — the CLIs and TUIs people actually use — across a matrix of open-weight models on a suite of coding tasks, and reports pass rate for each (harness × model) pair. Because it varies both the harness and the model, it's one of the few boards that measures the full agent rather than a bare LLM.
How we ranked it
composite 98 / 100Every benchmark is scored on four weighted criteria — how directly it measures a real coding agent, how much of the stack it covers, how real its tasks are, and how open and reproducible it is. Those blend into the composite that ranks it #1 of 19.
Scores a real (harness × model) pair — the full agent — not just the bare model.
Data-driven: this benchmark exposes 140 harness × model pairs · 3,465 runs.
Executable, real-world coding tasks with hard pass/fail — not human preference or aggregate scores.
Open data with per-run receipts you can audit.
Our review
Our #1 benchmark. It's the most agent-native board we track and by far the broadest matrix of agent × model combinations, all on executable tasks with public results — exactly the thing this site exists to measure. The main limitation is scope: it leans toward open-weight models, so the frontier closed models are under-represented here and have to be triangulated from other boards.
Top on CodingAgentBench
14 models · Pass rateSource: codingagentbench.com/leaderboard ↗ · aggregated 2026-07-19