AA Coding Index
How it works
Artificial Analysis's coding-specific composite — a blend of coding evals (SWE-bench, Terminal-Bench, SciCode, LiveCodeBench and others) into one index per model.
How we ranked it
composite 29 / 100Every benchmark is scored on four weighted criteria — how directly it measures a real coding agent, how much of the stack it covers, how real its tasks are, and how open and reproducible it is. Those blend into the composite that ranks it #16 of 19.
Scores a real (harness × model) pair — the full agent — not just the bare model.
Data-driven: this benchmark exposes 27 models.
Executable, real-world coding tasks with hard pass/fail — not human preference or aggregate scores.
Open data with per-run receipts you can audit.
Our review
More on-target than the general Intelligence Index because it's coding-only, and a reasonable quick proxy. But it's a composite of other benchmarks (some of which we already track directly), so it double-counts, and it's still model-level rather than agent-level.
Top on AA Coding Index
27 models · Coding IndexSource: artificialanalysis.ai/models ↗ · aggregated 2026-07-19