Artificial Analysis
How it works
Artificial Analysis runs a standard battery of evals (reasoning, knowledge, math, etc.) and blends them into a single 'Intelligence Index' per model. A general-capability proxy, not coding-specific.
How we ranked it
composite 26 / 100Every benchmark is scored on four weighted criteria — how directly it measures a real coding agent, how much of the stack it covers, how real its tasks are, and how open and reproducible it is. Those blend into the composite that ranks it #17 of 17.
Scores a real (harness × model) pair — the full agent — not just the bare model.
Data-driven: this benchmark exposes 22 models.
Executable, real-world coding tasks with hard pass/fail — not human preference or aggregate scores.
Open data with per-run receipts you can audit.
Our review
A convenient one-number summary of raw model capability, but it's model-level, not coding-focused, and an aggregate-of-aggregates whose methodology is only semi-open. We show it for orientation; it barely touches the agentic-coding axis this site cares about.
Top on Artificial Analysis
22 models · Intelligence IndexSource: artificialanalysis.ai/models ↗ · aggregated 2026-09-06