Artificial Analysis
How it works
Artificial Analysis runs a standard battery of evals (reasoning, knowledge, math, etc.) and blends them into a single 'Intelligence Index' per model. A general-capability proxy, not coding-specific.
How we ranked it
composite 26 / 100Every benchmark is scored on four weighted criteria — how directly it measures a real coding agent, how much of the stack it covers, how real its tasks are, and how open and reproducible it is. Those blend into the composite that ranks it #19 of 19.
Scores a real (harness × model) pair — the full agent — not just the bare model.
Data-driven: this benchmark exposes 27 models.
Executable, real-world coding tasks with hard pass/fail — not human preference or aggregate scores.
Open data with per-run receipts you can audit.
Our review
A convenient one-number summary of raw model capability, but it's model-level, not coding-focused, and an aggregate-of-aggregates whose methodology is only semi-open. We show it for orientation; it barely touches the agentic-coding axis this site cares about.
Top on Artificial Analysis
27 models · Intelligence IndexSource: artificialanalysis.ai/models ↗ · aggregated 2026-07-19