ARC-AGI
How it works
ARC-AGI (the ARC Prize) poses novel grid-transformation puzzles designed to resist memorization: infer the rule from a few examples and apply it. Scored as percent solved.
How we ranked it
composite 30 / 100Every benchmark is scored on four weighted criteria — how directly it measures a real coding agent, how much of the stack it covers, how real its tasks are, and how open and reproducible it is. Those blend into the composite that ranks it #14 of 19.
Scores a real (harness × model) pair — the full agent — not just the bare model.
Data-driven: this benchmark exposes 53 models.
Executable, real-world coding tasks with hard pass/fail — not human preference or aggregate scores.
Open data with per-run receipts you can audit.
Our review
A superb, contamination-resistant test of fluid reasoning — but almost orthogonal to shipping code. It's not agentic and not a coding task, so it ranks low on our axis despite being an excellent benchmark of what it actually measures.
Top on ARC-AGI
53 models · ScoreSource: openlm.ai/chatbot-arena ↗ · aggregated 2026-07-19