ixio
← Top Benchmark
benchmark · model-level

ARC-AGI

Score·53 models·homepage ↗
Rank · benchmarks#14 / 19
Score30

How it works

ARC-AGI (the ARC Prize) poses novel grid-transformation puzzles designed to resist memorization: infer the rule from a few examples and apply it. Scored as percent solved.

Metric
Score
Models
53
Type
Model-level

How we ranked it

composite 30 / 100

Every benchmark is scored on four weighted criteria — how directly it measures a real coding agent, how much of the stack it covers, how real its tasks are, and how open and reproducible it is. Those blend into the composite that ranks it #14 of 19.

Agent-nativeweight 30%
0

Scores a real (harness × model) pair — the full agent — not just the bare model.

Stack coverageweight 30%
0

Data-driven: this benchmark exposes 53 models.

Task realismweight 25%
66

Executable, real-world coding tasks with hard pass/fail — not human preference or aggregate scores.

Open & reproducibleweight 15%
90

Open data with per-run receipts you can audit.

Our review

A superb, contamination-resistant test of fluid reasoning — but almost orthogonal to shipping code. It's not agentic and not a coding task, so it ranks low on our axis despite being an excellent benchmark of what it actually measures.

Top on ARC-AGI

53 models · Score

Source: openlm.ai/chatbot-arena · aggregated 2026-07-19