ixio
← Top Benchmark
benchmark · model-level

MathArena Apex

Accuracy·46 models·homepage ↗
Rank · benchmarks#17 / 19
Score27

How it works

MathArena Apex (ETH Zürich / INSAIT) evaluates models on the hardest recent competition-math problems, chosen to be uncontaminated at release, and reports accuracy.

Metric
Accuracy
Models
46
Type
Model-level

How we ranked it

composite 27 / 100

Every benchmark is scored on four weighted criteria — how directly it measures a real coding agent, how much of the stack it covers, how real its tasks are, and how open and reproducible it is. Those blend into the composite that ranks it #17 of 19.

Agent-nativeweight 30%
0

Scores a real (harness × model) pair — the full agent — not just the bare model.

Stack coverageweight 30%
0

Data-driven: this benchmark exposes 46 models.

Task realismweight 25%
55

Executable, real-world coding tasks with hard pass/fail — not human preference or aggregate scores.

Open & reproducibleweight 15%
90

Open data with per-run receipts you can audit.

Our review

Rigorous and refreshingly uncontaminated, with open methodology. But like FrontierMath it's pure reasoning, not coding or agentic work, so it's here for completeness and sits low on the coding-agent axis.

Top on MathArena Apex

46 models · Accuracy

Source: matharena.ai/?comp=apex--apex_2025 · aggregated 2026-07-19