MathArena Apex
How it works
MathArena Apex (ETH Zürich / INSAIT) evaluates models on the hardest recent competition-math problems, chosen to be uncontaminated at release, and reports accuracy.
How we ranked it
composite 27 / 100Every benchmark is scored on four weighted criteria — how directly it measures a real coding agent, how much of the stack it covers, how real its tasks are, and how open and reproducible it is. Those blend into the composite that ranks it #17 of 19.
Scores a real (harness × model) pair — the full agent — not just the bare model.
Data-driven: this benchmark exposes 46 models.
Executable, real-world coding tasks with hard pass/fail — not human preference or aggregate scores.
Open data with per-run receipts you can audit.
Our review
Rigorous and refreshingly uncontaminated, with open methodology. But like FrontierMath it's pure reasoning, not coding or agentic work, so it's here for completeness and sits low on the coding-agent axis.
Top on MathArena Apex
46 models · AccuracySource: matharena.ai/?comp=apex--apex_2025 ↗ · aggregated 2026-07-19