FrontierMath (Tier 4)
How it works
Epoch AI's FrontierMath, Tier 4 (v2) — exceptionally difficult, research-level math problems, evaluated on a private held-out set so answers can't leak into training data. Verified numerically.
How we ranked it
composite 29 / 100Every benchmark is scored on four weighted criteria — how directly it measures a real coding agent, how much of the stack it covers, how real its tasks are, and how open and reproducible it is. Those blend into the composite that ranks it #15 of 19.
Scores a real (harness × model) pair — the full agent — not just the bare model.
Data-driven: this benchmark exposes 39 models.
Executable, real-world coding tasks with hard pass/fail — not human preference or aggregate scores.
Open data with per-run receipts you can audit.
Our review
Best-in-class for contamination resistance and openness of data and method. But it measures deep mathematical reasoning, not the coding-agent stack — a model can top FrontierMath and still be a mediocre coding agent — so it sits mid-table for our purposes.
Top on FrontierMath (Tier 4)
39 models · AccuracySource: epoch.ai/benchmarks/frontiermath-tier-4-v2 ↗ · aggregated 2026-07-19