ixio
← Top Benchmark
benchmark · model-level

FrontierMath (Tier 4)

Accuracy·39 models·homepage ↗
Rank · benchmarks#15 / 19
Score29

How it works

Epoch AI's FrontierMath, Tier 4 (v2) — exceptionally difficult, research-level math problems, evaluated on a private held-out set so answers can't leak into training data. Verified numerically.

Metric
Accuracy
Models
39
Type
Model-level

How we ranked it

composite 29 / 100

Every benchmark is scored on four weighted criteria — how directly it measures a real coding agent, how much of the stack it covers, how real its tasks are, and how open and reproducible it is. Those blend into the composite that ranks it #15 of 19.

Agent-nativeweight 30%
0

Scores a real (harness × model) pair — the full agent — not just the bare model.

Stack coverageweight 30%
0

Data-driven: this benchmark exposes 39 models.

Task realismweight 25%
60

Executable, real-world coding tasks with hard pass/fail — not human preference or aggregate scores.

Open & reproducibleweight 15%
92

Open data with per-run receipts you can audit.

Our review

Best-in-class for contamination resistance and openness of data and method. But it measures deep mathematical reasoning, not the coding-agent stack — a model can top FrontierMath and still be a mediocre coding agent — so it sits mid-table for our purposes.

Top on FrontierMath (Tier 4)

39 models · Accuracy

Source: epoch.ai/benchmarks/frontiermath-tier-4-v2 · aggregated 2026-07-19