model profile
Rank · overall#28
Score70.8
Benchmark scores
4 benchmarksSWE-bench Verified76.2%% Resolved · 2026-02-17
Chatbot Arena (Coding)1518Coding Elo
ARC-AGI65.1%Score
FrontierMath (Tier 4)17.1%Accuracy · ± 5.9
Across harnesses
0 agentsThis model appears only in model-level benchmarks — no harness ran it directly.
Scraped and aggregated from public leaderboards · 2026-07-23