model profile
Rank · overall#13
Score77.7
Benchmark scores
12 benchmarksSWE-bench Verified82%% Resolved · 2026-04-16
Terminal-Bench 2.168.9%Accuracy · May 1, 2026 · ± 1.4
Chatbot Arena (Coding)1559Coding Elo
ARC-AGI75.8%Score
Arena Agent7.94%Net Improvement · ± 1.24
WebDev Arena1559WebDev Elo · ± 7
Stratix Cup22.7Tournament score
FrontierMath (Tier 4)31.7%Accuracy · ± 7.4
MCP Atlas79.1Score
SkillsBench61.2Score
MathArena Apex40.6%Accuracy
BrowseComp79.3%Accuracy
Across harnesses
2 agentsPer-benchmark score · composite per harness. Tap a row to open that board.
Scraped and aggregated from public leaderboards · 2026-07-22