model profile
Rank · overall#24
Score72.9
Benchmark scores
9 benchmarksSWE-bench Verified79.8%% Resolved · 2026-02-19
Terminal-Bench 2.165.8%Accuracy · May 5, 2026 · ± 1.7
Chatbot Arena (Coding)1531Coding Elo
ARC-AGI77.1%Score
Stratix Cup15.7Tournament score
FrontierMath (Tier 4)26.8%Accuracy · ± 7
SWE-bench Pro46.1%% Resolved
SkillsBench60.8Score
BrowseComp85.9%Accuracy
Across harnesses
2 agentsPer-benchmark score · composite per harness. Tap a row to open that board.
Scraped and aggregated from public leaderboards · 2026-07-22