model profile
Rank · overall#12
Score77.8
Benchmark scores
9 benchmarksSWE-bench Verified79.8%% Resolved · 2026-06-30
Terminal-Bench 2.174.6%Accuracy · Jul 9, 2026 · ± 1.6
Chatbot Arena (Coding)1520Coding Elo
Artificial Analysis53.4Intelligence Index
AA Coding Index71.5Coding Index
Arena Agent8.66%Net Improvement · ± 1.89
WebDev Arena1544WebDev Elo · ± 11
FrontierMath (Tier 4)29.3%Accuracy · ± 7.2
BrowseComp86.6%Accuracy
Across harnesses
1 agentsPer-benchmark score · composite per harness. Tap a row to open that board.
Scraped and aggregated from public leaderboards · 2026-07-22