model profile
Rank · overall#17
Score74.7
Benchmark scores
9 benchmarksSWE-bench Verified82%% Resolved · 2026-07-09
Terminal-Bench 2.176.2%Accuracy · Jul 9, 2026 · ± 1.2
Chatbot Arena (Coding)1521Coding Elo
Artificial Analysis50.6Intelligence Index
AA Coding Index71.3Coding Index
Arena Agent0.67%Net Improvement · ± 0.89
WebDev Arena1539WebDev Elo · ± 12
SWE-bench Pro61.5%% Resolved
MCP Atlas88.1Score
Across harnesses
1 agentsPer-benchmark score · composite per harness. Tap a row to open that board.
Scraped and aggregated from public leaderboards · 2026-07-22