model profile
Rank · overall#47
Score66.1
Benchmark scores
7 benchmarksSWE-bench Verified74.4%% Resolved · 2025-11-18
Terminal-Bench 2.173.9%Accuracy · May 1, 2026 · ± 1.3
Chatbot Arena (Coding)1501Coding Elo
ARC-AGI33.6%Score
WebDev Arena1439WebDev Elo · ± 7
SWE-bench Multilingual68.7%% Resolved · 2026-02-13
BrowseComp59.2%Accuracy
Across harnesses
2 agentsPer-benchmark score · composite per harness. Tap a row to open that board.
Scraped and aggregated from public leaderboards · 2026-07-23