model profile
Rank · overall#359
Score32.5
Benchmark scores
6 benchmarksSWE-bench Verified59.8%% Resolved · 2025-08-07
Chatbot Arena (Coding)1419Coding Elo
ARC-AGI4.4%Score
FrontierMath (Tier 4)12.2%Accuracy · ± 5.2
SWE-bench Multilingual39.7%% Resolved · 2026-02-13
MathArena Apex1%Accuracy
Across harnesses
0 agentsThis model appears only in model-level benchmarks — no harness ran it directly.
Scraped and aggregated from public leaderboards · 2026-07-23