model profile
Rank · overall#164
Score49.6
Benchmark scores
3 benchmarksSWE-bench Verified70.1%% Resolved · 2025-08-05
Chatbot Arena (Coding)1465Coding Elo
FrontierMath (Tier 4)2.4%Accuracy · ± 2.4
Across harnesses
0 agentsThis model appears only in model-level benchmarks — no harness ran it directly.
Scraped and aggregated from public leaderboards · 2026-07-23