model profile
Rank · overall#374
Score31.8
Benchmark scores
3 benchmarksSWE-bench Verified45%% Resolved · 2025-04-16
FrontierMath (Tier 4)4.9%Accuracy · ± 3.4
BrowseComp51.5%Accuracy
Across harnesses
0 agentsThis model appears only in model-level benchmarks — no harness ran it directly.
Scraped and aggregated from public leaderboards · 2026-07-23