SWE-bench Pro
How it works
Scale's SEAL 'SWE-bench Pro': harder, more recent real-repository issues resolved against a private, held-out test set to resist contamination. Scale runs its own agent and reports % resolved.
How we ranked it
composite 57 / 100Every benchmark is scored on four weighted criteria — how directly it measures a real coding agent, how much of the stack it covers, how real its tasks are, and how open and reproducible it is. Those blend into the composite that ranks it #7 of 19.
Scores a real (harness × model) pair — the full agent — not just the bare model.
Data-driven: this benchmark exposes 25 models.
Executable, real-world coding tasks with hard pass/fail — not human preference or aggregate scores.
Open data with per-run receipts you can audit.
Our review
A stronger, harder cousin of SWE-bench Verified — same executable-realism strengths, better contamination hygiene thanks to the held-out set. It's Scale-run (semi-open) and model-level, which is why it lands just behind the open, agent-native boards.
Top on SWE-bench Pro
25 models · % ResolvedSource: labs.scale.com/leaderboard/swe_bench_pro_public ↗ · aggregated 2026-07-19