ixio
← Top Benchmark
benchmark · model-level

SWE-bench Pro

% Resolved·25 models·homepage ↗
Rank · benchmarks#7 / 19
Score57

How it works

Scale's SEAL 'SWE-bench Pro': harder, more recent real-repository issues resolved against a private, held-out test set to resist contamination. Scale runs its own agent and reports % resolved.

Metric
% Resolved
Models
25
Type
Model-level

How we ranked it

composite 57 / 100

Every benchmark is scored on four weighted criteria — how directly it measures a real coding agent, how much of the stack it covers, how real its tasks are, and how open and reproducible it is. Those blend into the composite that ranks it #7 of 19.

Agent-nativeweight 30%
65

Scores a real (harness × model) pair — the full agent — not just the bare model.

Stack coverageweight 30%
0

Data-driven: this benchmark exposes 25 models.

Task realismweight 25%
100

Executable, real-world coding tasks with hard pass/fail — not human preference or aggregate scores.

Open & reproducibleweight 15%
80

Open data with per-run receipts you can audit.

Our review

A stronger, harder cousin of SWE-bench Verified — same executable-realism strengths, better contamination hygiene thanks to the held-out set. It's Scale-run (semi-open) and model-level, which is why it lands just behind the open, agent-native boards.

Top on SWE-bench Pro

25 models · % Resolved

Source: labs.scale.com/leaderboard/swe_bench_pro_public · aggregated 2026-07-19