ixio
← Top Benchmark
benchmark · model-level

SWE-bench Verified

% Resolved·60 models·homepage ↗
Rank · benchmarks#5 / 19
Score58

How it works

The human-validated 'Verified' split of SWE-bench: an agent is given a real GitHub repository and a real issue, and must produce a patch that makes the project's hidden test suite pass. We take the best publicly submitted result per model.

Metric
% Resolved
Models
60
Type
Model-level

How we ranked it

composite 58 / 100

Every benchmark is scored on four weighted criteria — how directly it measures a real coding agent, how much of the stack it covers, how real its tasks are, and how open and reproducible it is. Those blend into the composite that ranks it #5 of 19.

Agent-nativeweight 30%
65

Scores a real (harness × model) pair — the full agent — not just the bare model.

Stack coverageweight 30%
0

Data-driven: this benchmark exposes 60 models.

Task realismweight 25%
100

Executable, real-world coding tasks with hard pass/fail — not human preference or aggregate scores.

Open & reproducibleweight 15%
90

Open data with per-run receipts you can audit.

Our review

The gold standard for task realism — real issues, real repos, hard pass/fail from the project's own tests. It rates lower than the agent boards only because we read it model-level (best submission per model), so it doesn't isolate the harness, and because its popularity makes contamination and submission-tuning a live concern. Still one of the most trustworthy signals here.

Top on SWE-bench Verified

60 models · % Resolved

Source: openlm.ai/swe-bench · aggregated 2026-07-19