SWE-bench Verified
How it works
The human-validated 'Verified' split of SWE-bench: an agent is given a real GitHub repository and a real issue, and must produce a patch that makes the project's hidden test suite pass. We take the best publicly submitted result per model.
How we ranked it
composite 58 / 100Every benchmark is scored on four weighted criteria — how directly it measures a real coding agent, how much of the stack it covers, how real its tasks are, and how open and reproducible it is. Those blend into the composite that ranks it #5 of 19.
Scores a real (harness × model) pair — the full agent — not just the bare model.
Data-driven: this benchmark exposes 60 models.
Executable, real-world coding tasks with hard pass/fail — not human preference or aggregate scores.
Open data with per-run receipts you can audit.
Our review
The gold standard for task realism — real issues, real repos, hard pass/fail from the project's own tests. It rates lower than the agent boards only because we read it model-level (best submission per model), so it doesn't isolate the harness, and because its popularity makes contamination and submission-tuning a live concern. Still one of the most trustworthy signals here.
Top on SWE-bench Verified
60 models · % ResolvedSource: openlm.ai/swe-bench ↗ · aggregated 2026-07-19