SWE-bench Multilingual
How it works
The Multilingual split of the official SWE-bench: resolve real GitHub issues across nine programming languages, graded by each project's hidden tests. Best submission per model.
How we ranked it
composite 58 / 100Every benchmark is scored on four weighted criteria — how directly it measures a real coding agent, how much of the stack it covers, how real its tasks are, and how open and reproducible it is. Those blend into the composite that ranks it #6 of 19.
Scores a real (harness × model) pair — the full agent — not just the bare model.
Data-driven: this benchmark exposes 13 models.
Executable, real-world coding tasks with hard pass/fail — not human preference or aggregate scores.
Open data with per-run receipts you can audit.
Our review
Real issues with real tests, broadened beyond Python — valuable coverage of languages most benchmarks ignore. Like the other SWE splits it's model-level here, so it scores on realism and openness rather than agent-nativeness.
Top on SWE-bench Multilingual
13 models · % ResolvedSource: raw.githubusercontent.com/swe-bench/swe-bench.github.io/master/data/leaderboards.json ↗ · aggregated 2026-07-19