ixio
← Top Benchmark
benchmark · model-level

SWE-bench Multilingual

% Resolved·13 models·homepage ↗
Rank · benchmarks#6 / 19
Score58

How it works

The Multilingual split of the official SWE-bench: resolve real GitHub issues across nine programming languages, graded by each project's hidden tests. Best submission per model.

Metric
% Resolved
Models
13
Type
Model-level

How we ranked it

composite 58 / 100

Every benchmark is scored on four weighted criteria — how directly it measures a real coding agent, how much of the stack it covers, how real its tasks are, and how open and reproducible it is. Those blend into the composite that ranks it #6 of 19.

Agent-nativeweight 30%
65

Scores a real (harness × model) pair — the full agent — not just the bare model.

Stack coverageweight 30%
0

Data-driven: this benchmark exposes 13 models.

Task realismweight 25%
100

Executable, real-world coding tasks with hard pass/fail — not human preference or aggregate scores.

Open & reproducibleweight 15%
90

Open data with per-run receipts you can audit.

Our review

Real issues with real tests, broadened beyond Python — valuable coverage of languages most benchmarks ignore. Like the other SWE splits it's model-level here, so it scores on realism and openness rather than agent-nativeness.

Top on SWE-bench Multilingual

13 models · % Resolved

Source: raw.githubusercontent.com/swe-bench/swe-bench.github.io/master/data/leaderboards.json · aggregated 2026-07-19