ixio
← Top Benchmark
benchmark · agent × model

ixio runs

Pass rate·26 harness × model pairs·homepage ↗
Rank · benchmarks#2 / 19
Score73

How it works

Our own runs. mini-SWE-agent drives a given model inside an isolated sandbox to solve a task, then a hidden test suite grades partial credit. It exists to fill the (harness × open-model) cells no public benchmark reports — e.g. Claude Code driving an open model.

Metric
Pass rate
Pairs
26
Type
Agent × model

How we ranked it

composite 73 / 100

Every benchmark is scored on four weighted criteria — how directly it measures a real coding agent, how much of the stack it covers, how real its tasks are, and how open and reproducible it is. Those blend into the composite that ranks it #2 of 19.

Agent-nativeweight 30%
100

Scores a real (harness × model) pair — the full agent — not just the bare model.

Stack coverageweight 30%
19

Data-driven: this benchmark exposes 26 harness × model pairs.

Task realismweight 25%
90

Executable, real-world coding tasks with hard pass/fail — not human preference or aggregate scores.

Open & reproducibleweight 15%
100

Open data with per-run receipts you can audit.

Our review

Fully open and agent-native, with per-run receipts we publish. The catch: our current task set is largely saturated — most capable models score ~100% — so it no longer discriminates. For that reason we keep it on the record here but exclude it from the composite scores. Harder tasks are in progress.

Top on ixio runs

18 models · Pass rate

Source: github.com/opencolin/ensemble/tree/main/runner · aggregated 2026-07-19