ixio runs
How it works
Our own runs. mini-SWE-agent drives a given model inside an isolated sandbox to solve a task, then a hidden test suite grades partial credit. It exists to fill the (harness × open-model) cells no public benchmark reports — e.g. Claude Code driving an open model.
How we ranked it
composite 73 / 100Every benchmark is scored on four weighted criteria — how directly it measures a real coding agent, how much of the stack it covers, how real its tasks are, and how open and reproducible it is. Those blend into the composite that ranks it #2 of 19.
Scores a real (harness × model) pair — the full agent — not just the bare model.
Data-driven: this benchmark exposes 26 harness × model pairs.
Executable, real-world coding tasks with hard pass/fail — not human preference or aggregate scores.
Open data with per-run receipts you can audit.
Our review
Fully open and agent-native, with per-run receipts we publish. The catch: our current task set is largely saturated — most capable models score ~100% — so it no longer discriminates. For that reason we keep it on the record here but exclude it from the composite scores. Harder tasks are in progress.
Top on ixio runs
18 models · Pass rateSource: github.com/opencolin/ensemble/tree/main/runner ↗ · aggregated 2026-07-19