ixio
← Top Benchmark
benchmark · agent × model

NL2Repo

Pass rate·2 harness × model pairs · 4 runs·homepage ↗
Rank · benchmarks#4 / 19
Score69

How it works

Our self-run reproduction of NL2RepoBench: mini-SWE-agent is given only a natural-language spec and must build a complete, installable Python library from scratch, which is then graded against that library's real upstream test suite (fixed-denominator pass rate). Run on ConTree sandboxes and cross-verified on Tenki.

Metric
Pass rate
Pairs
2
Runs
4
Type
Agent × model

How we ranked it

composite 69 / 100

Every benchmark is scored on four weighted criteria — how directly it measures a real coding agent, how much of the stack it covers, how real its tasks are, and how open and reproducible it is. Those blend into the composite that ranks it #4 of 19.

Agent-nativeweight 30%
100

Scores a real (harness × model) pair — the full agent — not just the bare model.

Stack coverageweight 30%
1

Data-driven: this benchmark exposes 2 harness × model pairs · 4 runs.

Task realismweight 25%
95

Executable, real-world coding tasks with hard pass/fail — not human preference or aggregate scores.

Open & reproducibleweight 15%
100

Open data with per-run receipts you can audit.

Our review

Right on theme — 0-to-1 repo generation, executable grading, fully open runs with two-backend verification. It genuinely discriminates (models that ace one library flop on another). Still early: a small task set and open-model cohort so far, which caps coverage until we widen it.

Top on NL2Repo

2 models · Pass rate

Source: github.com/opencolin/ensemble/tree/main/runner/nl2repo.py · aggregated 2026-07-19