NL2Repo
How it works
Our self-run reproduction of NL2RepoBench: mini-SWE-agent is given only a natural-language spec and must build a complete, installable Python library from scratch, which is then graded against that library's real upstream test suite (fixed-denominator pass rate). Run on ConTree sandboxes and cross-verified on Tenki.
How we ranked it
composite 69 / 100Every benchmark is scored on four weighted criteria — how directly it measures a real coding agent, how much of the stack it covers, how real its tasks are, and how open and reproducible it is. Those blend into the composite that ranks it #4 of 19.
Scores a real (harness × model) pair — the full agent — not just the bare model.
Data-driven: this benchmark exposes 2 harness × model pairs · 4 runs.
Executable, real-world coding tasks with hard pass/fail — not human preference or aggregate scores.
Open data with per-run receipts you can audit.
Our review
Right on theme — 0-to-1 repo generation, executable grading, fully open runs with two-backend verification. It genuinely discriminates (models that ace one library flop on another). Still early: a small task set and open-model cohort so far, which caps coverage until we widen it.
Top on NL2Repo
2 models · Pass rateSource: github.com/opencolin/ensemble/tree/main/runner/nl2repo.py ↗ · aggregated 2026-07-19