Arena Agent
How it works
Arena.ai's agentic coding evaluation, scored over real sessions as a 'Net Improvement' metric rather than a static pass rate.
How we ranked it
composite 43 / 100Every benchmark is scored on four weighted criteria — how directly it measures a real coding agent, how much of the stack it covers, how real its tasks are, and how open and reproducible it is. Those blend into the composite that ranks it #11 of 19.
Scores a real (harness × model) pair — the full agent — not just the bare model.
Data-driven: this benchmark exposes 29 models.
Executable, real-world coding tasks with hard pass/fail — not human preference or aggregate scores.
Open data with per-run receipts you can audit.
Our review
Genuinely session-based and agentic, which we like, and closer to real use than preference Elo. It's held back by a smaller sample and a proprietary methodology that's hard to audit, so openness and coverage are limited.
Top on Arena Agent
29 models · Net ImprovementSource: arena.ai/leaderboard/agent ↗ · aggregated 2026-07-19