ixio
← Top Benchmark
benchmark · model-level

Arena Agent

Net Improvement·29 models·homepage ↗
Rank · benchmarks#11 / 19
Score43

How it works

Arena.ai's agentic coding evaluation, scored over real sessions as a 'Net Improvement' metric rather than a static pass rate.

Metric
Net Improvement
Models
29
Type
Model-level

How we ranked it

composite 43 / 100

Every benchmark is scored on four weighted criteria — how directly it measures a real coding agent, how much of the stack it covers, how real its tasks are, and how open and reproducible it is. Those blend into the composite that ranks it #11 of 19.

Agent-nativeweight 30%
45

Scores a real (harness × model) pair — the full agent — not just the bare model.

Stack coverageweight 30%
0

Data-driven: this benchmark exposes 29 models.

Task realismweight 25%
80

Executable, real-world coding tasks with hard pass/fail — not human preference or aggregate scores.

Open & reproducibleweight 15%
60

Open data with per-run receipts you can audit.

Our review

Genuinely session-based and agentic, which we like, and closer to real use than preference Elo. It's held back by a smaller sample and a proprietary methodology that's hard to audit, so openness and coverage are limited.

Top on Arena Agent

29 models · Net Improvement

Source: arena.ai/leaderboard/agent · aggregated 2026-07-19