ixio
← Top Benchmark
benchmark · model-level

WebDev Arena

WebDev Elo·87 models·homepage ↗
Rank · benchmarks#13 / 19
Score33

How it works

Arena.ai's WebDev Arena: two models each generate a web app from the same prompt, and a human picks the better result; votes aggregate into an Elo rating.

Metric
WebDev Elo
Models
87
Type
Model-level

How we ranked it

composite 33 / 100

Every benchmark is scored on four weighted criteria — how directly it measures a real coding agent, how much of the stack it covers, how real its tasks are, and how open and reproducible it is. Those blend into the composite that ranks it #13 of 19.

Agent-nativeweight 30%
15

Scores a real (harness × model) pair — the full agent — not just the bare model.

Stack coverageweight 30%
0

Data-driven: this benchmark exposes 87 models.

Task realismweight 25%
78

Executable, real-world coding tasks with hard pass/fail — not human preference or aggregate scores.

Open & reproducibleweight 15%
60

Open data with per-run receipts you can audit.

Our review

Higher realism than chat-preference Elo because the models produce actual working web UIs that get judged — a real, if subjective, artifact. But it's single-shot preference voting (no iteration, no tools) on a proprietary board, so it's a useful signal for one slice of coding, weighted accordingly.

Top on WebDev Arena

87 models · WebDev Elo

Source: arena.ai/leaderboard/code/webdev · aggregated 2026-07-19