WebDev Arena
How it works
Arena.ai's WebDev Arena: two models each generate a web app from the same prompt, and a human picks the better result; votes aggregate into an Elo rating.
How we ranked it
composite 33 / 100Every benchmark is scored on four weighted criteria — how directly it measures a real coding agent, how much of the stack it covers, how real its tasks are, and how open and reproducible it is. Those blend into the composite that ranks it #13 of 19.
Scores a real (harness × model) pair — the full agent — not just the bare model.
Data-driven: this benchmark exposes 87 models.
Executable, real-world coding tasks with hard pass/fail — not human preference or aggregate scores.
Open data with per-run receipts you can audit.
Our review
Higher realism than chat-preference Elo because the models produce actual working web UIs that get judged — a real, if subjective, artifact. But it's single-shot preference voting (no iteration, no tools) on a proprietary board, so it's a useful signal for one slice of coding, weighted accordingly.
Top on WebDev Arena
87 models · WebDev EloSource: arena.ai/leaderboard/code/webdev ↗ · aggregated 2026-07-19