Chatbot Arena (Coding)
How it works
LMArena's coding split: humans are shown two anonymous model responses to a coding prompt and vote for the better one; votes are aggregated into an Elo rating. A huge, continuously-updated sample of human preference.
How we ranked it
composite 26 / 100Every benchmark is scored on four weighted criteria — how directly it measures a real coding agent, how much of the stack it covers, how real its tasks are, and how open and reproducible it is. Those blend into the composite that ranks it #18 of 19.
Scores a real (harness × model) pair — the full agent — not just the bare model.
Data-driven: this benchmark exposes 279 models.
Executable, real-world coding tasks with hard pass/fail — not human preference or aggregate scores.
Open data with per-run receipts you can audit.
Our review
Enormous sample size and hard to game, but it measures the wrong thing for us: preference on a chat snippet, not whether code runs. Nothing is executed and no harness is involved, so it's a proxy for perceived quality rather than task success. Useful context, low weight.
Top on Chatbot Arena (Coding)
279 models · Coding EloSource: openlm.ai/chatbot-arena ↗ · aggregated 2026-07-19