ixio
← Top Benchmark
benchmark · model-level

Chatbot Arena (Coding)

Coding Elo·279 models·homepage ↗
Rank · benchmarks#18 / 19
Score26

How it works

LMArena's coding split: humans are shown two anonymous model responses to a coding prompt and vote for the better one; votes are aggregated into an Elo rating. A huge, continuously-updated sample of human preference.

Metric
Coding Elo
Models
279
Type
Model-level

How we ranked it

composite 26 / 100

Every benchmark is scored on four weighted criteria — how directly it measures a real coding agent, how much of the stack it covers, how real its tasks are, and how open and reproducible it is. Those blend into the composite that ranks it #18 of 19.

Agent-nativeweight 30%
0

Scores a real (harness × model) pair — the full agent — not just the bare model.

Stack coverageweight 30%
0

Data-driven: this benchmark exposes 279 models.

Task realismweight 25%
60

Executable, real-world coding tasks with hard pass/fail — not human preference or aggregate scores.

Open & reproducibleweight 15%
70

Open data with per-run receipts you can audit.

Our review

Enormous sample size and hard to game, but it measures the wrong thing for us: preference on a chat snippet, not whether code runs. Nothing is executed and no harness is involved, so it's a proxy for perceived quality rather than task success. Useful context, low weight.

Top on Chatbot Arena (Coding)

279 models · Coding Elo

Source: openlm.ai/chatbot-arena · aggregated 2026-07-19