ixio
Coding-agent model standings· updated 2026-07-22

The best models for coding agents.

403models/18benchmarks/979results

An agent is a model in a harness. We rank the models — every public benchmark, blended into one honest composite.

Featured model
Kimi K3open weight#1 · WebDev Arena

The #1 model for building web apps — and the strongest open-weight model on the board.

Moonshot AI·open-weight #1·overall #3
1,678WebDev Elo · #1
Explore Kimi K3

Model Standingsby capability

1
Claude Fable 5
Anthropic11 / 18 benchmarks
97
Composite
Coding95Agentic95WebDev99Math100Reasoning99
rankings by harness →403 models · domain scores blend 18 benchmarks

One dataset, four rankings

An agent is a model × harness. The board above ranks the models; the same benchmarks, re-cut, rank the harnesses, the teams, and the benchmarks themselves.

Agent=Model+Harness

The board above holds the harness fixed and ranks models. Flip it and rank the harnesses themselves on Top Coder.

How it works

updated 2026-07-22
01
Aggregate public leaderboards

We scrape established benchmark leaderboards daily and normalize their messy model and harness names into one dataset.

02
Split agent vs model

Agent benchmarks score a (harness, model) pair; model benchmarks score the raw model. That split is what lets us rank harnesses and labs separately.

03
Three rankings, one dataset

Top Model (best model per harness), Top Agent (the harnesses), and Top Team (each team by its single best model).

Benchmarks scraped

Open coding agents (CLIs/TUIs) across open-weight models.

Pass rate

Resolve real GitHub issues; hidden tests must pass (best per model).

% Resolved

Complete real end-to-end terminal tasks.

Accuracy

Our own runs — any harness × any model via the proxy. Fills gaps nobody else measures.

Pass rate
NL2Repoagent

Build a complete Python library from a natural-language spec; graded on the library's real upstream tests. Our own runs, ConTree + Tenki verified.

Pass rate

Human preference Elo on coding prompts (LMArena).

Coding Elo

Composite intelligence index across evals (Artificial Analysis).

Intelligence Index

Artificial Analysis’s coding-evals composite (SWE-bench, Terminal-Bench, SciCode, LiveCodeBench…), per model.

Coding Index
ARC-AGImodel

Abstraction & reasoning puzzles (ARC Prize).

Score

Agentic coding eval over real sessions.

Net Improvement

Human-preference Elo for building real web apps — two models generate a UI from one prompt, a human picks the winner (arena.ai).

WebDev Elo

16 frontier models write their own soccer-strategy code and compete head-to-head (LayerLens).

Tournament score

Exceptionally difficult research-level math, scored by Epoch AI on a private held-out set.

Accuracy

Harder, contamination-resistant real-repo issues (Scale SEAL), held-out private set.

% Resolved

Resolve real GitHub issues across 9 programming languages (best submission per model).

% Resolved

Multi-step tool use over real MCP servers — plan, call tools, and act (Scale SEAL).

Score

How much curated Agent Skills lift an agent on real tasks (BenchFlow); best with-skills per model.

Score

Hardest recent competition math, uncontaminated (ETH/INSAIT MathArena).

Accuracy
Scoring
  • Cells show each benchmark's own headline number — % resolved, accuracy, and so on.
  • Composite is percentile-blended across a model's benchmarks, so a 60% on a hard benchmark isn't unfairly beaten by an 80% on an easy one.
  • Tiers follow the field: top 25% excellent, bottom 40% iffy, the rest solid.
  • Numbers are exactly what the sources report. We aggregate and attribute — we don't re-run anything.