The best models for coding agents.
An agent is a model in a harness. We rank the models — every public benchmark, blended into one honest composite.
The #1 model for building web apps — and the strongest open-weight model on the board.
Model Standingsby capability
One dataset, four rankings
An agent is a model × harness. The board above ranks the models; the same benchmarks, re-cut, rank the harnesses, the teams, and the benchmarks themselves.
The board above holds the harness fixed and ranks models. Flip it and rank the harnesses themselves on Top Coder.
How it works
updated 2026-07-22We scrape established benchmark leaderboards daily and normalize their messy model and harness names into one dataset.
Agent benchmarks score a (harness, model) pair; model benchmarks score the raw model. That split is what lets us rank harnesses and labs separately.
Top Model (best model per harness), Top Agent (the harnesses), and Top Team (each team by its single best model).
Open coding agents (CLIs/TUIs) across open-weight models.
Resolve real GitHub issues; hidden tests must pass (best per model).
Complete real end-to-end terminal tasks.
Our own runs — any harness × any model via the proxy. Fills gaps nobody else measures.
Build a complete Python library from a natural-language spec; graded on the library's real upstream tests. Our own runs, ConTree + Tenki verified.
Human preference Elo on coding prompts (LMArena).
Composite intelligence index across evals (Artificial Analysis).
Artificial Analysis’s coding-evals composite (SWE-bench, Terminal-Bench, SciCode, LiveCodeBench…), per model.
Abstraction & reasoning puzzles (ARC Prize).
Agentic coding eval over real sessions.
Human-preference Elo for building real web apps — two models generate a UI from one prompt, a human picks the winner (arena.ai).
16 frontier models write their own soccer-strategy code and compete head-to-head (LayerLens).
Exceptionally difficult research-level math, scored by Epoch AI on a private held-out set.
Harder, contamination-resistant real-repo issues (Scale SEAL), held-out private set.
Resolve real GitHub issues across 9 programming languages (best submission per model).
Multi-step tool use over real MCP servers — plan, call tools, and act (Scale SEAL).
How much curated Agent Skills lift an agent on real tasks (BenchFlow); best with-skills per model.
Hardest recent competition math, uncontaminated (ETH/INSAIT MathArena).
- Cells show each benchmark's own headline number — % resolved, accuracy, and so on.
- Composite is percentile-blended across a model's benchmarks, so a 60% on a hard benchmark isn't unfairly beaten by an 80% on an easy one.
- Tiers follow the field: top 25% excellent, bottom 40% iffy, the rest solid.
- Numbers are exactly what the sources report. We aggregate and attribute — we don't re-run anything.