ixio
coding-agent leaderboard · teams

Which team makes the best coding model?

64teams/452models/2026-07-22updated

Every team is only as good as its single best model. So each one here is ranked by that model's composite — a percentile-blended score across every public benchmark we scrape, fair across scales — not by how many models it ships. Flip to Open to rank by best open-weight model.

Team leaderboard

64 teams
#TeamBest model
1
Anthropic
Claude Fable 542
96
2
OpenAI
GPT-5.6 Sol54
95
3
Moonshot AI
open
Kimi K319
90
4
Meta
Muse Spark21
85
5
Google
Gemini-3.6-Flash36
78
6
xAI
Grok 4.517
78
7
NVIDIA
open
Nemotron-3 Ultra 550B12
75
8
Z.ai
open
GLM-5.215
73
9
Alibaba
Qwen3.6-Plus44
73
10
MiroMind
MiroThinker-H12
71
11
Google DeepMind
AI Co-Mathematician2
71
12
Baidu
ERNIE-5.12
71
13
ByteDance
Seed2.0 Pro3
71
14
Bytedance
seed-2.1-pro-preview1
68
15
DeepSeek
open
DeepSeek-V4-Pro-Max27
67
16
StepFun
open
Step-3.5-Flash5
67
17
Meituan
open
LongCat-Flash-Chat2
66
18
MiniMax
open
MiniMax M39
64
19
Xiaomi
MiMo-V2-Pro5
64
20
Tencent
Hunyuan-Hy311
63
21
Academic Research
Agents-A18
62
22
Exa
Exa Agent3
61
23
Amazon
Amazon-Nova-Chat-11-105
60
24
Mistral AI
Mistral Medium 3.117
58
25
Microsoft
MAI-1-Preview8
58
26
Parallel
Parallel Ultra8x6
54
27
Prime Intellect
open
INTELLECT-31
54
28
MIT
GLM-5.2 (Max) Z.ai ·2
53
29
Thinky
Thinking Machines Inkling1
53
30
Thinking Machines
Inkling1
51
31
Cohere
Command A (03-2025)6
49
32
Ai2
open
OLMo-3-32b-think5
49
33
01 AI
Yi-Lightning5
49
34
NexusFlow
Athene-v2-Chat-72B2
49
35
Sarvam AI
Sarvam-105B2
47
36
Perplexity
Perplexity Agent Advanced5
46
37
PolarSeeker
OpenSeeker-v22
44
38
LM-Provers
QED-Nano1
41
39
Alibaba Cloud / Tongyi Lab
Tongyi DeepResearch5
41
40
Zhipu AI
GLM-4.7-Flash1
40
41
AI21 Labs
open
Jamba-1.5-Large2
40
42
Poolside
open
laguna-m.12
40
43
Tavily
Tavily + GPT-5.4 harness1
40
44
Reka AI
Reka-Core-202409044
40
45
Princeton
open
Gemma-2-9B-it-SimPO1
38
46
IBM
open
Granite-3.1-8B-Instruct5
35
47
InternLM
InternLM2.5-20B-chat1
34
48
Proprietary
KAT-Coder-Pro-V11
33
49
HKUST NLP Group
WebExplorer-8B (RL)1
33
50
Nexusflow
open
Starling-LM-7B-beta1
32
51
THUDM / Tsinghua University
DeepDive-32B1
32
52
HuggingFace
open
Zephyr-ORPO-141b-A35b-v0.12
32
53
Databricks
DBRX-Instruct-Preview1
32
54
OpenChat
open
OpenChat-3.5-01062
30
55
UC Berkeley
Starling-LM-7B-alpha1
29
56
NousResearch
open
Nous-Hermes-2-Mixtral-8x7B-DPO2
29
57
Snowflake
open
Snowflake Arctic Instruct1
28
58
LMSYS
Vicuna-33B1
27
59
Inception AI
mercury-21
27
60
Upstage AI
SOLAR-10.7B-Instruct-v1.01
27
61
Perplexity AI
pplx-70B-online1
26
62
MosaicML
MPT-30B-chat1
26
63
Cognitive
open
Dolphin-2.2.1-Mistral-7B1
26
64
MBZUAI Institute of Foundation Models
K2 Think V21
13

A team's score is the composite of its top-ranked model — shipping more models never helps unless one is genuinely better. Best model → opens that model's full profile.