ixio
← Top Benchmark
benchmark · model-level

MCP Atlas

Score·26 models·homepage ↗
Rank · benchmarks#9 / 19
Score44

How it works

Scale's SEAL 'MCP Atlas': the model must complete multi-step tasks by planning and calling real tools exposed over MCP (Model Context Protocol) servers, then acting on the results.

Metric
Score
Models
26
Type
Model-level

How we ranked it

composite 44 / 100

Every benchmark is scored on four weighted criteria — how directly it measures a real coding agent, how much of the stack it covers, how real its tasks are, and how open and reproducible it is. Those blend into the composite that ranks it #9 of 19.

Agent-nativeweight 30%
40

Scores a real (harness × model) pair — the full agent — not just the bare model.

Stack coverageweight 30%
0

Data-driven: this benchmark exposes 26 models.

Task realismweight 25%
85

Executable, real-world coding tasks with hard pass/fail — not human preference or aggregate scores.

Open & reproducibleweight 15%
70

Open data with per-run receipts you can audit.

Our review

The closest model-level board to what an agent actually does — plan, call tools, use the output. That tool-use realism earns it a solid agent-native score for a model bench. It's Scale-run, so semi-open, and it isolates tool use rather than end-to-end coding.

Top on MCP Atlas

26 models · Score

Source: labs.scale.com/leaderboard/mcp_atlas · aggregated 2026-07-19