ixio
← Top Benchmark
benchmark · model-level

SkillsBench

Score·22 models·homepage ↗
Rank · benchmarks#8 / 19
Score45

How it works

BenchFlow's SkillsBench measures how much curated Agent Skills lift an agent: the same tasks are run with and without a skills bundle, and the delta is reported. We take each model's best with-skills score.

Metric
Score
Models
22
Type
Model-level

How we ranked it

composite 45 / 100

Every benchmark is scored on four weighted criteria — how directly it measures a real coding agent, how much of the stack it covers, how real its tasks are, and how open and reproducible it is. Those blend into the composite that ranks it #8 of 19.

Agent-nativeweight 30%
40

Scores a real (harness × model) pair — the full agent — not just the bare model.

Stack coverageweight 30%
0

Data-driven: this benchmark exposes 22 models.

Task realismweight 25%
80

Executable, real-world coding tasks with hard pass/fail — not human preference or aggregate scores.

Open & reproducibleweight 15%
88

Open data with per-run receipts you can audit.

Our review

A genuinely novel angle — it measures the harness/skills layer, not just the model, which is squarely on our theme. Open data and harness. It answers a narrow question (how much do skills help), so coverage is limited, but it's a welcome look at the part of the stack most boards ignore.

Top on SkillsBench

22 models · Score

Source: raw.githubusercontent.com/benchflow-ai/skillsbench/main/website/src/data/leaderboard-data.ts · aggregated 2026-07-19