SkillsBench
How it works
BenchFlow's SkillsBench measures how much curated Agent Skills lift an agent: the same tasks are run with and without a skills bundle, and the delta is reported. We take each model's best with-skills score.
How we ranked it
composite 45 / 100Every benchmark is scored on four weighted criteria — how directly it measures a real coding agent, how much of the stack it covers, how real its tasks are, and how open and reproducible it is. Those blend into the composite that ranks it #8 of 19.
Scores a real (harness × model) pair — the full agent — not just the bare model.
Data-driven: this benchmark exposes 22 models.
Executable, real-world coding tasks with hard pass/fail — not human preference or aggregate scores.
Open data with per-run receipts you can audit.
Our review
A genuinely novel angle — it measures the harness/skills layer, not just the model, which is squarely on our theme. Open data and harness. It answers a narrow question (how much do skills help), so coverage is limited, but it's a welcome look at the part of the stack most boards ignore.
Top on SkillsBench
22 models · ScoreSource: raw.githubusercontent.com/benchflow-ai/skillsbench/main/website/src/data/leaderboard-data.ts ↗ · aggregated 2026-07-19