ixio
← Top Benchmark
benchmark · model-level

Stratix Cup

Tournament score·16 models·homepage ↗
Rank · benchmarks#10 / 19
Score43

How it works

LayerLens's tournament: 16 frontier models each write code for a soccer-strategy bot, and the bots compete head-to-head in a round-robin, with every match traced and signed. Score is tournament performance.

Metric
Tournament score
Models
16
Type
Model-level

How we ranked it

composite 43 / 100

Every benchmark is scored on four weighted criteria — how directly it measures a real coding agent, how much of the stack it covers, how real its tasks are, and how open and reproducible it is. Those blend into the composite that ranks it #10 of 19.

Agent-nativeweight 30%
35

Scores a real (harness × model) pair — the full agent — not just the bare model.

Stack coverageweight 30%
0

Data-driven: this benchmark exposes 16 models.

Task realismweight 25%
80

Executable, real-world coding tasks with hard pass/fail — not human preference or aggregate scores.

Open & reproducibleweight 15%
85

Open data with per-run receipts you can audit.

Our review

A refreshing executable, adversarial format — models write real code that has to actually work against other code, and the receipts are public. The knock is domain narrowness (one game) and that it's model-level, so it's a fun, credible signal rather than a broad one.

Top on Stratix Cup

16 models · Tournament score

Source: layerlens.ai/stratix-cup/season-1/schedule · aggregated 2026-07-19