Stratix Cup
How it works
LayerLens's tournament: 16 frontier models each write code for a soccer-strategy bot, and the bots compete head-to-head in a round-robin, with every match traced and signed. Score is tournament performance.
How we ranked it
composite 43 / 100Every benchmark is scored on four weighted criteria — how directly it measures a real coding agent, how much of the stack it covers, how real its tasks are, and how open and reproducible it is. Those blend into the composite that ranks it #10 of 19.
Scores a real (harness × model) pair — the full agent — not just the bare model.
Data-driven: this benchmark exposes 16 models.
Executable, real-world coding tasks with hard pass/fail — not human preference or aggregate scores.
Open data with per-run receipts you can audit.
Our review
A refreshing executable, adversarial format — models write real code that has to actually work against other code, and the receipts are public. The knock is domain narrowness (one game) and that it's model-level, so it's a fun, credible signal rather than a broad one.
Top on Stratix Cup
16 models · Tournament scoreSource: layerlens.ai/stratix-cup/season-1/schedule ↗ · aggregated 2026-07-19