BrowseComp
How it works
BrowseComp (OpenAI) tests hard web-browsing tasks that require an agent to find and synthesize information across the live web. OpenAI publishes no standing leaderboard, so we read the reported scores aggregated by steel.dev.
How we ranked it
composite 34 / 100Every benchmark is scored on four weighted criteria — how directly it measures a real coding agent, how much of the stack it covers, how real its tasks are, and how open and reproducible it is. Those blend into the composite that ranks it #12 of 19.
Scores a real (harness × model) pair — the full agent — not just the bare model.
Data-driven: this benchmark exposes 95 models.
Executable, real-world coding tasks with hard pass/fail — not human preference or aggregate scores.
Open data with per-run receipts you can audit.
Our review
The task itself — real browsing under hard verification — is exactly the kind of agentic work we want to reward. The problem is provenance: the numbers are vendor-reported and third-party aggregated rather than independently re-run, so we rate its openness low and treat it with caution.
Top on BrowseComp
95 models · AccuracySource: leaderboard.steel.dev/leaderboards/browsecomp ↗ · aggregated 2026-07-19