ixio
← Top Benchmark
benchmark · model-level

BrowseComp

Accuracy·95 models·homepage ↗
Rank · benchmarks#12 / 19
Score34

How it works

BrowseComp (OpenAI) tests hard web-browsing tasks that require an agent to find and synthesize information across the live web. OpenAI publishes no standing leaderboard, so we read the reported scores aggregated by steel.dev.

Metric
Accuracy
Models
95
Type
Model-level

How we ranked it

composite 34 / 100

Every benchmark is scored on four weighted criteria — how directly it measures a real coding agent, how much of the stack it covers, how real its tasks are, and how open and reproducible it is. Those blend into the composite that ranks it #12 of 19.

Agent-nativeweight 30%
30

Scores a real (harness × model) pair — the full agent — not just the bare model.

Stack coverageweight 30%
0

Data-driven: this benchmark exposes 95 models.

Task realismweight 25%
75

Executable, real-world coding tasks with hard pass/fail — not human preference or aggregate scores.

Open & reproducibleweight 15%
40

Open data with per-run receipts you can audit.

Our review

The task itself — real browsing under hard verification — is exactly the kind of agentic work we want to reward. The problem is provenance: the numbers are vendor-reported and third-party aggregated rather than independently re-run, so we rate its openness low and treat it with caution.

Top on BrowseComp

95 models · Accuracy

Source: leaderboard.steel.dev/leaderboards/browsecomp · aggregated 2026-07-19