ixio
← Top Researcher
deep-research agent

Langchain-Open-Deep-Research-Gpt-5

DeepResearch-Bench·real multi-step web research·leaderboard ↗
Rank · agents#28 / 45
Overall49.3

Scores by dimension

5 scored · 0–100
Comprehensiveness
49.8

Breadth of coverage — how completely the report answers the question.

Insight
47.3

Depth and originality of the analysis beyond the surface facts.

Instruction-following
51.0

How closely the report adheres to the brief it was given.

Readability
49.0

Structure and clarity of the finished write-up.

Citation accuracy
34.7

Share of the report's citations that genuinely support their claims — audited for a subset.

The benchmark

DeepResearch-Bench runs each agent on a set of real, multi-step web-research tasks — the kind that require planning a search, reading across many sources, and writing a cited report. An LLM judge scores every report on comprehensiveness, insight, instruction-following, and readability, and — for a subset of agents — audits citation accuracy, the share of a report's citations that genuinely support their claims. The overall score is the composite across those dimensions.

Source: DeepResearch-Bench ↗ · updated 2026-07-21