RecallRadar
Scores by dimension
4 scored · 0–100Breadth of coverage — how completely the report answers the question.
Depth and originality of the analysis beyond the surface facts.
How closely the report adheres to the brief it was given.
Structure and clarity of the finished write-up.
The benchmark
DeepResearch-Bench runs each agent on a set of real, multi-step web-research tasks — the kind that require planning a search, reading across many sources, and writing a cited report. An LLM judge scores every report on comprehensiveness, insight, instruction-following, and readability, and — for a subset of agents — audits citation accuracy, the share of a report's citations that genuinely support their claims. The overall score is the composite across those dimensions.
Source: DeepResearch-Bench ↗ · updated 2026-07-20