Deep research · investigating the web · updated 2026-07-22
Which agent researches your topic best?
45agents/4 quality + 2 citationscored dims/13 of 45citation-audited
A deep-research agent doesn't write code — it plans a multi-step web investigation and returns a cited report. This ranks 45 of them on real research tasks, scored for comprehensiveness, insight, instruction-following, readability, and — where measured — citation accuracy.
Research agentsranked by overall score
source: DeepResearch-Bench ↗#AgentCompreh.InsightInstr.Readab.CitationOverall
1Qianfan Deepresearch 043059.561.553.954.3—58.032ZTE-Nebula-DeepResearch-V2026051958.459.854.154.7—57.273Link58.259.753.255.0—57.084Zhipu Deep Research58.160.153.553.9—57.065Xiaoyi58.659.453.654.0—57.006WhaleCloud-DocChain57.159.354.055.0—56.817Cellcog-Max57.460.053.353.2—56.6781688AILab-DeepResearch-042857.359.353.553.4—56.539Octen-Deepresearch-050856.959.053.453.8—56.3110Grep-V556.858.953.453.4—56.2311Nvidia-Aiq-Nemotron-Gpt52-Updated56.958.552.953.4—55.95121688AILab-DeepResearch-032555.557.653.453.5—55.3913Ms Deepresearch Gpt52mixqwen35 09 Edit Restart09 Think Medium56.856.853.152.3—55.3114Drb Cellcog55.458.252.553.1—55.3115Deepinsight55.758.752.550.9—55.2416Ms Deepresearch56.556.253.351.7—54.9717TrajectoryKit54.157.952.952.7—54.9218Onyx54.756.453.152.0—54.5419Deepsynth54.256.152.951.8—54.2220Deepdog53.156.151.851.2—53.5221RecallRadar53.953.552.252.4—53.1922MindDR-V1.551.555.350.551.3—52.5423Tavily-Research52.853.651.949.2—52.4424Thinkdepthai-Deepresearch52.053.952.050.1—52.4325Salesforce-Air-Deep-Research50.051.150.850.3—50.6526Gensee-Search-Gpt-550.150.851.349.732.950.6027Gemini-2.5-Pro-Deepresearch49.549.550.150.078.349.7128Langchain-Open-Deep-Research-Gpt-549.847.351.049.034.749.3329Openai-Deepresearch46.543.749.447.275.046.4530Raaa-Deep-Research43.848.347.243.8—46.1331Dr-Tulu44.144.649.642.3—45.4932Claude-Research45.342.847.644.7—45.0033Kimi-Researcher45.042.047.145.6—44.6434Doubao-Deepresearch44.840.648.044.752.944.3435Langchain-Open-Deep-Research43.039.248.145.249.143.4436Nvidia-Aiq-Research-Assistant38.038.444.642.6—40.5237tongyi-deepresearch-30B-A3B39.534.446.244.3—40.4638perplexity-Research39.135.646.143.182.640.4639Grok-Deeper-Search36.130.946.642.273.138.2240Sonar-Reasoning-Pro35.031.644.942.445.237.7641Sonar-Reasoning34.732.644.442.452.637.7542Claude-3-7-Sonnet-With-Search36.031.344.036.187.336.6343Sonar-Pro33.929.743.441.179.736.1944Gemini-2.5-Pro-Preview-05-0631.824.640.232.8—31.9045Gpt-4o-Search-Preview27.820.441.037.686.630.74
How it's measured
Each agent runs the same set of real research prompts; an LLM-judge rubric scores the reports on comprehensiveness, insight, instruction-following, and readability, and the overall is their composite. Citation accuracy — the share of a report's citations that actually support their claims — is audited for a subset (13 agents here).
A different axis
These are research agents, not coding agents — ranked here for completeness of the agent landscape, sourced from DeepResearch-Bench and read straight from its open leaderboard, refreshed daily.