ixio
Deep research · investigating the web · updated 2026-07-22

Which agent researches your topic best?

45agents/4 quality + 2 citationscored dims/13 of 45citation-audited
A deep-research agent doesn't write code — it plans a multi-step web investigation and returns a cited report. This ranks 45 of them on real research tasks, scored for comprehensiveness, insight, instruction-following, readability, and — where measured — citation accuracy.

Research agentsranked by overall score

source: DeepResearch-Bench ↗
#AgentCompreh.InsightInstr.Readab.CitationOverall
1Qianfan Deepresearch 043059.561.553.954.358.032ZTE-Nebula-DeepResearch-V2026051958.459.854.154.757.273Link58.259.753.255.057.084Zhipu Deep Research58.160.153.553.957.065Xiaoyi58.659.453.654.057.006WhaleCloud-DocChain57.159.354.055.056.817Cellcog-Max57.460.053.353.256.6781688AILab-DeepResearch-042857.359.353.553.456.539Octen-Deepresearch-050856.959.053.453.856.3110Grep-V556.858.953.453.456.2311Nvidia-Aiq-Nemotron-Gpt52-Updated56.958.552.953.455.95121688AILab-DeepResearch-032555.557.653.453.555.3913Ms Deepresearch Gpt52mixqwen35 09 Edit Restart09 Think Medium56.856.853.152.355.3114Drb Cellcog55.458.252.553.155.3115Deepinsight55.758.752.550.955.2416Ms Deepresearch56.556.253.351.754.9717TrajectoryKit54.157.952.952.754.9218Onyx54.756.453.152.054.5419Deepsynth54.256.152.951.854.2220Deepdog53.156.151.851.253.5221RecallRadar53.953.552.252.453.1922MindDR-V1.551.555.350.551.352.5423Tavily-Research52.853.651.949.252.4424Thinkdepthai-Deepresearch52.053.952.050.152.4325Salesforce-Air-Deep-Research50.051.150.850.350.6526Gensee-Search-Gpt-550.150.851.349.732.950.6027Gemini-2.5-Pro-Deepresearch49.549.550.150.078.349.7128Langchain-Open-Deep-Research-Gpt-549.847.351.049.034.749.3329Openai-Deepresearch46.543.749.447.275.046.4530Raaa-Deep-Research43.848.347.243.846.1331Dr-Tulu44.144.649.642.345.4932Claude-Research45.342.847.644.745.0033Kimi-Researcher45.042.047.145.644.6434Doubao-Deepresearch44.840.648.044.752.944.3435Langchain-Open-Deep-Research43.039.248.145.249.143.4436Nvidia-Aiq-Research-Assistant38.038.444.642.640.5237tongyi-deepresearch-30B-A3B39.534.446.244.340.4638perplexity-Research39.135.646.143.182.640.4639Grok-Deeper-Search36.130.946.642.273.138.2240Sonar-Reasoning-Pro35.031.644.942.445.237.7641Sonar-Reasoning34.732.644.442.452.637.7542Claude-3-7-Sonnet-With-Search36.031.344.036.187.336.6343Sonar-Pro33.929.743.441.179.736.1944Gemini-2.5-Pro-Preview-05-0631.824.640.232.831.9045Gpt-4o-Search-Preview27.820.441.037.686.630.74
How it's measured

Each agent runs the same set of real research prompts; an LLM-judge rubric scores the reports on comprehensiveness, insight, instruction-following, and readability, and the overall is their composite. Citation accuracy — the share of a report's citations that actually support their claims — is audited for a subset (13 agents here).

A different axis

These are research agents, not coding agents — ranked here for completeness of the agent landscape, sourced from DeepResearch-Bench and read straight from its open leaderboard, refreshed daily.