DeepResearch Bench — leaderboard

Benchmark for deep research agents, evaluating long-form research reports across comprehensiveness, insight, instruction following, readability, and citation quality.

Metric: Score (%). Source: huggingface.co. Status: years away from saturation. 45 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro49.71
2Claude 3.7 Sonnet36.63
3Sonar Pro36.19
4Gemini 2.5 Pro (Preview 05-06)31.9

Interactive version: theaggregate.ai/benchmark?slug=deepresearch-bench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.