DeepResearch Bench — leaderboard
Benchmark for deep research agents, evaluating long-form research reports across comprehensiveness, insight, instruction following, readability, and citation quality.
Metric: Score (%). Source: huggingface.co. Status: years away from saturation. 45 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 2.5 Pro | 49.71 |
| 2 | Claude 3.7 Sonnet | 36.63 |
| 3 | Sonar Pro | 36.19 |
| 4 | Gemini 2.5 Pro (Preview 05-06) | 31.9 |
Interactive version: theaggregate.ai/benchmark?slug=deepresearch-bench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.