DeepResearch Bench - Comprehensiveness — leaderboard
Metric: Score (%). Source: huggingface.co. 45 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 2.5 Pro | 49.51 |
| 2 | Claude 3.7 Sonnet | 35.95 |
| 3 | Sonar Pro | 33.92 |
| 4 | Gemini 2.5 Pro (Preview 05-06) | 31.75 |
Interactive version: theaggregate.ai/benchmark?slug=deepresearch-bench-comprehensiveness · How the rankings work · Data refreshed daily, snapshot 2026-07-22.