DeepResearchBench — leaderboard

Evaluates deep research agents on multi-step information gathering and synthesis tasks. Tests ability to conduct thorough research across complex topics.

Metric: Average Score. Source: deepresearch-bench.github.io. Status: saturation imminent. 36 models tracked.

Top models

#ModelScore
1Claude Opus 4.6 (High)55.31
2GPT-555.13
3Claude Sonnet 4.6 (High)54.87
4GPT-5.5 (High)54.01
5Claude Opus 4.6 (Medium)53.24
6GPT-5 (Low)51
7Claude Opus 4.8 (High)50.23
8Gemini 2.5 Pro49.71
9GPT-5 (Minimal)49.7
10Claude Opus 4.1 (20250805)49.7
11GPT-5.5 (Medium)49.55
12GPT-5 (Medium)49.5
13GPT-5.5 (Low)48.73
14GPT-5 (High)48.6
15Grok 447.9

Interactive version: theaggregate.ai/benchmark?slug=deepresearchbench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.