DeepResearchBench: leaderboard

Evaluates deep research agents on multi-step information gathering and synthesis tasks. Tests ability to conduct thorough research across complex topics.

Metric: Average Score. Source: deepresearch-bench.github.io. Status: years away from saturation. 41 models tracked.

Top models

#ModelScore
1Claude Opus 4.6 (High)55.3
2Claude Sonnet 4.6 (High)54.9
3Claude Opus 4.5 (20251101) (High)54.8
4GPT-5.5 (High)54
5Claude Opus 4.8 (High)50.2
6GPT-5 (Low)49.6
7GPT-5.5 (Medium)49.6
8Claude Opus 4.8 (Low)49.3
9Gemini 3 Flash49
10GPT-5 (Minimal)48.9
11GPT-5.5 (Low)48.7
12GPT-5 (Medium)48.6
13Claude Opus 4.1 (20250805)48.3
14GPT-5 (High)48.1
15Gemini 3 Flash (Preview) (High)47.9

Interactive version: theaggregate.ai/benchmark?slug=deepresearchbench · How It Works · Data refreshed daily, snapshot 2026-09-05.