DeepWeb-Bench - Retrieval: leaderboard

Metric: Mean cell score (%) over the Retrieval capability family (finding the evidence cells), DeepWeb-Bench deep-research tasks answered cell by cell with the benchmark's web search, page visit and PDF tools (native browsing disabled), runs hosted in Claude Code CLI or Codex CLI at default settings; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 9 models tracked.

Top models

#ModelScore
1GPT-5.537.84
2Claude Opus 4.736.52
3GLM-5.134.19
4Claude Sonnet 4.633.8
5DeepSeek V4 Flash33.72
6DeepSeek V4 Pro32.89
7Qwen 3.6 Plus32.25
8MiniMax-M2.728.56
9Kimi K2.626.21

Interactive version: theaggregate.ai/benchmark?slug=deepweb-bench-retrieval · How It Works · Data refreshed daily, snapshot 2026-10-07.