DeepWeb-Bench - Retrieval: leaderboard
Metric: Mean cell score (%) over the Retrieval capability family (finding the evidence cells), DeepWeb-Bench deep-research tasks answered cell by cell with the benchmark's web search, page visit and PDF tools (native browsing disabled), runs hosted in Claude Code CLI or Codex CLI at default settings; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.5 | 37.84 |
| 2 | Claude Opus 4.7 | 36.52 |
| 3 | GLM-5.1 | 34.19 |
| 4 | Claude Sonnet 4.6 | 33.8 |
| 5 | DeepSeek V4 Flash | 33.72 |
| 6 | DeepSeek V4 Pro | 32.89 |
| 7 | Qwen 3.6 Plus | 32.25 |
| 8 | MiniMax-M2.7 | 28.56 |
| 9 | Kimi K2.6 | 26.21 |
Interactive version: theaggregate.ai/benchmark?slug=deepweb-bench-retrieval · How It Works · Data refreshed daily, snapshot 2026-10-07.