DeepWeb-Bench - Derivation: leaderboard
Metric: Mean cell score (%) over the Derivation capability family (computing derived values from the evidence), DeepWeb-Bench deep-research tasks answered cell by cell with the benchmark's web search, page visit and PDF tools (native browsing disabled), runs hosted in Claude Code CLI or Codex CLI at default settings; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.5 | 32.55 |
| 2 | Claude Opus 4.7 | 30.97 |
| 3 | DeepSeek V4 Pro | 27.73 |
| 4 | GLM-5.1 | 27.06 |
| 5 | Claude Sonnet 4.6 | 26.87 |
| 6 | DeepSeek V4 Flash | 26.77 |
| 7 | Qwen 3.6 Plus | 25.34 |
| 8 | MiniMax-M2.7 | 22.94 |
| 9 | Kimi K2.6 | 15.36 |
Interactive version: theaggregate.ai/benchmark?slug=deepweb-bench-derivation · How It Works · Data refreshed daily, snapshot 2026-10-07.