DeepWeb-Bench - Derivation: leaderboard

Metric: Mean cell score (%) over the Derivation capability family (computing derived values from the evidence), DeepWeb-Bench deep-research tasks answered cell by cell with the benchmark's web search, page visit and PDF tools (native browsing disabled), runs hosted in Claude Code CLI or Codex CLI at default settings; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 9 models tracked.

Top models

#ModelScore
1GPT-5.532.55
2Claude Opus 4.730.97
3DeepSeek V4 Pro27.73
4GLM-5.127.06
5Claude Sonnet 4.626.87
6DeepSeek V4 Flash26.77
7Qwen 3.6 Plus25.34
8MiniMax-M2.722.94
9Kimi K2.615.36

Interactive version: theaggregate.ai/benchmark?slug=deepweb-bench-derivation · How It Works · Data refreshed daily, snapshot 2026-10-07.