DeepWeb-Bench - Calibration: leaderboard

Metric: Mean cell score (%) over the Calibration capability family (calibrated answers), DeepWeb-Bench deep-research tasks answered cell by cell with the benchmark's web search, page visit and PDF tools (native browsing disabled), runs hosted in Claude Code CLI or Codex CLI at default settings; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 9 models tracked.

Top models

#ModelScore
1GPT-5.534.16
2Claude Opus 4.731.14
3DeepSeek V4 Pro29.77
4GLM-5.129.56
5Claude Sonnet 4.628.89
6DeepSeek V4 Flash28.39
7Qwen 3.6 Plus27
8MiniMax-M2.724.69
9Kimi K2.616.39

Interactive version: theaggregate.ai/benchmark?slug=deepweb-bench-calibration · How It Works · Data refreshed daily, snapshot 2026-10-07.