AutoResearchBench - Deep Research: leaderboard

Metric: Accuracy (%) of pinpointing the single target paper on the Deep Research queries (progressive multi-step probing over full-text scientific literature), a ReAct agent with the DeepXiv scientific-literature search tool, at most 30 turns; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 13 models tracked.

Top models

#ModelScore
1Claude Opus 4.69.39
2Gemini 3.1 Pro (Preview)7.93
3GPT-5.47.44
4Qwen 3.5 397B A17B6.97
5Claude Sonnet 4.66.96
6Seed 2.0 Pro6.8
7Kimi K2.54.69
8DeepSeek V3.24.21
9Qwen 3.5 122B A10B3.88
10Qwen 3 Max3.24
11MiniMax-M2.52.91
12Qwen 3.5 35B A3B1.94

Interactive version: theaggregate.ai/benchmark?slug=autoresearchbench-deep-research · How It Works · Data refreshed daily, snapshot 2026-10-07.