AutoResearchBench - Wide Research: leaderboard

Metric: Intersection over union (%) between the collected paper set and the gold set of papers satisfying the query conditions on the Wide Research queries, a ReAct agent with the DeepXiv scientific-literature search tool, at most 30 turns; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 13 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)9.31
2GPT-5.48.12
3Seed 2.0 Pro7.87
4DeepSeek V3.27.7
5Qwen 3 Max6.89
6Claude Opus 4.66.56
7Kimi K2.56.23
8Claude Sonnet 4.65.83
9Qwen 3.5 397B A17B3.83
10Qwen 3.5 122B A10B2.76
11Qwen 3.5 35B A3B2.71
12MiniMax-M2.51.39

Interactive version: theaggregate.ai/benchmark?slug=autoresearchbench-wide-research · How It Works · Data refreshed daily, snapshot 2026-10-07.