EvoBrowseComp-EN: leaderboard

Metric: Accuracy (%) on the 400 English EvoBrowseComp questions (complex multi-hop questions synthesized from fresh live-web knowledge to resist parametric recall), search agent with web search and page visit tools (Tongyi DeepResearch tool definitions), answers judged correct against the reference by a GLM-5-Chat judge; temperature 0.6, top_p 0.95, 128K context, at most 40 tool calls; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 14 models tracked.

Top models

#ModelScore
1Claude Opus 4.8 (Max)46.2
2Claude Opus 4.6 (Max)44.8
3Qwen 3.5 397B A17B42
4GLM-539.2
5DeepSeek V4 Flash (High)34.5
6Kimi K2.631
7Qwen 3.5 35B A3B29.2
8Qwen 3.5 122B A10B29.2
9DeepSeek V4 Flash (Non-reasoning)27.2
10Qwen 3.5 27B25
11DeepSeek V3.2 (Thinking)23
12Qwen 3 235B A22B 2507 (Thinking)16.5
13DeepSeek V4 Pro (Max)16.5
14DeepSeek V4 Flash (Max)16.5

Interactive version: theaggregate.ai/benchmark?slug=evobrowsecomp-en · How It Works · Data refreshed daily, snapshot 2026-09-29.