EvoBrowseComp-ZH: leaderboard

Metric: Accuracy (%) on the 400 Chinese EvoBrowseComp questions (complex multi-hop questions synthesized from fresh live-web knowledge to resist parametric recall), search agent with web search and page visit tools (Tongyi DeepResearch tool definitions), answers judged correct against the reference by a GLM-5-Chat judge; temperature 0.6, top_p 0.95, 128K context, at most 40 tool calls; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 14 models tracked.

Top models

#ModelScore
1Claude Opus 4.8 (Max)38.3
2Claude Opus 4.6 (Max)36.8
3Qwen 3.5 397B A17B34.5
4GLM-530.5
5DeepSeek V3.2 (Thinking)30.5
6DeepSeek V4 Flash (High)24.8
7Qwen 3.5 122B A10B21
8DeepSeek V4 Flash (Non-reasoning)20
9Kimi K2.619.8
10Qwen 3.5 27B18.6
11Qwen 3 235B A22B 2507 (Thinking)17.5
12DeepSeek V4 Pro (Max)15
13Qwen 3.5 35B A3B14.2
14DeepSeek V4 Flash (Max)10.8

Interactive version: theaggregate.ai/benchmark?slug=evobrowsecomp-zh · How It Works · Data refreshed daily, snapshot 2026-09-29.