OneMillion-Bench (CN, Web Search): leaderboard

Metric: Pass rate (%): share of tasks whose expert score reaches the 0.7 professional threshold on the 200-task CN set (Chinese-context tasks; Finance, Law, Healthcare, Industry and Natural Science, 40 tasks each) of OneMillion-Bench's expert-curated professional tasks, each model equipped with web-search tool calling (deep-research agents as they are); rubrics graded by an LLM judge; higher is better. Source: arxiv.org. Saturation forecast: Around March 2027. 19 models tracked.

Top models

#ModelScoreOverall rank
1Claude Opus 4.6 (High)48.5#60 (Claude Opus 4.6)
2GPT-5.4 (High)35.5#76 (GPT-5.4)
3Gemini 3 Pro (Preview)30#64
4GPT-5.2 (High)28#105 (GPT-5.2)
5Doubao-Seed-2-0-Pro-260215 (High)25
6GPT-5.3 Codex (High)23#68 (GPT-5.3 Codex)
7DeepSeek V3.2 Speciale16.5#125
8O3 Deep Research16#109
9Kimi K2.5 (High)15#139 (Kimi K2.5)
10Sonar Deep Research13.5#191
11Grok 411.5#169
12GLM-511#137
13O4 Mini Deep Research9#219
14MiniMax-M2.57.5#295
15Step 3.5 Flash6.5#262

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=onemillion-bench-cn-web-search · How It Works · Data refreshed daily, snapshot 2026-10-11.