QuestBench - Mean Score: leaderboard

Metric: Mean score (0-100) against question-specific grading criteria over the 256 Chinese deep research questions, with identical web search and visit tools and a nominal 50-call budget per question; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 13 models tracked.

Top models

#ModelScore
1GPT-5.567.12
2Claude Opus 4.757.79
3Gemini 3.1 Pro (Preview)43.07
4GLM-5.142.3
5DeepSeek V4 Pro38.24
6Kimi K2.637.74
7Qwen 3.6 Plus36.05
8MiMo-V2.5-Pro32.05
9Kimi K2.530.48
10MiniMax-M2.729.29
11DeepSeek V3.214.58

Interactive version: theaggregate.ai/benchmark?slug=questbench-mean-score · How It Works · Data refreshed daily, snapshot 2026-10-07.