KDR-Bench: leaderboard

Metric: RACE general-purpose score (0-100): equal-weight mean of comprehensiveness, depth, coherence and readability, each scored against a reference report (50 means parity with the reference), on the 41 expert-level research questions of KDR-Bench in nine domains backed by 1,252 Statista tables, each report judged by DeepSeek-V3.2; higher is better. Source: arxiv.org. Saturation forecast: Around March 2027. 9 models tracked.

Top models

#ModelScore
1MiniMax-M244.7
2Qwen 3 Max44.6
3Tongyi DeepResearch 30B A3B41.8
4Deep Research41.3
5GLM-4.640.8

Interactive version: theaggregate.ai/benchmark?slug=kdr-bench · How It Works · Data refreshed daily, snapshot 2026-10-07.