KDR-Bench: leaderboard
Metric: RACE general-purpose score (0-100): equal-weight mean of comprehensiveness, depth, coherence and readability, each scored against a reference report (50 means parity with the reference), on the 41 expert-level research questions of KDR-Bench in nine domains backed by 1,252 Statista tables, each report judged by DeepSeek-V3.2; higher is better. Source: arxiv.org. Saturation forecast: Around March 2027. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | MiniMax-M2 | 44.7 |
| 2 | Qwen 3 Max | 44.6 |
| 3 | Tongyi DeepResearch 30B A3B | 41.8 |
| 4 | Deep Research | 41.3 |
| 5 | GLM-4.6 | 40.8 |
Interactive version: theaggregate.ai/benchmark?slug=kdr-bench · How It Works · Data refreshed daily, snapshot 2026-10-07.