LongJudgeBench (Rubric) - RealDeepResearch: leaderboard
Metric: Judge accuracy (%) on RealDeepResearch (640 expert-annotated Chinese and English deep-research reports in DOCX form, pointwise): agreement of the LLM judge's pairwise preferences (from pointwise scores, pairwise choices or listwise ranks) or correctness labels with the human reference judgments, averaged within query groups, temperature 0, judging with task-specific rubrics; higher is better. Source: arxiv.org. Saturation forecast: Around December 2027. 8 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | DeepSeek V4 Flash | 57.1 |
| 2 | GLM-5.1 | 56.06 |
| 3 | Kimi K2.6 | 51.75 |
| 4 | Qwen 3 Max | 50.19 |
| 5 | Qwen 3 32B (Thinking) | 50.17 |
| 6 | GPT-5.2 | 44.29 |
| 7 | Qwen 3 32B (Non-reasoning) | 38.81 |
| 8 | GPT-4o Mini | 34.58 |
Interactive version: theaggregate.ai/benchmark?slug=longjudgebench-rubric-realdeepresearch · How It Works · Data refreshed daily, snapshot 2026-09-29.