LongJudgeBench: leaderboard
Metric: Judge accuracy (%) averaged over the six datasets (DeepResearch Bench, RealDeepResearch, SurGE structure and content, WritingPreferenceBench, VerifyBench, MA): agreement of the LLM judge's pairwise preferences (from pointwise scores, pairwise choices or listwise ranks) or correctness labels with the human reference judgments, averaged within query groups, temperature 0, direct judging with no rubric or reference; higher is better. Source: arxiv.org. Saturation forecast: Around June 2027. 8 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GLM-5.1 | 60.86 |
| 2 | DeepSeek V4 Flash | 59.58 |
| 3 | Kimi K2.6 | 58.9 |
| 4 | Qwen 3 Max | 58.09 |
| 5 | Qwen 3 32B (Non-reasoning) | 54.62 |
| 6 | Qwen 3 32B (Thinking) | 49.79 |
| 7 | GPT-5.2 | 46.05 |
| 8 | GPT-4o Mini | 38.86 |
Interactive version: theaggregate.ai/benchmark?slug=longjudgebench · How It Works · Data refreshed daily, snapshot 2026-09-29.