LongJudgeBench (Reference + Rubric): leaderboard

Metric: Judge accuracy (%) averaged over five datasets (DeepResearch Bench, RealDeepResearch, SurGE structure and content, VerifyBench, MA; WritingPreferenceBench has no references): agreement of the LLM judge's pairwise preferences (from pointwise scores, pairwise choices or listwise ranks) or correctness labels with the human reference judgments, averaged within query groups, temperature 0, judging with both references and rubrics; higher is better. Source: arxiv.org. Saturation forecast: Around 2030. 8 models tracked.

Top models

#ModelScore
1DeepSeek V4 Flash64.31
2Qwen 3 Max62.54
3Qwen 3 32B (Non-reasoning)60.81
4GLM-5.160.32
5Kimi K2.659.37
6GPT-4o Mini55.53
7Qwen 3 32B (Thinking)54.48
8GPT-5.245.48

Interactive version: theaggregate.ai/benchmark?slug=longjudgebench-reference-plus-rubric · How It Works · Data refreshed daily, snapshot 2026-09-29.