JudgeSense - Coherence: leaderboard

Metric: Judge Sensitivity Score: share of items on which the judge returns the same verdict under two reworded instruction templates that ask for the same judgement, over the pairs it answered (refusals and malformed outputs left out); temperature 0 (claude-opus-4-7 at the provider default), 1,024 output tokens, on 250 SummEval summaries rated 1 to 5 for coherence, agreement on the exact rating; higher is better. Source: arxiv.org. Saturation forecast: Around March 2027. 25 models tracked.

Top models

#ModelScore
1Qwen 2.5 72B Instruct0.86
2Gemma 4 31B (IT)0.85
3Gemini 3.7 Flash0.84
4Kimi K30.83
5Claude Sonnet 4.50.83
6Claude Haiku 4.5 (20251001)0.79
7Llama 3.3 70B Instruct0.77
8Llama 4 Maverick Instruct FP80.76
9Qwen 3.6 35B A3B0.76
10Qwen 3.8 27B0.76
11Claude Opus 4.70.74
12DeepSeek V4 Pro (0813)0.72
13Qwen 3 32B0.7
14Gemini 2.5 Flash0.69
15Llama 3.1 8B Instruct0.67

Interactive version: theaggregate.ai/benchmark?slug=judgesense-coherence · How It Works · Data refreshed daily, snapshot 2026-10-07.