JudgeSense - Coherence: leaderboard
Metric: Judge Sensitivity Score: share of items on which the judge returns the same verdict under two reworded instruction templates that ask for the same judgement, over the pairs it answered (refusals and malformed outputs left out); temperature 0 (claude-opus-4-7 at the provider default), 1,024 output tokens, on 250 SummEval summaries rated 1 to 5 for coherence, agreement on the exact rating; higher is better. Source: arxiv.org. Saturation forecast: Around March 2027. 25 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen 2.5 72B Instruct | 0.86 |
| 2 | Gemma 4 31B (IT) | 0.85 |
| 3 | Gemini 3.7 Flash | 0.84 |
| 4 | Kimi K3 | 0.83 |
| 5 | Claude Sonnet 4.5 | 0.83 |
| 6 | Claude Haiku 4.5 (20251001) | 0.79 |
| 7 | Llama 3.3 70B Instruct | 0.77 |
| 8 | Llama 4 Maverick Instruct FP8 | 0.76 |
| 9 | Qwen 3.6 35B A3B | 0.76 |
| 10 | Qwen 3.8 27B | 0.76 |
| 11 | Claude Opus 4.7 | 0.74 |
| 12 | DeepSeek V4 Pro (0813) | 0.72 |
| 13 | Qwen 3 32B | 0.7 |
| 14 | Gemini 2.5 Flash | 0.69 |
| 15 | Llama 3.1 8B Instruct | 0.67 |
Interactive version: theaggregate.ai/benchmark?slug=judgesense-coherence · How It Works · Data refreshed daily, snapshot 2026-10-07.