JudgeSense - Preference: leaderboard
Metric: Judge Sensitivity Score: share of items on which the judge returns the same verdict under two reworded instruction templates that ask for the same judgement, over the pairs it answered (refusals and malformed outputs left out); temperature 0 (claude-opus-4-7 at the provider default), 1,024 output tokens, on 130 MT-Bench response pairs with a clear human majority, judged in both candidate orders (260 rows); higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 25 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemma 4 31B (IT) | 1 |
| 2 | Qwen 3.6 35B A3B | 0.97 |
| 3 | Gemini 3.7 Flash | 0.97 |
| 4 | Llama 3.3 70B Instruct | 0.97 |
| 5 | Qwen 2.5 72B Instruct | 0.97 |
| 6 | Kimi K3 | 0.97 |
| 7 | Qwen 3 32B | 0.95 |
| 8 | Llama 4 Maverick Instruct FP8 | 0.94 |
| 9 | Gemini 2.5 Flash | 0.94 |
| 10 | DeepSeek V4 Pro (0813) | 0.92 |
| 11 | Llama 3.1 8B Instruct | 0.91 |
| 12 | Qwen 3.8 27B | 0.91 |
| 13 | Claude Haiku 4.5 (20251001) | 0.9 |
| 14 | DeepSeek V4 Flash (0731) | 0.89 |
| 15 | GLM-5.2 | 0.87 |
Interactive version: theaggregate.ai/benchmark?slug=judgesense-preference · How It Works · Data refreshed daily, snapshot 2026-10-07.