JudgeSense - Preference: leaderboard

Metric: Judge Sensitivity Score: share of items on which the judge returns the same verdict under two reworded instruction templates that ask for the same judgement, over the pairs it answered (refusals and malformed outputs left out); temperature 0 (claude-opus-4-7 at the provider default), 1,024 output tokens, on 130 MT-Bench response pairs with a clear human majority, judged in both candidate orders (260 rows); higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 25 models tracked.

Top models

#ModelScore
1Gemma 4 31B (IT)1
2Qwen 3.6 35B A3B0.97
3Gemini 3.7 Flash0.97
4Llama 3.3 70B Instruct0.97
5Qwen 2.5 72B Instruct0.97
6Kimi K30.97
7Qwen 3 32B0.95
8Llama 4 Maverick Instruct FP80.94
9Gemini 2.5 Flash0.94
10DeepSeek V4 Pro (0813)0.92
11Llama 3.1 8B Instruct0.91
12Qwen 3.8 27B0.91
13Claude Haiku 4.5 (20251001)0.9
14DeepSeek V4 Flash (0731)0.89
15GLM-5.20.87

Interactive version: theaggregate.ai/benchmark?slug=judgesense-preference · How It Works · Data refreshed daily, snapshot 2026-10-07.