JudgeSense - Relevance: leaderboard

Metric: Judge Sensitivity Score: share of items on which the judge returns the same verdict under two reworded instruction templates that ask for the same judgement, over the pairs it answered (refusals and malformed outputs left out); temperature 0 (claude-opus-4-7 at the provider default), 1,024 output tokens, on 250 BEIR TREC-COVID topics, each with a fully relevant passage and a human-rejected one, judged in both candidate orders (500 rows); higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 25 models tracked.

Top models

#ModelScore
1Gemini 3.7 Flash0.98
2Kimi K30.97
3Claude Opus 4.70.96
4Llama 4 Maverick Instruct FP80.96
5DeepSeek V4 Pro (0813)0.96
6Qwen 3 32B0.95
7GLM-5.20.95
8Qwen 3.8 27B0.95
9DeepSeek V4 Flash (0731)0.95
10Qwen 3.6 35B A3B0.94
11Qwen 2.5 72B Instruct0.94
12Qwen 3 14B0.93
13Claude Haiku 4.5 (20251001)0.93
14Llama 3.1 8B Instruct0.9
15Llama 3.3 70B Instruct0.89

Interactive version: theaggregate.ai/benchmark?slug=judgesense-relevance · How It Works · Data refreshed daily, snapshot 2026-10-07.