JudgeSense - Relevance: leaderboard
Metric: Judge Sensitivity Score: share of items on which the judge returns the same verdict under two reworded instruction templates that ask for the same judgement, over the pairs it answered (refusals and malformed outputs left out); temperature 0 (claude-opus-4-7 at the provider default), 1,024 output tokens, on 250 BEIR TREC-COVID topics, each with a fully relevant passage and a human-rejected one, judged in both candidate orders (500 rows); higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 25 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3.7 Flash | 0.98 |
| 2 | Kimi K3 | 0.97 |
| 3 | Claude Opus 4.7 | 0.96 |
| 4 | Llama 4 Maverick Instruct FP8 | 0.96 |
| 5 | DeepSeek V4 Pro (0813) | 0.96 |
| 6 | Qwen 3 32B | 0.95 |
| 7 | GLM-5.2 | 0.95 |
| 8 | Qwen 3.8 27B | 0.95 |
| 9 | DeepSeek V4 Flash (0731) | 0.95 |
| 10 | Qwen 3.6 35B A3B | 0.94 |
| 11 | Qwen 2.5 72B Instruct | 0.94 |
| 12 | Qwen 3 14B | 0.93 |
| 13 | Claude Haiku 4.5 (20251001) | 0.93 |
| 14 | Llama 3.1 8B Instruct | 0.9 |
| 15 | Llama 3.3 70B Instruct | 0.89 |
Interactive version: theaggregate.ai/benchmark?slug=judgesense-relevance · How It Works · Data refreshed daily, snapshot 2026-10-07.