JudgeSense - Factuality: leaderboard
Metric: Judge Sensitivity Score: share of items on which the judge returns the same verdict under two reworded instruction templates that ask for the same judgement, over the pairs it answered (refusals and malformed outputs left out); temperature 0 (claude-opus-4-7 at the provider default), 1,024 output tokens, on 250 TruthfulQA statements judged accurate or inaccurate; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 25 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemma 4 31B (IT) | 0.98 |
| 2 | Gemini 3.7 Flash | 0.98 |
| 3 | Claude Opus 4.7 | 0.96 |
| 4 | Claude Haiku 4.5 (20251001) | 0.96 |
| 5 | Llama 3.3 70B Instruct | 0.95 |
| 6 | Claude Sonnet 4.5 | 0.95 |
| 7 | Qwen 3 32B | 0.95 |
| 8 | Kimi K3 | 0.95 |
| 9 | Qwen 3.6 35B A3B | 0.93 |
| 10 | Llama 4 Maverick Instruct FP8 | 0.92 |
| 11 | Qwen 2.5 72B Instruct | 0.91 |
| 12 | DeepSeek V4 Pro (0813) | 0.9 |
| 13 | Gemini 2.5 Flash | 0.9 |
| 14 | Llama 4 Scout Instruct | 0.89 |
| 15 | Qwen 3 14B | 0.89 |
Interactive version: theaggregate.ai/benchmark?slug=judgesense-factuality · How It Works · Data refreshed daily, snapshot 2026-10-07.