JudgeSense - Factuality: leaderboard

Metric: Judge Sensitivity Score: share of items on which the judge returns the same verdict under two reworded instruction templates that ask for the same judgement, over the pairs it answered (refusals and malformed outputs left out); temperature 0 (claude-opus-4-7 at the provider default), 1,024 output tokens, on 250 TruthfulQA statements judged accurate or inaccurate; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 25 models tracked.

Top models

#ModelScore
1Gemma 4 31B (IT)0.98
2Gemini 3.7 Flash0.98
3Claude Opus 4.70.96
4Claude Haiku 4.5 (20251001)0.96
5Llama 3.3 70B Instruct0.95
6Claude Sonnet 4.50.95
7Qwen 3 32B0.95
8Kimi K30.95
9Qwen 3.6 35B A3B0.93
10Llama 4 Maverick Instruct FP80.92
11Qwen 2.5 72B Instruct0.91
12DeepSeek V4 Pro (0813)0.9
13Gemini 2.5 Flash0.9
14Llama 4 Scout Instruct0.89
15Qwen 3 14B0.89

Interactive version: theaggregate.ai/benchmark?slug=judgesense-factuality · How It Works · Data refreshed daily, snapshot 2026-10-07.