FactArena - Justification Informativeness: leaderboard

Metric: Elo rating from pairwise judged comparisons. Source: arxiv.org. Saturation forecast: Estimated already saturated. 16 models tracked.

Top models

#ModelScore
1O3 (2025-04-16)1331.77
2DeepSeek R11160.02
3O4 Mini (2025-04-16)1101.49
4Gemini 2.5 Pro (Preview 06-05)1078.02
5GPT-4.5 (Preview)1039.95
6GPT-4.11032.77
7Grok 31026.11
8Grok 3 Mini Beta966.4
9Claude Opus 4 (20250514)961.3
10Gemini 2.5 Flash945.27
11Claude Sonnet 4 (20250514)930.52
12Llama 4 Maverick Instruct926.36
13DeepSeek V3923.45
14Qwen 3 235B A22B882.68
15GPT-4o872.55

Interactive version: theaggregate.ai/benchmark?slug=factarena-justification-informativeness · How It Works · Data refreshed daily, snapshot 2026-09-25.