FactArena - Justification Soundness: leaderboard

Metric: Elo rating from pairwise judged comparisons. Source: arxiv.org. Saturation forecast: Estimated already saturated. 16 models tracked.

Top models

#ModelScore
1O3 (2025-04-16)1292.55
2DeepSeek R11131.83
3O4 Mini (2025-04-16)1097.33
4Gemini 2.5 Pro (Preview 06-05)1081.98
5GPT-4.5 (Preview)1053.24
6GPT-4.11044.03
7Grok 31026.05
8Grok 3 Mini Beta958.97
9Claude Opus 4 (20250514)958.32
10Gemini 2.5 Flash956.76
11DeepSeek V3938.83
12Claude Sonnet 4 (20250514)919.29
13Llama 4 Maverick Instruct916.31
14GPT-4o891.96
15Qwen 3 235B A22B891.82

Interactive version: theaggregate.ai/benchmark?slug=factarena-justification-soundness · How It Works · Data refreshed daily, snapshot 2026-09-25.