FactArena - Justification Readability: leaderboard

Metric: Elo rating from pairwise judged comparisons. Source: arxiv.org. Saturation forecast: Estimated already saturated. 16 models tracked.

Top models

#ModelScore
1O3 (2025-04-16)1268.83
2DeepSeek R11117.53
3O4 Mini (2025-04-16)1097.51
4Gemini 2.5 Pro (Preview 06-05)1077.44
5GPT-4.5 (Preview)1041.03
6GPT-4.11029.11
7Grok 31019.52
8Gemini 2.5 Flash980.5
9Grok 3 Mini Beta974.06
10DeepSeek V3951.61
11Claude Opus 4 (20250514)933.39
12Qwen 3 235B A22B922.65
13GPT-4o913.03
14Llama 4 Maverick Instruct909.13
15Claude Sonnet 4 (20250514)892.76

Interactive version: theaggregate.ai/benchmark?slug=factarena-justification-readability · How It Works · Data refreshed daily, snapshot 2026-09-25.