FactArena - Justification Helpfulness: leaderboard

Metric: Elo rating from pairwise judged comparisons. Source: arxiv.org. Saturation forecast: Around September 2026. 16 models tracked.

Top models

#ModelScore
1O3 (2025-04-16)1294.67
2DeepSeek R11141.88
3O4 Mini (2025-04-16)1094.46
4Gemini 2.5 Pro (Preview 06-05)1085.29
5GPT-4.5 (Preview)1052.08
6GPT-4.11044
7Grok 31027
8Grok 3 Mini Beta965.29
9Gemini 2.5 Flash962.62
10Claude Opus 4 (20250514)953.03
11DeepSeek V3937.19
12Llama 4 Maverick Instruct922.33
13Claude Sonnet 4 (20250514)905.29
14Qwen 3 235B A22B893.44
15GPT-4o888.45

Interactive version: theaggregate.ai/benchmark?slug=factarena-justification-helpfulness · How It Works · Data refreshed daily, snapshot 2026-09-25.