FactArena - Claim Extraction: leaderboard

Metric: Elo rating from pairwise judged comparisons. Source: arxiv.org. Saturation forecast: Estimated already saturated. 16 models tracked.

Top models

#ModelScore
1O3 (2025-04-16)1188.11
2DeepSeek R11096.22
3O4 Mini (2025-04-16)1081
4GPT-4.5 (Preview)1050.66
5Gemini 2.5 Pro (Preview 06-05)1046.19
6GPT-4.11018.24
7Grok 3995.75
8Grok 3 Mini Beta979.61
9Claude Opus 4 (20250514)973.2
10Gemini 2.5 Flash969.48
11DeepSeek V3968.84
12Claude Sonnet 4 (20250514)940.14
13GPT-4o937.6
14Qwen 3 235B A22B936.05
15Llama 4 Maverick Instruct919.58

Interactive version: theaggregate.ai/benchmark?slug=factarena-claim-extraction · How It Works · Data refreshed daily, snapshot 2026-09-25.