LLM-AggreFact — leaderboard

Aggregated factual consistency benchmark testing grounded factuality checking across 11 datasets including CNN, XSum, RAGTruth, and more.

Metric: Balanced Accuracy (%). Source: llm-aggrefact.github.io. Status: saturation imminent. 39 models tracked.

Top models

#ModelScore
1Claude 3.5 Sonnet77.21
2Mistral Large 2 (Jul)76.48
3GPT-4 Turbo (Preview)76.16
4GPT-4o (2024-05-13)75.92
5Qwen 2.5 72B Instruct75.57
6Llama 3.1 70B Instruct75.15
7Claude 3 Opus74.84
8Llama 3.3 70B Instruct74.55
9Llama 3.1 405B Instruct74.37
10Mistral Large74.17
11GPT-4o Mini (2024-07-18)74.05
12Llama 3 70B Instruct73.74
13Qwen 2.5 7B Instruct72.83
14QwQ 32B-Preview71.77
15Mixtral 8x22B71.49

Interactive version: theaggregate.ai/benchmark?slug=llm-aggrefact · How the rankings work · Data refreshed daily, snapshot 2026-07-22.