LLM-AggreFact — leaderboard
Aggregated factual consistency benchmark testing grounded factuality checking across 11 datasets including CNN, XSum, RAGTruth, and more.
Metric: Balanced Accuracy (%). Source: llm-aggrefact.github.io. Status: saturation imminent. 39 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude 3.5 Sonnet | 77.21 |
| 2 | Mistral Large 2 (Jul) | 76.48 |
| 3 | GPT-4 Turbo (Preview) | 76.16 |
| 4 | GPT-4o (2024-05-13) | 75.92 |
| 5 | Qwen 2.5 72B Instruct | 75.57 |
| 6 | Llama 3.1 70B Instruct | 75.15 |
| 7 | Claude 3 Opus | 74.84 |
| 8 | Llama 3.3 70B Instruct | 74.55 |
| 9 | Llama 3.1 405B Instruct | 74.37 |
| 10 | Mistral Large | 74.17 |
| 11 | GPT-4o Mini (2024-07-18) | 74.05 |
| 12 | Llama 3 70B Instruct | 73.74 |
| 13 | Qwen 2.5 7B Instruct | 72.83 |
| 14 | QwQ 32B-Preview | 71.77 |
| 15 | Mixtral 8x22B | 71.49 |
Interactive version: theaggregate.ai/benchmark?slug=llm-aggrefact · How the rankings work · Data refreshed daily, snapshot 2026-07-22.