TriviaQA — leaderboard
Large-scale reading comprehension dataset with 95K question-answer pairs authored by trivia enthusiasts. Tests factual knowledge retrieval and reasoning over evidence documents.
Metric: Accuracy (%). Source: epoch.ai. Status: saturated. 39 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Llama 2 70B Base | 87.6 |
| 2 | Claude 2 | 87.5 |
| 3 | LLaMA-65B | 86 |
| 4 | GPT-3.5 Turbo (1106) | 85.8 |
| 5 | GPT-4 (0613) | 84.8 |
| 6 | DeepSeek V3 | 82.9 |
| 7 | Llama 3.1 405B | 82.7 |
| 8 | Mixtral 8x7B (v0.1) | 82.2 |
| 9 | DeepSeek V2 | 80 |
| 10 | falcon-40B | 79.9 |
| 11 | Median Human | 79.7 |
| 12 | Llama 2 13B | 79.6 |
| 13 | Claude Instant 1.1 | 78.9 |
| 14 | Claude Instant 1.2 | 78.7 |
| 15 | LLaMA-13B | 77.9 |
Interactive version: theaggregate.ai/benchmark?slug=triviaqa · How the rankings work · Data refreshed daily, snapshot 2026-07-22.