TriviaQA — leaderboard

Large-scale reading comprehension dataset with 95K question-answer pairs authored by trivia enthusiasts. Tests factual knowledge retrieval and reasoning over evidence documents.

Metric: Accuracy (%). Source: epoch.ai. Status: saturated. 39 models tracked.

Top models

#ModelScore
1Llama 2 70B Base87.6
2Claude 287.5
3LLaMA-65B86
4GPT-3.5 Turbo (1106)85.8
5GPT-4 (0613)84.8
6DeepSeek V382.9
7Llama 3.1 405B82.7
8Mixtral 8x7B (v0.1)82.2
9DeepSeek V280
10falcon-40B79.9
11Median Human79.7
12Llama 2 13B79.6
13Claude Instant 1.178.9
14Claude Instant 1.278.7
15LLaMA-13B77.9

Interactive version: theaggregate.ai/benchmark?slug=triviaqa · How the rankings work · Data refreshed daily, snapshot 2026-07-22.