HELM Classic - TruthfulQA — leaderboard

Metric: Exact Match (%). Source: crfm.stanford.edu. 67 models tracked.

Top models

#ModelScore
1text-davinci-00260.96
2GPT-3.5 Turbo (0301)60.86
3text-davinci-00359.33
4Llama 2 70B55.35
5LLaMA-65B50.76
6Mistral-7B-v0.142.2
7falcon-40B35.32
8LLaMA-30B34.4
9GPT-3.5 Turbo (0613)33.94
10Llama 2 13B33.03
11LLaMA-13B32.42
12LLaMA-7B27.98
13Llama 2 7B27.22
14text-curie-00125.74
15alpaca-7B24.31

Interactive version: theaggregate.ai/benchmark?slug=helm-classic-truthfulqa · How the rankings work · Data refreshed daily, snapshot 2026-07-22.