HELM Classic - NaturalQuestions Open Book — leaderboard

Metric: F1 (%). Source: crfm.stanford.edu. 66 models tracked.

Top models

#ModelScore
1text-davinci-00377.02
2text-davinci-00271.32
3Mistral-7B-v0.168.66
4falcon-40B67.53
5GPT-3.5 Turbo (0613)67.48
6Llama 2 70B67.42
7mpt-30B67.29
8LLaMA-65B67.21
9LLaMA-30B66.56
10Llama 2 13B63.73
11davinci62.46
12GPT-3.5 Turbo (0301)62.43
13LLaMA-13B61.43
14Llama 2 7B61.13
15gpt-neox-20B59.61

Interactive version: theaggregate.ai/benchmark?slug=helm-classic-naturalquestions-open-book · How the rankings work · Data refreshed daily, snapshot 2026-07-22.