HELM NaturalQuestions (Closed) — leaderboard

Closed-book question answering from Google search queries without passage access, testing parametric knowledge recall. Part of Stanford HELM.

Metric: F1 (%). Source: crfm.stanford.edu. Status: saturation imminent. 91 models tracked.

Top models

#ModelScore
1Claude 3.5 Sonnet (20240620)50.16
2GPT-4o (2024-05-13)50.13
3GPT-4o (2024-08-06)49.58
4GPT-4 Turbo48.21
5Mixtral 8x22B47.76
6Llama 3 70B47.53
7Claude 3.5 Sonnet (20241022)46.7
8DeepSeek V346.7
9Llama 2 70B45.96
10GPT-4 (0613)45.69
11Llama 3.2 90B Vision Instruct45.68
12Palmyra-X-00445.66
13Llama 3.1 405B Instruct45.6
14Gemini 1.5 Pro (002)45.55
15Mistral Large 2 (Jul)45.27

Interactive version: theaggregate.ai/benchmark?slug=helm-naturalquestions-closed · How the rankings work · Data refreshed daily, snapshot 2026-07-22.