HELM NaturalQuestions (Open) — leaderboard
Open-book question answering from Google search queries with access to Wikipedia passages, scored by F1. Part of Stanford HELM.
Metric: F1 (%). Source: crfm.stanford.edu. Status: saturation imminent. 91 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Nova Pro | 82.9 |
| 2 | Nova Lite | 81.52 |
| 3 | PaLM-2 (Bison) | 81.26 |
| 4 | GPT-4o (2024-05-13) | 80.32 |
| 5 | GPT-4 Turbo | 79.52 |
| 6 | GPT-4o (2024-08-06) | 79.26 |
| 7 | GPT-4 (0613) | 78.97 |
| 8 | Nova Micro | 77.86 |
| 9 | Qwen 1.5 32B | 77.71 |
| 10 | Qwen 2 72B Instruct | 77.58 |
| 11 | Yi 34B (Base) | 77.51 |
| 12 | Qwen 1.5 14B | 77.21 |
| 13 | DeepSeek V3 | 76.53 |
| 14 | GPT-4 Turbo (Preview) | 76.29 |
| 15 | Qwen 1.5 72B | 75.85 |
Interactive version: theaggregate.ai/benchmark?slug=helm-naturalquestions-open · How the rankings work · Data refreshed daily, snapshot 2026-07-22.