HELM NarrativeQA: leaderboard

Reading comprehension on full stories and movie scripts requiring free-form answers, scored by F1 overlap. Part of the Stanford HELM evaluation suite.

Metric: F1 (%). Source: crfm.stanford.edu. Status: saturated. 91 models tracked.

Top models

#ModelScore
1GPT-4o (2024-05-13)80.39
2Llama 3 70B79.82
3DeepSeek V379.63
4GPT-4o (2024-08-06)79.54
5Llama 3.3 70B Instruct79.13
6Nova Pro79.13
7Gemma 2 27B (IT)78.97
8Gemini 1.5 Flash (001)78.29
9Gemini 1.5 Pro (001)78.25
10Yi 34B (Base)78.22
11Mistral Large 2 (Jul)77.87
12Mixtral 8x22B77.87
13Llama 3.2 90B Vision Instruct77.69
14Palmyra-X-00477.26
15Llama 3.1 70B Instruct77.24

Interactive version: theaggregate.ai/benchmark?slug=helm-narrativeqa · How It Works · Data refreshed daily, snapshot 2026-09-05.