HELM NarrativeQA — leaderboard

Reading comprehension on full stories and movie scripts requiring free-form answers, scored by F1 overlap. Part of Stanford HELM's holistic evaluation suite.

Metric: F1 (%). Source: crfm.stanford.edu. Status: saturation imminent. 91 models tracked.

Top models

#ModelScore
1GPT-4o (2024-05-13)80.39
2Llama 3 70B79.82
3DeepSeek V379.63
4GPT-4o (2024-08-06)79.54
5Llama 3.3 70B Instruct79.13
6Nova Pro79.13
7Gemma 2 27B (IT)78.97
8Gemini 2.0 Flash (Preview)78.3
9Gemini 1.5 Flash (001)78.29
10Gemini 1.5 Pro (001)78.25
11Yi 34B (Base)78.22
12Mistral Large 2 (Jul)77.87
13Mixtral 8x22B77.87
14Llama 3.2 90B Vision Instruct77.69
15Palmyra-X-00477.26

Interactive version: theaggregate.ai/benchmark?slug=helm-narrativeqa · How the rankings work · Data refreshed daily, snapshot 2026-07-22.