HELM Classic - NarrativeQA — leaderboard

Metric: F1 (%). Source: crfm.stanford.edu. 66 models tracked.

Top models

#ModelScore
1Llama 2 70B76.99
2LLaMA-65B75.49
3LLaMA-30B75.25
4Llama 2 13B74.4
5mpt-30B73.15
6text-davinci-00272.72
7text-davinci-00372.71
8Mistral-7B-v0.171.65
9LLaMA-13B71.12
10Llama 2 7B69.12
11davinci68.69
12falcon-40B67.26
13LLaMA-7B66.91
14GPT-3.5 Turbo (0301)66.3
15GPT-3.5 Turbo (0613)62.51

Interactive version: theaggregate.ai/benchmark?slug=helm-classic-narrativeqa · How the rankings work · Data refreshed daily, snapshot 2026-07-22.