HELM Classic - Entity Data Imputation — leaderboard

Metric: Exact Match (%). Source: crfm.stanford.edu. 67 models tracked.

Top models

#ModelScore
1mpt-30B85.52
2LLaMA-30B84.36
3Llama 2 13B84.36
4text-davinci-00284.16
5text-davinci-00383.91
6Llama 2 70B83.77
7GPT-3.5 Turbo (0613)83.59
8davinci83.59
9LLaMA-7B83.4
10falcon-40B83.01
11LLaMA-13B82.42
12LLaMA-65B82.24
13Llama 2 7B80.1
14Mistral-7B-v0.179.52
15text-curie-00179.13

Interactive version: theaggregate.ai/benchmark?slug=helm-classic-entity-data-imputation · How the rankings work · Data refreshed daily, snapshot 2026-07-22.