HELM Classic - CivilComments — leaderboard

Metric: Exact Match (%). Source: crfm.stanford.edu. 67 models tracked.

Top models

#ModelScore
1GPT-3.5 Turbo (0613)69.64
2text-davinci-00368.44
3GPT-3.5 Turbo (0301)67.41
4text-davinci-00266.84
5LLaMA-65B65.46
6Llama 2 70B65.19
7Mistral-7B-v0.162.44
8LLaMA-13B59.98
9mpt-30B59.9
10Llama 2 13B58.77
11alpaca-7B56.58
12LLaMA-7B56.28
13Llama 2 7B56.16
14falcon-40B55.2
15LLaMA-30B54.93

Interactive version: theaggregate.ai/benchmark?slug=helm-classic-civilcomments · How the rankings work · Data refreshed daily, snapshot 2026-07-22.