HELM — leaderboard

HELM: Evaluates broad language-model knowledge, reasoning, commonsense, instruction following, or exam-style accuracy.

Metric: Mean win rate (self-reported). Source: benchmarklist.com. Status: saturation imminent. 65 models tracked.

Top models

#ModelScore
1Llama 2 70B94.35
2LLaMA-65B90.83
3Mistral-7B-v0.188.4
4text-davinci-00387.16
5Llama 2 13B82.3
6GPT-3.5 Turbo78.3
7LLaMA-30B78.13
8falcon-40B72.94
9mpt-30B71.45
10Llama 2 7B60.73
11LLaMA-13B59.47
12davinci53.77
13LLaMA-7B53.27
14alpaca-7B38.09
15falcon-7B37.83

Interactive version: theaggregate.ai/benchmark?slug=helm · How the rankings work · Data refreshed daily, snapshot 2026-07-22.