HELM: leaderboard

HELM: Evaluates broad language-model knowledge, reasoning, commonsense, instruction following, or exam-style accuracy.

Metric: Mean win rate (self-reported). Source: benchmarklist.com. Status: saturated. 66 models tracked.

Top models

#ModelScore
1Llama 2 70B94.35
2LLaMA-65B90.83
3text-davinci-00290.5
4Mistral-7B-v0.188.4
5text-davinci-00387.16
6Llama 2 13B82.3
7GPT-3.5 Turbo (0613)78.3
8LLaMA-30B78.13
9GPT-3.5 Turbo (0301)76.03
10falcon-40B72.94
11mpt-30B71.45
12Llama 2 7B60.73
13LLaMA-13B59.47
14davinci53.77
15LLaMA-7B53.27

Interactive version: theaggregate.ai/benchmark?slug=helm · How It Works · Data refreshed daily, snapshot 2026-09-05.