HELM (Stanford) — leaderboard

Holistic Evaluation of Language Models across 42 scenarios and 7 metrics: accuracy, fairness, bias, toxicity, efficiency, robustness, and calibration.

Metric: Mean Win Rate (%). Source: crfm.stanford.edu. Status: saturated. 91 models tracked.

Top models

#ModelScore
1GPT-4o (2024-05-13)93.76
2GPT-4o (2024-08-06)92.76
3DeepSeek V390.83
4Claude 3.5 Sonnet (20240620)88.5
5Nova Pro88.48
6GPT-4 (0613)86.71
7GPT-4 Turbo86.41
8Llama 3.1 405B Instruct85.39
9Claude 3.5 Sonnet (20241022)84.61
10Gemini 1.5 Pro (002)84.19
11Llama 3.2 90B Vision Instruct81.94
12Gemini 2.0 Flash (Preview)81.27
13Llama 3.3 70B Instruct81.22
14Llama 3.1 70B Instruct80.83
15Palmyra-X-00480.82

Interactive version: theaggregate.ai/benchmark?slug=helm-stanford · How the rankings work · Data refreshed daily, snapshot 2026-07-22.