HELM MMLU - Overall: leaderboard

Metric: EM. Source: crfm.stanford.edu. 79 models tracked.

Top models

#ModelScore
1Claude 3.5 Sonnet (20241022)87.27
2DeepSeek V387.2
3Gemini 1.5 Pro (002)86.88
4Claude 3.5 Sonnet (20240620)86.54
5Claude 3 Opus (20240229)84.57
6Llama 3.1 405B Instruct84.53
7GPT-4o (2024-08-06)84.33
8GPT-4o (2024-05-13)84.17
9Qwen 2.5 72B Instruct83.37
10Gemini 1.5 Pro (001)82.71
11GPT-4 (0613)82.44
12Qwen 2 72B Instruct82.35
13Nova Pro82
14GPT-4 Turbo81.27
15Llama 3.2 90B Vision Instruct80.27

Interactive version: theaggregate.ai/benchmark?slug=helm-mmlu-overall · How It Works · Data refreshed daily, snapshot 2026-09-08.