HELM MMLU - Moral Scenarios: leaderboard

Metric: EM. Source: crfm.stanford.edu. 79 models tracked.

Top models

#ModelScore
1GPT-4 (0613)90.17
2Claude 3.5 Sonnet (20241022)88.83
3Claude 3.5 Sonnet (20240620)88.16
4Llama 3.1 405B Instruct87.6
5Llama 3.2 90B Vision Instruct84.13
6GPT-4o (2024-05-13)84.13
7Mistral Large 2 (Jul)83.91
8Llama 3.1 70B Instruct83.35
9Yi Large (Preview)83.13
10Claude 3 Opus (20240229)82.57
11GPT-4 Turbo (Preview)81.56
12Gemini 2.0 Flash (Preview)81.45
13Qwen 2 72B Instruct81.45
14DeepSeek V380.78
15GPT-4 Turbo80.34

Interactive version: theaggregate.ai/benchmark?slug=helm-mmlu-moral-scenarios · How It Works · Data refreshed daily, snapshot 2026-09-08.