HELM MMLU - Moral Disputes: leaderboard

Metric: EM. Source: crfm.stanford.edu. 79 models tracked.

Top models

#ModelScore
1Claude 3.5 Sonnet (20240620)88.73
2GPT-4o (2024-05-13)88.15
3Claude 3 Opus (20240229)88.15
4DeepSeek V387.28
5Claude 3.5 Sonnet (20241022)86.99
6GPT-4o (2024-08-06)86.99
7Llama 3.1 405B Instruct86.99
8GPT-4 (0613)86.71
9Gemini 2.0 Flash (Preview)86.42
10GPT-4 Turbo86.13
11GPT-4 Turbo (Preview)85.84
12Gemini 1.5 Pro (002)85.55
13Llama 3.3 70B Instruct85.26
14Qwen 2 72B Instruct85.26
15Llama 3.2 90B Vision Instruct84.97

Interactive version: theaggregate.ai/benchmark?slug=helm-mmlu-moral-disputes · How It Works · Data refreshed daily, snapshot 2026-09-08.