HELM MMLU - Professional Psychology: leaderboard

Metric: EM. Source: crfm.stanford.edu. 79 models tracked.

Top models

#ModelScore
1Claude 3.5 Sonnet (20241022)92.16
2Claude 3.5 Sonnet (20240620)92.16
3Gemini 1.5 Pro (002)91.18
4GPT-4o (2024-05-13)90.52
5Claude 3 Opus (20240229)90.36
6GPT-4o (2024-08-06)89.87
7Gemini 1.5 Pro (001)89.38
8GPT-4 (0613)89.05
9DeepSeek V388.73
10GPT-4 Turbo (Preview)88.73
11Qwen 2 72B Instruct88.56
12Gemini 2.0 Flash (Preview)87.58
13GPT-4 Turbo87.25
14Llama 3 70B87.09
15Qwen 2.5 72B Instruct86.44

Interactive version: theaggregate.ai/benchmark?slug=helm-mmlu-professional-psychology · How It Works · Data refreshed daily, snapshot 2026-09-08.