MMLU — leaderboard

Massive Multitask Language Understanding: 15,908 questions across 57 subjects from STEM to humanities. The original broad knowledge benchmark for LLMs, now considered saturated at the frontier.

Metric: Accuracy (%). Source: epoch.ai. Status: saturated. 136 models tracked.

Top models

#ModelScore
1Human Expert90
2GPT-4o (2024-11-20)88.1
3Claude 3.5 Sonnet (20241022)87.3
4DeepSeek V387.2
5Gemini 1.5 Pro (002)86.9
6Claude 3.5 Sonnet (20240620)86.5
7GPT-4 (0314)86.4
8Llama 3.3 70B Instruct86.3
9Gemini 1.5 Pro (001)85.9
10Qwen 2.5 72B Instruct85.3
11Qwen 2.5 72B85
12Phi-484.8
13Claude 3 Opus (20240229)84.6
14Llama 3.1 405B Instruct84.5
15Llama 3.1 405B84.4

Interactive version: theaggregate.ai/benchmark?slug=mmlu · How the rankings work · Data refreshed daily, snapshot 2026-07-22.