Humanity's Last Exam: leaderboard

PhD-level questions at the frontier of human expert knowledge. Made by the same team behind MASK. Even top models show high calibration errors, indicating reasoning gaps.

Metric: Score. Source: scale.com. Status: years away from saturation. 51 models tracked.

Top models

#ModelScore
1Claude Fable 5.1 (xHigh)46.5
2Gemini 3.1 Pro (Preview) (High)46.44
3GPT-5.4 Pro (xHigh)44.32
4Muse Spark40.56
5Gemini 3 Pro (Preview)37.52
6GPT-5.4 (xHigh)36.24
7Claude Opus 4.736.2
8Claude Opus 4.6 (Adaptive Reasoning, Max Effort)34.44
9GPT-5 Pro31.64
10GPT-5.227.8
11GPT-525.32
12Claude Opus 4.5 (20251101) (Thinking)25.2
13Kimi K2.524.37
14GPT-5.1 (Thinking)23.68
15Gemini 2.5 Pro (Preview 06-05)21.64

Interactive version: theaggregate.ai/benchmark?slug=humanity-s-last-exam · How It Works · Data refreshed daily, snapshot 2026-09-05.