SEAL - Humanity's Last Exam — leaderboard

Metric: Score. Source: scale.com. 50 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview) (High)46.44
2GPT-5.4 Pro (xHigh)44.32
3Muse Spark40.56
4Gemini 3 Pro (Preview)37.52
5GPT-5.4 (xHigh)36.24
6Claude Opus 4.736.2
7Claude Opus 4.6 (Adaptive Reasoning, Max Effort)34.44
8GPT-5 Pro31.64
9GPT-5.227.8
10GPT-525.32
11Claude Opus 4.5 (20251101) (Thinking)25.2
12Kimi K2.524.37
13GPT-5.1 (Thinking)23.68
14Gemini 2.5 Pro (Preview 06-05)21.64
15O3 (High)20.32

Interactive version: theaggregate.ai/benchmark?slug=seal-humanity-s-last-exam · How the rankings work · Data refreshed daily, snapshot 2026-07-22.