Humanity's Last Exam (Text Only): leaderboard

Humanity's Last Exam text-only leaderboard evaluates frontier LLMs using text-based expert questions, excluding multimodal content.

Metric: Score. Source: scale.com. Status: saturation imminent. 61 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview) (High)47.31
2Claude Fable 5.1 (xHigh)46.8
3GPT-5.4 Pro (xHigh)45.32
4Muse Spark40.92
5Gemini 3 Pro (Preview)37.72
6GPT-5.4 (xHigh)36.47
7Claude Opus 4.6 (Adaptive Reasoning, Max Effort)36.24
8GPT-5 Pro33.32
9GPT-5.228.5
10GPT-526.32
11Claude Opus 4.5 (20251101) (Thinking)26.32
12GPT-5.1 (Thinking)24.65
13Gemini 2.5 Pro (Preview 06-05)22.06
14O3 (High)20.57
15O3 (Medium)19.78

Interactive version: theaggregate.ai/benchmark?slug=humanity-s-last-exam-text-only · How It Works · Data refreshed daily, snapshot 2026-09-05.