Vellum - Humanity's Last Exam: leaderboard

Vellum's independently-run Humanity's Last Exam evaluation. Expert-level questions across academic domains.

Metric: Accuracy (%). Source: www.vellum.ai. Status: years away from saturation. 35 models tracked.

Top models

#ModelScore
1Claude Fable 5.165
2Claude Opus 564.7
3Claude Mythos 564.5
4Claude Opus 4.857.9
5Claude Sonnet 557.4
6GPT-657.2
7Kimi K356
8GLM-5.3 Flash55.3
9GLM-5.254.7
10Kimi K2.654
11DeepSeek V4 Flash51.6
12DeepSeek V4 Pro48.2
13GPT-5.6 Sol47.2
14Gemini 3 Pro45.8
15Kimi K2 (Thinking)44.9

Interactive version: theaggregate.ai/benchmark?slug=vellum-humanity-s-last-exam · How It Works · Data refreshed daily, snapshot 2026-09-05.