SEAL - Humanity's Last Exam (Text Only) — leaderboard
Metric: Score. Source: scale.com. 60 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | gemini-3.1-pro-preview (thinking high) | 47.31 |
| 2 | gpt-5.4-pro-2026-03-05 | 45.32 |
| 3 | Muse Spark | 40.92 |
| 4 | gemini-3-pro-preview | 37.72 |
| 5 | gpt-5.4-2026-03-05 (xhigh thinking) | 36.47 |
| 6 | claude-opus-4-6-thinking-max | 36.24 |
| 7 | gpt-5-pro-2025-10-06 | 33.32 |
| 8 | gpt-5.2-2025-12-11 | 28.5 |
| 9 | claude-opus-4-5-20251101-thinking | 26.32 |
| 10 | gpt-5-2025-08-07 | 26.32 |
Interactive version: theaggregate.ai/benchmark?slug=seal-humanity-s-last-exam-text-only · How the rankings work · Data refreshed daily, snapshot 2026-07-22.