MATH Level 5 — leaderboard

The hardest tier (Level 5) of the Hendrycks MATH dataset - 1,324 competition-level problems spanning algebra, number theory, geometry, and combinatorics. Evaluated by Epoch AI.

Metric: Accuracy (%). Source: epoch.ai. Status: saturated. 108 models tracked.

Top models

#ModelScore
1Human Expert100
2GPT-5 (High)98.13
3GPT-5 (Medium)97.92
4GPT-5 Mini (High)97.85
5O4 Mini (2025-04-16) (High)97.83
6O3 (2025-04-16) (High)97.77
7Claude Sonnet 4.5 (Thinking 32K)97.73
8Qwen 3 Max (2025-09-23)97.13
9GPT-5 Mini (Medium)96.79
10DeepSeek R1 052896.64
11O3 Mini (2025-01-31) (High)96.49
12Gemini 2.5 Pro (Preview 05-06)95.9
13Gemini 2.5 Pro (Preview 03-25)95.56
14GPT-5 Nano (Medium)95.24
15O3 Mini (Medium)95.17

Interactive version: theaggregate.ai/benchmark?slug=math-level-5 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.