FrontierMath: leaderboard

Epoch AI's expert-crafted math benchmark featuring original, unpublished problems that test genuine mathematical reasoning rather than memorization.

Metric: Accuracy (%, 285 private v2 problems). Source: epoch.ai. Status: years away from saturation. 101 models tracked.

Top models

#ModelScore
1GPT-6 (Max)93.68
2Claude Fable 5.1 (Max)90.18
3GPT-5.6 Sol (Max)89.12
4GPT-5.5 Pro (xHigh)87.72
5Claude Fable 5 (Max)87.02
6GPT-5.6 Terra (Max)85.96
7Claude Opus 5 (Max)85.61
8GPT-5.5 (xHigh)85.26
9GPT-5.4 Pro (xHigh)82.46
10GPT-5.6 Luna (Max)82.11
11Claude Opus 4.8 (Max)80
12GPT-5.4 (xHigh)78.6
13Qwen 3.8 Max (xHigh)74.74
14Kimi K3 (Max)72.18
15Gemini 3.7 Flash (High)71.58

Interactive version: theaggregate.ai/benchmark?slug=frontiermath · How It Works · Data refreshed daily, snapshot 2026-09-05.