FrontierMath - Tier 4 — leaderboard

Metric: Accuracy (%, 48 problems). Source: epoch.ai. 72 models tracked.

Top models

#ModelScore
1GPT-5.5 Pro (xHigh)39.6
2GPT-5.4 Pro (xHigh)37.5
3GPT-5.5 (xHigh)35.4
4GPT-5.2 Pro31.3
5Claude Opus 4.8 (Max)31.25
6GPT-5.4 (xHigh)27.1
7Claude Opus 4.7 (xHigh)22.92
8Claude Opus 4.6 (Max)22.9
9Claude Opus 4.6 (Thinking 32K)20.83
10GPT-5.2 (High)18.8
11GPT-5.2 (xHigh)18.8
12Gemini 3 Pro (Preview)18.75
13Gemini 3.1 Pro (Preview)16.7
14GPT-5.2 (Medium)16.7
15Muse Spark14.6

Interactive version: theaggregate.ai/benchmark?slug=frontiermath-tier-4 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.