MathArena - APEX 2025 — leaderboard

Metric: Accuracy (%). Source: matharena.ai. 47 models tracked.

Top models

#ModelScore
1Claude Opus 4.8 (Max)81.25
2GPT-5.580.21
3GPT-5.5 (xHigh)80.21
4GPT-5.4 Pro (xHigh)69.79
5Gemini 3.1 Pro (Preview)60.94
6GPT-5.454.17
7GPT-5.4 (xHigh)54.17
8Qwen 3.7 Max (Max)44.5
9Claude Opus 4.740.62
10Claude Opus 4.7 (xHigh)40.62
11Nex N2 Pro36.5
12Claude Opus 4.6 (Max)34.5
13Claude Opus 4.6 (High)34.45
14Gemini 3.5 Flash32.29
15DeepSeek V4 Pro28.12

Interactive version: theaggregate.ai/benchmark?slug=matharena-apex-2025 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.