MathArena - APEX Shortlist 2025 — leaderboard

Metric: Accuracy (%). Source: matharena.ai. 37 models tracked.

Top models

#ModelScore
1GPT-5.5 (xHigh)98.4
2Gemini 3.1 Pro (Preview)92.02
3Claude Opus 4.8 (Max)90.43
4DeepSeek V4 Flash (Max)89.36
5DeepSeek V4 Pro (Max)87.77
6Claude Opus 4.6 (High)86.7
7Gemini 3.5 Flash82.45
8GPT-5.4 (xHigh)81.38
9GPT-5.2 (High)79.26
10Kimi K2.6 (Thinking)77.13
11Step 3.7 Flash76.6
12GLM-5.171.81
13Step 3.5 Flash70.74
14DeepSeek V3.2 Speciale69.68
15GLM-568.62

Interactive version: theaggregate.ai/benchmark?slug=matharena-apex-shortlist-2025 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.