MathArena - ARXIV April — leaderboard

Metric: Accuracy (%). Source: matharena.ai. 15 models tracked.

Top models

#ModelScore
1Claude Fable 5 (Max)70.73
2GPT-5.5 (xHigh)67.07
3Gemini 3.1 Pro (Preview)62.2
4Claude Opus 4.8 (Max)60.98
5Claude Opus 4.7 (xHigh)58.54
6DeepSeek V4 Pro (Max)52.44
7Gemini 3.5 Flash51.22
8DeepSeek V4 Flash (Max)48.78
9GLM-5.247.97
10Step 3.7 Flash37.4
11GLM-5.136.59
12Qwen 3.6 35B A3B34.15
13Step 3.5 Flash34.15
14Qwen 3.5 2B2.44

Interactive version: theaggregate.ai/benchmark?slug=matharena-arxiv-april · How the rankings work · Data refreshed daily, snapshot 2026-07-22.