MathArena - ARXIV March — leaderboard

Metric: Accuracy (%). Source: matharena.ai. 15 models tracked.

Top models

#ModelScore
1GPT-5.5 (xHigh)77.5
2Claude Opus 4.8 (Max)75
3Gemini 3.1 Pro (Preview)68.33
4GPT-5.4 (xHigh)67.78
5Claude Opus 4.6 (High)62.5
6GLM-5.261.67
7Gemini 3.5 Flash55.83
8Kimi K2.6 (Thinking)55.83
9DeepSeek V4 Pro (Max)55.83
10DeepSeek V4 Flash (Max)52.5
11Claude Opus 4.7 (xHigh)50.83
12GLM-5.149.17
13GLM-545.83
14Step 3.7 Flash45
15Step 3.5 Flash43.33

Interactive version: theaggregate.ai/benchmark?slug=matharena-arxiv-march · How the rankings work · Data refreshed daily, snapshot 2026-07-22.