MathArena - ARXIV February — leaderboard

Metric: Accuracy (%). Source: matharena.ai. 26 models tracked.

Top models

#ModelScore
1GPT-5.4 Pro (xHigh)75.78
2GPT-5.5 (xHigh)73.44
3Gemini 3.1 Pro (Preview)62.5
4Claude Opus 4.8 (Max)60.16
5GPT-5.4 (xHigh)59.38
6DeepSeek V4 Pro (Max)51.56
7Gemini 3.5 Flash50.78
8GLM-5.246.09
9DeepSeek V4 Flash (Max)43.75
10Kimi K2.6 (Thinking)42.97
11GLM-541.41
12Claude Opus 4.6 (High)40.62
13Claude Opus 4.7 (xHigh)40.62
14Gemini 3.1 Pro (Preview) (Low)40.62
15GLM-5.139.06

Interactive version: theaggregate.ai/benchmark?slug=matharena-arxiv-february · How the rankings work · Data refreshed daily, snapshot 2026-07-22.