MathArena - ARXIV_FALSE February — leaderboard

Metric: Accuracy (%). Source: matharena.ai. 17 models tracked.

Top models

#ModelScore
1GPT-5.5 (xHigh)68.15
2GPT-5.4 (xHigh)37.9
3Claude Opus 4.8 (Max)35.48
4GPT-5.2 (xHigh)25.81
5Gemini 3.1 Pro (Preview)18.55
6DeepSeek V4 Flash (Max)17.74
7DeepSeek V4 Pro (Max)13.31
8Step 3.7 Flash12.1
9GLM-511.69
10Kimi K2.6 (Thinking)11.69
11Step 3.5 Flash11.29
12GLM-5.210.48
13GLM-5.19.27
14Gemini 3.5 Flash6.45
15Claude Opus 4.7 (xHigh)4.03

Interactive version: theaggregate.ai/benchmark?slug=matharena-arxiv-false-february · How the rankings work · Data refreshed daily, snapshot 2026-07-22.