MathArena - ARXIV_FALSE March — leaderboard

Metric: Accuracy (%). Source: matharena.ai. 15 models tracked.

Top models

#ModelScore
1GPT-5.5 (xHigh)73.66
2GPT-5.4 (xHigh)36.61
3Claude Opus 4.8 (Max)35.71
4DeepSeek V4 Flash (Max)19.64
5Gemini 3.5 Flash16.07
6Kimi K2.6 (Thinking)15.62
7DeepSeek V4 Pro (Max)15.18
8Gemini 3.1 Pro (Preview)13.84
9GLM-512.05
10GLM-5.212.05
11Step 3.7 Flash10.71
12Step 3.5 Flash7.14
13GLM-5.16.47
14Claude Opus 4.6 (High)5.8
15Claude Opus 4.7 (xHigh)5.8

Interactive version: theaggregate.ai/benchmark?slug=matharena-arxiv-false-march · How the rankings work · Data refreshed daily, snapshot 2026-07-22.