MathArena - ArXiv Math Dec 2025 — leaderboard

Metric: Accuracy (%). Source: matharena.ai. 20 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)66.18
2GPT-5.4 (xHigh)60.29
3Claude Opus 4.6 (High)57.35
4GPT-5.2 (High)52.21
5Gemini 3 Pro (Preview)51.1
6Grok 4.1 Fast (Reasoning)50
7Gemini 3 Flash43.38
8Step 3.5 Flash41.91
9Kimi K2.5 (Thinking)41.91
10DeepSeek V3.2 (Thinking)41.54
11Qwen 3.5 27B41.18
12Qwen 3.5 9B39.71
13Qwen 3.5 35B A3B39.71
14GLM-538.24
15Qwen 3.5 397B A17B38.24

Interactive version: theaggregate.ai/benchmark?slug=matharena-arxiv-math-dec-2025 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.