AA MATH-500: leaderboard

Artificial Analysis independent evaluation of MATH-500: 500 competition-level math problems across algebra, geometry, number theory, and more.

Metric: Accuracy (%). Source: artificialanalysis.ai. Status: saturated. 61 models tracked.

Top models

#ModelScore
1Best Expert100
2GPT-5 (High)99.4
3O399.2
4GPT-5 (Medium)99.13
5Grok 499
6GPT-5 (Low)98.73
7Qwen 3 235B A22B 2507 (Thinking)98.4
8DeepSeek R1 052898.27
9Qwen 3 235B A22B 2507 Instruct98
10GLM-4.5 (Reasoning)97.87
11Qwen 3 30B A3B 2507 (Thinking)97.6
12Qwen 3 30B A3B 2507 Instruct97.53
13Kimi K297.13
14Gemini 2.5 Pro96.73
15GLM 4.5 Air96.53

Interactive version: theaggregate.ai/benchmark?slug=aa-math-500 · How It Works · Data refreshed daily, snapshot 2026-09-05.