AA MATH-500 — leaderboard

Artificial Analysis independent evaluation of MATH-500: 500 competition-level math problems across algebra, geometry, number theory, and more.

Metric: Accuracy (%). Source: artificialanalysis.ai. Status: saturated. 202 models tracked.

Top models

#ModelScore
1Human Expert100
2GPT-5 (High)99.4
3O399.2
4GPT-5 (Medium)99.13
5Claude Sonnet 4 (Thinking)99.07
6Grok 499
7O4 Mini (High)98.87
8GPT-5 (Low)98.73
9Gemini 2.5 Pro (Preview 05-06)98.6
10O3 Mini (High)98.47
11Qwen 3 235B A22B 2507 (Thinking)98.4
12DeepSeek R1 052898.27
13Claude Opus 4 (Thinking)98.2
14Gemini 2.5 Flash (Thinking)98.13
15Qwen 3 235B A22B 2507 Instruct98

Interactive version: theaggregate.ai/benchmark?slug=aa-math-500 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.