AA AIME 2025 — leaderboard

Artificial Analysis independent evaluation of all 30 problems from the 2025 American Invitational Mathematics Examination (olympiad-level).

Metric: Accuracy (%). Source: artificialanalysis.ai. Status: saturated. 269 models tracked.

Top models

#ModelScore
1Human Expert100
2GPT-5.2 (xHigh)99
3GPT-5 Codex (High)98.67
4GPT-5.2 (Medium)96.67
5DeepSeek V3.2 Speciale96.67
6MiMo-V2-Flash (Reasoning)96.33
7Gemini 3 Pro (Preview) (High)95.67
8GPT-5.1 Codex (High)95.67
9GLM-4.7 (Reasoning)95
10Kimi K2 (Thinking)94.67
11KAT-Coder-Pro V194.67
12GPT-5 (High)94.33
13GPT-5.1 (High)94
14GPT-OSS-120B (High)93.44
15Grok 492.67

Interactive version: theaggregate.ai/benchmark?slug=aa-aime-2025 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.