Vellum - AIME 2025 — leaderboard
Vellum's independently-run AIME 2025 evaluation. American Invitational Mathematics Examination problems.
Metric: Accuracy (%). Source: www.vellum.ai. Status: saturated. 28 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.2 | 100 |
| 2 | Gemini 3 Pro | 100 |
| 3 | Human Expert | 100 |
| 4 | Claude Opus 4.6 | 99.8 |
| 5 | Kimi K2 (Thinking) | 99.1 |
| 6 | GPT-OSS-20B | 98.7 |
| 7 | O3 | 98.4 |
| 8 | GPT-OSS-120B | 97.9 |
| 9 | Claude Haiku 4.5 | 96.3 |
| 10 | Kimi K2.5 | 96.1 |
| 11 | GPT-5.1 | 94 |
| 12 | Grok 3 | 93.3 |
| 13 | O4 Mini | 92.7 |
| 14 | Grok 4 | 91.7 |
| 15 | GPT-5 | 91 |
Interactive version: theaggregate.ai/benchmark?slug=vellum-aime-2025 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.