Vellum - MATH — leaderboard

Vellum's independently-run MATH evaluation. Competition-level math problems across algebra, geometry, number theory, and more.

Metric: Accuracy (%). Source: www.vellum.ai. Status: saturated. 21 models tracked.

Top models

#ModelScore
1Kimi K2.598
2O3 Mini97.9
3Claude Sonnet 4.697.8
4Claude Opus 4.697.6
5DeepSeek R197.3
6O196.4
7Claude 3.7 Sonnet (Thinking)96.2
8DeepSeek V3 (0324)94
9O1 Mini90
10Gemini 2.0 Flash89.7
11Gemma 3 27B89
12Claude 3.7 Sonnet82.2
13Claude 3.5 Sonnet78
14Llama 3.3 70B77
15Nova Pro76.6

Interactive version: theaggregate.ai/benchmark?slug=vellum-math · How the rankings work · Data refreshed daily, snapshot 2026-07-22.