MATH-Perturb (Hard) — leaderboard

Perturbed versions of MATH competition problems testing genuine mathematical reasoning vs. pattern matching on memorized solutions.

Metric: Accuracy (%). Source: math-perturb.github.io. Status: saturated. 37 models tracked.

Top models

#ModelScore
1Grok 487.1
2Grok 3 Mini (High)87.1
3O4 Mini (Medium)87.1
4QwQ-32B86.02
5DeepSeek R185.19
6O3 Mini (High)84.35
7O3 Mini (Medium)82.92
8DeepSeek R1 Distill Qwen 14B81.84
9O1 Mini79.69
10DeepSeek R1 Distill Llama 70B79.57
11Grok 379.21
12O3 Mini (Low)78.49
13O177.18
14Claude 3.7 Sonnet76.11
15O3 (Medium)75.63

Interactive version: theaggregate.ai/benchmark?slug=math-perturb-hard · How the rankings work · Data refreshed daily, snapshot 2026-07-22.