MATH-Perturb (Hard): leaderboard

Perturbed versions of MATH competition problems testing genuine mathematical reasoning vs. pattern matching on memorized solutions.

Metric: Accuracy (%). Source: math-perturb.github.io. Status: saturated. 37 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro (Preview)88.89
2Grok 487.1
3Grok 3 Mini (High)87.1
4O4 Mini (Medium)87.1
5QwQ-32B86.02
6DeepSeek R185.19
7O3 Mini (High)84.35
8O3 Mini (Medium)82.92
9DeepSeek R1 Distill Qwen 14B81.84
10O1 Mini79.69
11DeepSeek R1 Distill Llama 70B79.57
12Grok 379.21
13O3 Mini (Low)78.49
14O177.18
15Claude 3.7 Sonnet76.11

Interactive version: theaggregate.ai/benchmark?slug=math-perturb-hard · How It Works · Data refreshed daily, snapshot 2026-09-05.