μ-MATH — leaderboard

Metric: Accuracy (%). Source: huggingface.co. 18 models tracked.

Top models

#ModelScore
1O189.5
2O1 Mini84.8
3QwQ 32B-Preview83.3
4DeepSeek R182.2
5Gemini 2.0 Flash (01-21) (Thinking)81.2
6Gemini 1.5 Pro80.7
7GPT-4o (2024-08-06)77.4
8Mistral Large 2 (Nov) Instruct (2411)76.7
9Qwen 2.5 72B Instruct75.7
10Claude 3.5 Sonnet75
11Gemini 1.5 Flash74.9
12Qwen2.5-Math-72B-Instruct74.4
13GPT-4o Mini (2024-07-18)72.5
14Qwen 2.5 7B Instruct69.9
15Qwen2.5-Math-7B-Instruct63.3

Interactive version: theaggregate.ai/benchmark?slug=math · How the rankings work · Data refreshed daily, snapshot 2026-07-22.