GSM8K — leaderboard
Grade School Math 8K: 8,500 grade-school math word problems requiring multi-step arithmetic reasoning. A foundational math benchmark, now largely saturated by frontier models.
Metric: Accuracy (%). Source: epoch.ai. Status: saturated. 93 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Human Expert | 100 |
| 2 | Median Human | 95 |
| 3 | Qwen 2.5 Coder 14B Instruct | 94.2 |
| 4 | Qwen 2.5 Coder 32B Instruct | 93 |
| 5 | GPT-4 (0314) | 92 |
| 6 | GPT-4o Mini (2024-07-18) | 91.3 |
| 7 | Qwen 2.5 Coder 32B | 91.1 |
| 8 | GPT-4 (0613) | 89.99 |
| 9 | Phi-3.5-MoE-instruct | 88.7 |
| 10 | Qwen 2.5 Coder 14B | 88.7 |
| 11 | DeepSeek Coder V2 Lite Instruct | 87.6 |
| 12 | Qwen 2.5 Coder 7B Instruct | 86.7 |
| 13 | Claude Instant 1.2 | 86.7 |
| 14 | Phi-3.5-mini-instruct | 86.2 |
| 15 | Gemma 2 9B | 84.9 |
Interactive version: theaggregate.ai/benchmark?slug=gsm8k · How the rankings work · Data refreshed daily, snapshot 2026-07-22.