MathBench — leaderboard

Hierarchical math evaluation from arithmetic through college level, testing across 5 progressive difficulty stages with both theoretical and application questions.

Metric: Average Score. Source: github.com. Status: saturated. 34 models tracked.

Top models

#ModelScore
1GPT-4o (2024-05-13)70.9
2Qwen 2 72B Instruct67
3Claude 3 Opus63
4GPT-4 Preview (0125)58.8
5Qwen 1.5 110B Chat58.4
6Llama 3 70B Instruct56.4
7Qwen 2 7B Instruct53.4
8Yi 1.5 34B Chat52.2
9Yi-1.5-9B Chat51.1
10GLM-4 9B Chat47.7
11internlm2-chat-20B42.1
12GPT-3.5 Turbo (0125)41
13Qwen-14B-Chat39.5
14Llama 3 8B Instruct36.7
15internlm2-chat-7B34.1

Interactive version: theaggregate.ai/benchmark?slug=mathbench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.