MGSM: leaderboard

A multilingual benchmark for mathematical questions.

Metric: Accuracy (%). Source: www.vals.ai. Status: saturated. 75 models tracked.

Top models

#ModelScore
1Claude Opus 4.5 (20251101) (Thinking)95.2
2Claude Opus 4.5 (20251101)94.76
3Claude Opus 4.1 (20250805) (Thinking)94.44
4Claude Sonnet 4.5 (Thinking)94.33
5Claude Sonnet 4.5 20250929 (Thinking)94.33
6Claude Opus 4.1 (20250805)94.22
7GPT-5.2 (xHigh)94
8Gemini 3 Pro (Preview) (High)93.93
9Claude Opus 493.78
10Claude Opus 4 (20250514)93.78
11O4 Mini (2025-04-16) (High)93.42
12Gemini 3 Flash (Preview) (High)93.31
13Claude Sonnet 493.02
14Claude Sonnet 4 (20250514)93.02
15GPT-5.1 (High)92.98

Interactive version: theaggregate.ai/benchmark?slug=mgsm · How It Works · Data refreshed daily, snapshot 2026-09-05.