MGSM — leaderboard

A multilingual benchmark for mathematical questions.

Metric: Score (self-reported). Source: vals.ai. Status: saturated. 66 models tracked.

Top models

#ModelScore
1Claude Opus 4.5 Anthropic95.2
2Claude Opus 4.1 Anthropic94.44
3Claude Sonnet 4.5 Anthropic94.33
4GPT-5.2 x-highOpenAI94
5Gemini 3 highGoogle93.93
6Claude Opus 4 Anthropic93.78
7o4 Mini highOpenAI93.42
8Gemini 3 Flash Preview highGoogle93.31
9Claude Sonnet 4 Anthropic93.02
10Claude 3.7 Sonnet (thinking) Anthropic92.98

Interactive version: theaggregate.ai/benchmark?slug=mgsm · How the rankings work · Data refreshed daily, snapshot 2026-07-22.