IMO-AnswerBench — leaderboard

International Mathematical Olympiad answer benchmark evaluating final-answer correctness on high-difficulty olympiad-style mathematical problems.

Metric: Score (self-reported). Source: benchmarklist.com. Status: saturation imminent. 14 models tracked.

Top models

#ModelScore
1DeepSeek V4 Flash91.1
2GLM-5.291
3Qwen 3.7 Max (Max)90
4DeepSeek V4 Pro89.8
5Nemotron 3 Ultra88.6
6Qwen 3.7 Plus86
7Qwen 3.6 Plus83.8
8GLM-5.183.8
9Claude Opus 4.883.5
10Qwen 3.5 397B A17B83.1
11Gemini 3.1 Pro (Preview)81
12Claude Opus 4.6 (Max)75.3
13MiniMax-M2.768.3

Interactive version: theaggregate.ai/benchmark?slug=imo-answerbench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.