VMMU - Mathematics: leaderboard

Metric: Accuracy (%) on the VMMU Mathematics questions (Vietnamese page-image multiple choice); higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 17 models tracked.

Top models

#ModelScore
1Gemini 3 Pro92.54
2GPT-5 (Medium)91.23
3O3 (Medium)84.87
4Claude Opus 4.6 (Thinking)83.33
5Gemini 2.5 Flash82.46
6Claude Sonnet 464.25
7Qwen 2.5 VL 32B Instruct61.18
8Llama 4 Scout59.65
9Qwen 2.5 VL 72B Instruct58.33
10Gemma 3 27B47.37
11GPT-4.146.27
12Mistral Medium 344.74
13Llama 4 Maverick42.54
14Mistral Small 3.234.65
15Gemma 3 4B26.32

Interactive version: theaggregate.ai/benchmark?slug=vmmu-mathematics · How It Works · Data refreshed daily, snapshot 2026-09-29.