Math-VR: leaderboard

Mathematical visual reasoning benchmark for VLMs, unified models, and LLMs, reporting answer correctness and process scores on text and multimodal questions.

Metric: Overall Answer Correctness (self-reported). Source: benchmarklist.com. Status: saturation imminent. 30 models tracked.

Top models

#ModelScore
1Qwen 3 VL 235B A22B (Thinking)66.8
2Qwen 3 VL 235B A22B Instruct65
3Gemini 2.5 Pro64.7
4Gemini 2.5 Flash60.5
5O359.3
6GPT-5 (Thinking)58.1
7Gemini 2.5 Flash (Thinking)52.3
8GLM-4.5V49.6
9GPT-4.1 Mini33.3
10Claude Sonnet 428.1
11GPT-4.126
12Gemini 2.0 Flash20.6
13GPT-4.1 Nano9.1
14GPT-4o4.3

Interactive version: theaggregate.ai/benchmark?slug=math-vr · How It Works · Data refreshed daily, snapshot 2026-09-05.