Math-VR — leaderboard

Mathematical visual reasoning benchmark for VLMs, unified models, and LLMs, reporting answer correctness and process scores on text and multimodal questions.

Metric: Overall Answer Correctness (self-reported). Source: benchmarklist.com. Status: saturation imminent. 31 models tracked.

Top models

#ModelScore
1Qwen 3 VL 235B A22B (Thinking)66.8
2Qwen 3 VL 235B A22B Instruct65
3Gemini 2.5 Pro64.7
4Gemini 2.5 Flash60.5
5O359.3
6GPT-558.1
7Claude Opus 4.154.3
8Gemini 2.5 Flash (Thinking)52.3
9GLM-4.5V49.6
10InternVL3.5-8B40.8
11GPT-4.1 Mini33.3
12Claude Sonnet 428.1
13GPT-4.126
14Gemini 2.0 Flash20.6
15GPT-4.1 Nano9.1

Interactive version: theaggregate.ai/benchmark?slug=math-vr · How the rankings work · Data refreshed daily, snapshot 2026-07-22.