MMVU — leaderboard

MMVU evaluates model capability on multimodal tasks from the linked upstream source with Score as the primary reported metric.

Metric: Score (self-reported). Source: benchmarklist.com. Status: saturation imminent. 36 models tracked.

Top models

#ModelScore
1O176.1
2Gemini 2.0 Flash (Thinking)69.5
3GPT-4o66.7
4Gemini 2.0 Flash66.5
5Gemini 1.5 Pro65.8
6Claude 3.5 Sonnet64.1
7Grok 263.4
8GPT-4o Mini61.5
9Gemini 1.5 Flash58.8
10Qwen 2 VL 72B53.2
11Qwen 2 VL 7B Instruct42.5
12Qwen 2 VL 2B36.5
13InternVL2-8B36.2
14Pixtral-12B32.2

Interactive version: theaggregate.ai/benchmark?slug=mmvu · How the rankings work · Data refreshed daily, snapshot 2026-07-22.