VideoMME w sub. — leaderboard

VideoMME w sub. evaluates model capability on multimodal tasks from the linked upstream source with Score as the primary reported metric.

Metric: Score (self-reported). Source: benchmarklist.com. Status: saturation imminent. 56 models tracked.

Top models

#ModelScore
1GPT-5.489.5
2Gemini 3.1 Pro (Preview)88.4
3Qwen 3.7 Plus88
4Qwen 3.6 Plus87.8
5Claude Opus 4.6 (Max)86.1
6Gemini 1.5 Pro81.3
7GPT-4o77.2
8Gemini 1.5 Flash75
9GPT-4o Mini68.9
10MiniCPM-V-2.663.7
11GPT-463.3
12Claude 3.5 Sonnet62.9

Interactive version: theaggregate.ai/benchmark?slug=videomme-w-sub · How the rankings work · Data refreshed daily, snapshot 2026-07-22.