OpenVLM MME: leaderboard

OpenCompass OpenVLM evaluation of MME: perception and cognition tests for multimodal models covering OCR, object existence, position, color, celebrity, artwork, commonsense, numerical calculation, and text translation.

Metric: Overall Score. Source: huggingface.co. Status: saturated. 235 models tracked.

Top models

#ModelScore
1InternVL3-78B2538.6
2Qwen 2 VL 72B2504.2
3InternVL3-38B2500.7
4InternVL2.5-78B2467.6
5InternVL3-8B2422
6Gemini 2.0 Flash2386.1
7GPT-4.1 (2025-04-14)2337.6
8Qwen 2 VL 7B2276.3
9InternVL2-8B2215.1
10Grok 2 (1212)2154.7
11Gemini 1.0 Pro2148.9
12Gemma 3 27B2148.5
13GPT-4.1 Mini2110.9
14Gemini 1.5 Pro2110.6
15Claude 3.5 Sonnet (20241022)2110

Interactive version: theaggregate.ai/benchmark?slug=openvlm-mme · How It Works · Data refreshed daily, snapshot 2026-09-05.