OpenVLM SEED-Bench 2 Plus — leaderboard

OpenCompass OpenVLM evaluation of SEED-Bench-2-Plus: image and video understanding across charts, maps, webpages, navigation, OCR, spatial reasoning, and multimodal knowledge.

Metric: Accuracy (%). Source: huggingface.co. Status: saturation imminent. 211 models tracked.

Top models

#ModelScore
1GPT-4.1 (2025-04-14)73.1
2Qwen 2 VL 72B72.3
3GPT-4.1 Mini71.9
4InternVL3-38B71.8
5InternVL3-78B71.7
6Claude 3.5 Sonnet71.7
7Gemini 1.5 Pro70.8
8Ovis2-8B70.4
9InternVL3-14B70.1
10InternVL3-8B69.4
11Qwen 2 VL 7B68.6
12Gemini 1.5 Flash68.6
13Llama 3.2 90B Vision Instruct68.2
14Claude 3.7 Sonnet67.6
15Pixtral-12B67.4

Interactive version: theaggregate.ai/benchmark?slug=openvlm-seed-bench-2-plus · How the rankings work · Data refreshed daily, snapshot 2026-07-22.