OpenVLM A-Bench: leaderboard

OpenCompass OpenVLM evaluation of A-Bench: attribute and compositional image understanding with questions about object attributes, spatial arrangement, occlusion, orientation, size, and relations.

Metric: Accuracy (%). Source: huggingface.co. Status: saturated. 160 models tracked.

Top models

#ModelScore
1GPT-4.1 (2025-04-14)79.6
2Qwen 2 VL 72B79.4
3InternVL3-78B78.9
4GPT-4.1 Mini77.4
5InternVL3-38B77.1
6Qwen 2 VL 7B76.4
7InternVL3-8B76.3
8Gemma 3 27B76.3
9Grok 2 (1212)76.3
10Aquila-VL-2B75.4
11Gemini 1.5 Pro75.4
12Gemini 1.5 Flash73.6
13Pixtral-12B73.1
14InternVL2-8B72.9
15GPT-4.1 Nano72.9

Interactive version: theaggregate.ai/benchmark?slug=openvlm-a-bench · How It Works · Data refreshed daily, snapshot 2026-09-05.