MMSI-Bench: leaderboard

Multi-image spatial intelligence benchmark testing VLMs on positional relationships, attribute reasoning, and motion reasoning across multiple images.

Metric: Accuracy (%). Source: github.com. Status: saturation imminent. 38 models tracked.

Top models

#ModelScore
1Median Expert97.2
2Gemini 3 Pro49.2
3GPT-541.9
4O341
5GPT-4.540.3
6Gemini 2.5 Pro (Thinking)37
7Gemini 2.5 Pro36.9
8Doubao-1.5-Pro33
9GPT-4.130.9
10GPT-4o30.3
11Claude 3.7 Sonnet (Thinking)30.2
12InternVL2.5-2B29
13InternVL3-78B28.5
14InternVL2.5-78B28.5
15InternVL3-8B25.7

Interactive version: theaggregate.ai/benchmark?slug=mmsi-bench · How It Works · Data refreshed daily, snapshot 2026-09-05.