MMSI-Bench — leaderboard

Multi-image spatial intelligence benchmark testing VLMs on positional relationships, attribute reasoning, and motion reasoning across multiple images.

Metric: Accuracy (%). Source: github.com. Status: saturation imminent. 38 models tracked.

Top models

#ModelScore
1Human Expert97.2
2Median Human97.2
3Gemini 3 Pro49.2
4GPT-541.9
5O341
6GPT-4.540.3
7Gemini 2.5 Pro36.9
8Doubao-1.5-Pro33
9GPT-4.130.9
10GPT-4o30.3
11Claude 3.7 Sonnet (Thinking)30.2
12InternVL2.5-2B29
13InternVL3-78B28.5
14InternVL2.5-78B28.5
15InternVL3-8B25.7

Interactive version: theaggregate.ai/benchmark?slug=mmsi-bench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.