DiCoBench - Spatial Commonality: leaderboard

Metric: Accuracy (%) on the spatial-commonality task (locate the shared element on a 6 x 6 grid) of DiCoBench (765 multi-image samples at near-2K resolution): 5-option multiple choice (four candidate image pairs or instructions plus a no-visible-commonality option, chance 20%), answer letter matched exactly, VLMEvalKit settings at temperature 0; higher is better. Source: arxiv.org. Saturation forecast: Around August 2027. 18 models tracked.

Top models

#ModelScore
1Gemini 3 Pro58.4
2Qwen 3.5 35B A3B49.6
3O4 Mini48
4Gemini 3 Flash41.7
5Gemma 3 27B (IT)36.8
6GPT-536
7Gemma 3 12B (IT)29.6
8GPT-4.127.2
9GPT-4.1 Mini25.6
10Qwen 2.5 VL 7B Instruct23.2
11Qwen 2.5 VL 32B Instruct21.6

Interactive version: theaggregate.ai/benchmark?slug=dicobench-spatial-commonality · How It Works · Data refreshed daily, snapshot 2026-09-29.