DiCoBench - Category Commonality: leaderboard

Metric: Accuracy (%) on the category-commonality task (match objects of the same category under changed appearance) of DiCoBench (765 multi-image samples at near-2K resolution): 5-option multiple choice (four candidate image pairs or instructions plus a no-visible-commonality option, chance 20%), answer letter matched exactly, VLMEvalKit settings at temperature 0; higher is better. Source: arxiv.org. Saturation forecast: Around February 2028. 18 models tracked.

Top models

#ModelScore
1Gemini 3 Flash63.6
2Gemini 3 Pro63.6
3GPT-4.1 Mini60.9
4O4 Mini60.9
5Qwen 3.5 35B A3B59.1
6Qwen 2.5 VL 7B Instruct59.1
7GPT-553.6
8Qwen 2.5 VL 32B Instruct48.2
9Gemma 3 27B (IT)44.6
10Gemma 3 12B (IT)39.1
11GPT-4.126.4
12GPT-4o24.6

Interactive version: theaggregate.ai/benchmark?slug=dicobench-category-commonality · How It Works · Data refreshed daily, snapshot 2026-09-29.