DiCoBench - Reasoning Difference: leaderboard

Metric: Accuracy (%) on the reasoning-difference task (traces that imply a latent physical event) of DiCoBench (765 multi-image samples at near-2K resolution): 5-option multiple choice (four candidate image pairs or instructions plus a no-visible-difference option, chance 20%), answer letter matched exactly, VLMEvalKit settings at temperature 0; higher is better. Source: arxiv.org. Saturation forecast: Around 2037. 18 models tracked.

Top models

#ModelScore
1GPT-4.1 Mini28
2Gemini 3 Pro28
3Qwen 3.5 35B A3B26.7
4GPT-4o25.33
5Qwen 2.5 VL 7B Instruct25.3
6GPT-522.7
7O4 Mini22.7
8Gemma 3 27B (IT)21.3
9GPT-4.121.3
10Gemma 3 12B (IT)20
11Qwen 2.5 VL 32B Instruct20
12Gemini 3 Flash14.3

Interactive version: theaggregate.ai/benchmark?slug=dicobench-reasoning-difference · How It Works · Data refreshed daily, snapshot 2026-09-29.