DO-Bench: leaderboard

Metric: Accuracy (%) of binary object-existence answers over all ten queries per scene (60 percent of answers are yes) on DO-Bench, 124 scenes with ten controlled yes/no object-existence queries each (1,240 queries: a present-but-anomalous object under four prior strengths and two zoomed views, an absent-but-expected object removed by local inpainting under four prior strengths), expert-verified images, temperature 0 or the mean of three API runs; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 17 models tracked.

Top models

#ModelScore
1Gemini 3 Flash (Preview)88.53
2GPT-5.285.66
3Claude Opus 4.685.22
4Qwen 2.5 VL 32B Instruct74.4
5InternVL2.5-78B72.8
6Qwen 2.5 VL 72B Instruct72.31
7Qwen 2.5 VL 7B Instruct65.81
8InternVL2.5-2B47.58

Interactive version: theaggregate.ai/benchmark?slug=do-bench · How It Works · Data refreshed daily, snapshot 2026-10-07.