DO-Bench: leaderboard
Metric: Accuracy (%) of binary object-existence answers over all ten queries per scene (60 percent of answers are yes) on DO-Bench, 124 scenes with ten controlled yes/no object-existence queries each (1,240 queries: a present-but-anomalous object under four prior strengths and two zoomed views, an absent-but-expected object removed by local inpainting under four prior strengths), expert-verified images, temperature 0 or the mean of three API runs; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 17 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3 Flash (Preview) | 88.53 |
| 2 | GPT-5.2 | 85.66 |
| 3 | Claude Opus 4.6 | 85.22 |
| 4 | Qwen 2.5 VL 32B Instruct | 74.4 |
| 5 | InternVL2.5-78B | 72.8 |
| 6 | Qwen 2.5 VL 72B Instruct | 72.31 |
| 7 | Qwen 2.5 VL 7B Instruct | 65.81 |
| 8 | InternVL2.5-2B | 47.58 |
Interactive version: theaggregate.ai/benchmark?slug=do-bench · How It Works · Data refreshed daily, snapshot 2026-10-07.