DO-Bench - Prior Robustness: leaderboard
Metric: PriorRobust (%): one minus the normalized area under the false-negative curve (present object, prior strength 0-3) and under the false-positive curve (absent object, prior strength 0-3), averaged; any constant answer scores 50, on DO-Bench, 124 scenes with ten controlled yes/no object-existence queries each (1,240 queries: a present-but-anomalous object under four prior strengths and two zoomed views, an absent-but-expected object removed by local inpainting under four prior strengths), expert-verified images, temperature 0 or the mean of three API runs; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 17 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3 Flash (Preview) | 91.45 |
| 2 | GPT-5.2 | 84.64 |
| 3 | Claude Opus 4.6 | 81.45 |
| 4 | InternVL2.5-78B | 63.78 |
| 5 | Qwen 2.5 VL 32B Instruct | 63.31 |
| 6 | Qwen 2.5 VL 72B Instruct | 61.62 |
| 7 | Qwen 2.5 VL 7B Instruct | 56.45 |
| 8 | InternVL2.5-2B | 36.7 |
Interactive version: theaggregate.ai/benchmark?slug=do-bench-prior-robustness · How It Works · Data refreshed daily, snapshot 2026-10-07.