MIOH - Existence: leaderboard

Metric: Accuracy (%; existence questions about whether objects are present across several images; multi-image questions built from COCO-ReM, PACO and Visual Genome scenes in three reasoning patterns (comprehensive, comparative, selective), averaged over the easy, hard-negative, hard-positive and eight-image conditions). Source: arxiv.org. Saturation forecast: Around December 2026. 29 models tracked.

Top models

#ModelScore
1GPT-578.4
2Gemini 2.5 Pro75.4
3MiniCPM-V-2.668.5
4Qwen 2 VL 7B67.5
5Qwen 2.5 VL 7B65.4
6Qwen 2 VL 2B47.6
7InternVL3.5-8B35.9
8Phi-4 Multimodal Instruct35.1

Interactive version: theaggregate.ai/benchmark?slug=mioh-existence · How It Works · Data refreshed daily, snapshot 2026-09-29.