MIOH: leaderboard

Metric: Accuracy (%; mean of the existence, counting, attribute and position task averages; multi-image questions built from COCO-ReM, PACO and Visual Genome scenes in three reasoning patterns (comprehensive, comparative, selective), averaged over the easy, hard-negative, hard-positive and eight-image conditions). Source: arxiv.org. Saturation forecast: Around July 2027. 29 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro64.4
2GPT-563.1
3Qwen 2 VL 7B49.1
4MiniCPM-V-2.648.5
5Qwen 2.5 VL 7B46.5
6Qwen 2 VL 2B38
7InternVL3.5-8B34.1
8Phi-4 Multimodal Instruct31

Interactive version: theaggregate.ai/benchmark?slug=mioh · How It Works · Data refreshed daily, snapshot 2026-09-29.