MIOH - Attribute: leaderboard

Metric: Accuracy (%; attribute questions about object attributes across several images; multi-image questions built from COCO-ReM, PACO and Visual Genome scenes in three reasoning patterns (comprehensive, comparative, selective), averaged over the easy, hard-negative, hard-positive and eight-image conditions). Source: arxiv.org. Saturation forecast: Around May 2028. 29 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro57.9
2GPT-557.8
3Qwen 2 VL 7B41.9
4Qwen 2.5 VL 7B39.9
5MiniCPM-V-2.638.2
6Qwen 2 VL 2B37.9
7InternVL3.5-8B32
8Phi-4 Multimodal Instruct29.3

Interactive version: theaggregate.ai/benchmark?slug=mioh-attribute · How It Works · Data refreshed daily, snapshot 2026-09-29.