AMVICC - MMVP Individual Accuracy - Positional and Relational Context: leaderboard

Metric: Individual accuracy (%; the 32 MMVP questions on 16 image pairs in the Positional and Relational Context category, free-form answers graded by GPT-4 against the authors' rubric). Source: arxiv.org. Saturation forecast: Around January 2027. 11 models tracked.

Top models

#ModelScore
1Llama 3.2 90B Vision Instruct90.63
2Claude Opus 4.181.25
3Gemini 2.5 Pro78.13
4GPT-4o78.13
5Claude Sonnet 475
6Llama 4 Scout75
7Llama 4 Maverick71.88
8Qwen 2.5 VL 72B Instruct71.88
9pixtral-large-241171.88
10Gemma 3 27B68.75
11Grok 450

Interactive version: theaggregate.ai/benchmark?slug=amvicc-mmvp-individual-accuracy-positional-and-relational-context · How It Works · Data refreshed daily, snapshot 2026-09-29.