MUN-vis (5-Shot): leaderboard
Metric: Win rate (%; share of MUN-vis items where a GPT-4o judge, AlpacaEval-style, prefers the model's explanation (explain why a visually odd image leads to an ordinary outcome) over the human-written, GPT-4o-refined reference; five randomly chosen in-context examples). Source: arxiv.org. Saturation forecast: Estimated already saturated. 7 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Phi-4 Multimodal Instruct | 63 |
| 2 | Qwen 2.5 VL 7B Instruct | 43.9 |
| 3 | Gemma 3 4B (IT) | 39.3 |
| 4 | Qwen 2 VL 7B Instruct | 27.2 |
Interactive version: theaggregate.ai/benchmark?slug=mun-vis-5-shot · How It Works · Data refreshed daily, snapshot 2026-09-26.