MUN-lang (Zero-Shot): leaderboard

Metric: Win rate (%; share of MUN-lang items where a GPT-4o judge, AlpacaEval-style, prefers the model's explanation (explain how an ordinary-looking image leads to an unusual outcome) over the human-written, GPT-4o-refined reference; zero-shot). Source: arxiv.org. Saturation forecast: Estimated already saturated. 7 models tracked.

Top models

#ModelScore
1Qwen 2.5 VL 7B Instruct42.2
2Phi-4 Multimodal Instruct35.7
3Qwen 2 VL 7B Instruct34.9
4Gemma 3 4B (IT)25.7

Interactive version: theaggregate.ai/benchmark?slug=mun-lang-zero-shot · How It Works · Data refreshed daily, snapshot 2026-09-26.