MMOOC - Multimodal Ambiguity (Yes/No): leaderboard

Metric: Out-of-context score (%; mean of the refusal rate, the share of the Multimodal Ambiguity Yes/No questions, which the image and context cannot answer, that the model correctly declines, and refusal rationality, the mean 0-1 judge score of its abstention reasoning in steps of 0.25; judge scores averaged over GPT-5.6, Claude Opus 5 and DeepSeek-V4-Pro). Source: arxiv.org. Saturation forecast: Around 2032. 18 models tracked.

Top models

#ModelScore
1O157.5
2Gemma 4 31B54.5
3InternVL3-8B48.25
4Ministral 3 8B46.75
5Ministral 3 14B45.25
6GPT-4o43
7Llama 4 Maverick40
8Gemma 4 26B39.5
9Qwen 3.5 27B36.25
10Claude Opus 4.634.25
11Qwen 3.5 122B A10B30.5
12O329.75
13Gemini 3.1 Pro (Preview)4.75

Interactive version: theaggregate.ai/benchmark?slug=mmooc-multimodal-ambiguity-yes-no · How It Works · Data refreshed daily, snapshot 2026-09-29.