MMOOC - Image-Question Mismatch (Yes/No): leaderboard

Metric: Shifted in-context score (%; mean of accuracy on the answerable Image-Question Mismatch Yes/No questions and answer rationality, the mean 0-1 judge score of the answer's correctness and evidence consistency in steps of 0.25; judge scores averaged over GPT-5.6, Claude Opus 5 and DeepSeek-V4-Pro). Source: arxiv.org. Saturation forecast: Around August 2028. 18 models tracked.

Top models

#ModelScore
1Qwen 3.5 27B79.25
2Qwen 3.5 122B A10B75.75
3Gemma 4 26B74
4Claude Opus 4.672.5
5Ministral 3 14B72.5
6GPT-4o72
7Llama 4 Maverick71.75
8O171
9Ministral 3 8B71
10Gemma 4 31B67.75
11O363.75
12InternVL3-8B62
13Gemini 3.1 Pro (Preview)59.5

Interactive version: theaggregate.ai/benchmark?slug=mmooc-image-question-mismatch-yes-no · How It Works · Data refreshed daily, snapshot 2026-09-29.