MMOOC - Image-Question Mismatch (Multiple Choice): leaderboard

Metric: Shifted in-context score (%; mean of accuracy on the answerable Image-Question Mismatch Multiple Choice questions and answer rationality, the mean 0-1 judge score of the answer's correctness and evidence consistency in steps of 0.25; judge scores averaged over GPT-5.6, Claude Opus 5 and DeepSeek-V4-Pro). Source: arxiv.org. Saturation forecast: Around December 2026. 18 models tracked.

Top models

#ModelScore
1Qwen 3.5 27B94.5
2Qwen 3.5 122B A10B94.5
3Gemma 4 31B93.25
4Ministral 3 8B90
5Llama 4 Maverick85.75
6O185.25
7Gemma 4 26B83.25
8GPT-4o82
9Ministral 3 14B79.25
10O379
11Gemini 3.1 Pro (Preview)75.75
12InternVL3-8B69.25
13Claude Opus 4.640

Interactive version: theaggregate.ai/benchmark?slug=mmooc-image-question-mismatch-multiple-choice · How It Works · Data refreshed daily, snapshot 2026-09-29.