MMOOC - Misleading Premise (Open-Ended VQA): leaderboard

Metric: Shifted in-context score (%; mean of accuracy on the answerable Misleading Premise Open-Ended VQA questions and answer rationality, the mean 0-1 judge score of the answer's correctness and evidence consistency in steps of 0.25; judge scores averaged over GPT-5.6, Claude Opus 5 and DeepSeek-V4-Pro). Source: arxiv.org. Saturation forecast: Around December 2026. 18 models tracked.

Top models

#ModelScore
1Qwen 3.5 27B90.5
2Qwen 3.5 122B A10B90.25
3Llama 4 Maverick77.5
4GPT-4o76.75
5Ministral 3 8B76.75
6Gemma 4 31B75.75
7O174.75
8Gemma 4 26B74.25
9O371
10Ministral 3 14B65.25
11InternVL3-8B62.5
12Gemini 3.1 Pro (Preview)61.5
13Claude Opus 4.621.75

Interactive version: theaggregate.ai/benchmark?slug=mmooc-misleading-premise-open-ended-vqa · How It Works · Data refreshed daily, snapshot 2026-09-29.