MMOOC - Unclear Logical and Symbolic (Multiple Choice): leaderboard

Metric: Out-of-context score (%; mean of the refusal rate, the share of the Unclear Logical and Symbolic Multiple Choice questions, which the image and context cannot answer, that the model correctly declines, and refusal rationality, the mean 0-1 judge score of its abstention reasoning in steps of 0.25; judge scores averaged over GPT-5.6, Claude Opus 5 and DeepSeek-V4-Pro). Source: arxiv.org. Saturation forecast: Around 2034. 18 models tracked.

Top models

#ModelScore
1Gemma 4 26B57.25
2Gemma 4 31B47.5
3Ministral 3 14B42
4Qwen 3.5 122B A10B35.25
5GPT-4o34
6O130
7Ministral 3 8B27.75
8Llama 4 Maverick27
9InternVL3-8B21.25
10Qwen 3.5 27B16.75
11O316
12Gemini 3.1 Pro (Preview)12.5
13Claude Opus 4.66.75

Interactive version: theaggregate.ai/benchmark?slug=mmooc-unclear-logical-and-symbolic-multiple-choice · How It Works · Data refreshed daily, snapshot 2026-09-29.