MMOOC - Missing Knowledge and Background (Multiple Choice): leaderboard
Metric: Out-of-context score (%; mean of the refusal rate, the share of the Missing Knowledge and Background Multiple Choice questions, which the image and context cannot answer, that the model correctly declines, and refusal rationality, the mean 0-1 judge score of its abstention reasoning in steps of 0.25; judge scores averaged over GPT-5.6, Claude Opus 5 and DeepSeek-V4-Pro). Source: arxiv.org. Saturation forecast: Around 2033. 18 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemma 4 26B | 61.25 |
| 2 | O1 | 45.75 |
| 3 | Ministral 3 8B | 44.5 |
| 4 | Gemma 4 31B | 43.5 |
| 5 | Ministral 3 14B | 40.5 |
| 6 | Llama 4 Maverick | 36.25 |
| 7 | Qwen 3.5 27B | 35.5 |
| 8 | GPT-4o | 34.25 |
| 9 | Qwen 3.5 122B A10B | 23 |
| 10 | O3 | 19.75 |
| 11 | InternVL3-8B | 18.5 |
| 12 | Gemini 3.1 Pro (Preview) | 15 |
| 13 | Claude Opus 4.6 | 4.75 |
Interactive version: theaggregate.ai/benchmark?slug=mmooc-missing-knowledge-and-background-multiple-choice · How It Works · Data refreshed daily, snapshot 2026-09-29.