MMOOC - Partial Answerability (Yes/No): leaderboard

Metric: Shifted in-context score (%; mean of accuracy on the answerable Partial Answerability Yes/No questions and answer rationality, the mean 0-1 judge score of the answer's correctness and evidence consistency in steps of 0.25; judge scores averaged over GPT-5.6, Claude Opus 5 and DeepSeek-V4-Pro). Source: arxiv.org. Saturation forecast: Around June 2027. 18 models tracked.

Top models

#ModelScore
1Claude Opus 4.682.75
2O181.25
3Ministral 3 14B81
4Gemma 4 26B79
5GPT-4o78.25
6Qwen 3.5 122B A10B78.25
7Gemma 4 31B77.25
8O377.25
9Qwen 3.5 27B75.75
10Llama 4 Maverick74.25
11Ministral 3 8B72.75
12InternVL3-8B66.75
13Gemini 3.1 Pro (Preview)62.75

Interactive version: theaggregate.ai/benchmark?slug=mmooc-partial-answerability-yes-no · How It Works · Data refreshed daily, snapshot 2026-09-29.