SDABench (MCQ): leaderboard

Metric: Accuracy (%) of four-option multiple-choice answers with perturbation-based distractors on the held-out SDA-Synth test split (2,400 instances, macro-averaged over the six task types; zero-shot; 25% chance); higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 13 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)69.46
2GPT-5.466.58
3Claude Sonnet 4.666.38
4Qwen 3.5 397B A17B62.46
5DeepSeek V3.262.21
6GLM-560.79
7GPT-5.4 Mini58.46
8Kimi K2.557.04
9DeepSeek R156.38
10Qwen 3 235B A22B 2507 Instruct52.42
11Llama 3.3 70B Instruct30.46
12Llama 3.1 8B Instruct25

Interactive version: theaggregate.ai/benchmark?slug=sdabench-mcq · How It Works · Data refreshed daily, snapshot 2026-09-29.