SDABench - Causal (MCQ): leaderboard

Metric: Accuracy (%) of four-option multiple-choice answers with perturbation-based distractors on the held-out SDA-Synth test split (the 400 causal analysis instances; zero-shot; 25% chance); higher is better. Source: arxiv.org. Saturation forecast: Around June 2028. 13 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)59.75
2Qwen 3.5 397B A17B52.75
3GPT-5.4 Mini50.25
4Claude Sonnet 4.650
5DeepSeek V3.245.75
6GLM-542
7GPT-5.440.75
8DeepSeek R140.5
9Qwen 3 235B A22B 2507 Instruct36.5
10Kimi K2.536.25
11Llama 3.1 8B Instruct30.5
12Llama 3.3 70B Instruct24

Interactive version: theaggregate.ai/benchmark?slug=sdabench-causal-mcq · How It Works · Data refreshed daily, snapshot 2026-09-29.