SDABench - Causal: leaderboard

Metric: Accuracy (%) of open-ended answers on the held-out SDA-Synth test split (the 400 causal analysis instances; zero-shot, JSON reasoning and answer fields, the answer scored by relative-error numeric, exact or structured matching with a GPT-4o fallback judge); higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 13 models tracked.

Top models

#ModelScore
1GPT-5.452.75
2GLM-548
3Claude Sonnet 4.646
4DeepSeek V3.246
5Gemini 3.1 Pro (Preview)43.5
6DeepSeek R140
7Kimi K2.539.75
8Qwen 3 235B A22B 2507 Instruct38.75
9Qwen 3.5 397B A17B37
10GPT-5.4 Mini23
11Llama 3.1 8B Instruct7.75
12Llama 3.3 70B Instruct4.5

Interactive version: theaggregate.ai/benchmark?slug=sdabench-causal · How It Works · Data refreshed daily, snapshot 2026-09-29.