SDABench - Predictive: leaderboard

Metric: Accuracy (%) of open-ended answers on the held-out SDA-Synth test split (the 400 predictive analysis instances; zero-shot, JSON reasoning and answer fields, the answer scored by relative-error numeric, exact or structured matching with a GPT-4o fallback judge); higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 13 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)53.75
2GLM-542.75
3Qwen 3 235B A22B 2507 Instruct42
4Qwen 3.5 397B A17B37
5Claude Sonnet 4.636
6GPT-5.434.5
7Kimi K2.534.5
8DeepSeek V3.231
9DeepSeek R123.5
10GPT-5.4 Mini17.5
11Llama 3.3 70B Instruct2.75
12Llama 3.1 8B Instruct0.5

Interactive version: theaggregate.ai/benchmark?slug=sdabench-predictive · How It Works · Data refreshed daily, snapshot 2026-09-29.