InterveneBench: leaderboard

Metric: Final score (times 100): the 45-point rubric total normalized to 0-1 (model type, core independent variable and group definition 10 points each, control and dependent variables 5 each, reasoning 2, explanation 3), on InterveneBench's 74-study 2025 test set (each instance an empirical social-science study of a real policy intervention; the model proposes the causal study design from the policy context, without causal graphs); a Gemini-3-Pro grader compares the design with the expert-verified ground truth; direct inference; higher is better. Source: arxiv.org. Saturation forecast: Around January 2028. 12 models tracked.

Top models

#ModelScoreOverall rank
1GPT-5.157.8#131
2Claude 3.7 Sonnet54.4#241
3Claude Sonnet 453.9#194
4GLM-4.653.1#246
5Gemini 2.5 Pro49.7#145
6Qwen 3 235B A22B48.2#304
7Gemini 3 Flash46.6#93
8Grok 446.4#169
9GPT-4.145.6#240
10GPT-OSS-120B44.3#330
11DeepSeek V3.242.5#198
12Kimi K241.5#236

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=intervenebench · How It Works · Data refreshed daily, snapshot 2026-10-11.