InterveneBench - Control Variables: leaderboard

Metric: Rubric component score (0-1, times 100) for coverage of the ground-truth control variables, on InterveneBench's 74-study 2025 test set (each instance an empirical social-science study of a real policy intervention; the model proposes the causal study design from the policy context, without causal graphs); a Gemini-3-Pro grader compares the design with the expert-verified ground truth; direct inference; higher is better. Source: arxiv.org. Saturation forecast: Not forecast. 12 models tracked.

Top models

#ModelScoreOverall rank
1GPT-5.165.2#131
2Claude 3.7 Sonnet57#241
3Gemini 2.5 Pro50.8#145
4DeepSeek V3.249.2#198
5Claude Sonnet 448.8#194
6GLM-4.648#246
7GPT-OSS-120B46.6#330
8Grok 446.2#169
9GPT-4.145.8#240
10Qwen 3 235B A22B45.2#304
11Kimi K243.8#236
12Gemini 3 Flash43.4#93

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=intervenebench-control-variables · How It Works · Data refreshed daily, snapshot 2026-10-11.