InterveneBench - Dependent Variable: leaderboard

Metric: Rubric component score (0-1, times 100) for identification of the outcome variable, on InterveneBench's 74-study 2025 test set (each instance an empirical social-science study of a real policy intervention; the model proposes the causal study design from the policy context, without causal graphs); a Gemini-3-Pro grader compares the design with the expert-verified ground truth; direct inference; higher is better. Source: arxiv.org. Saturation forecast: Not forecast. 12 models tracked.

Top models

#ModelScoreOverall rank
1GPT-5.165.8#131
2Claude Sonnet 465.6#194
3Claude 3.7 Sonnet65.6#241
4GLM-4.664.4#246
5Qwen 3 235B A22B60.4#304
6Gemini 2.5 Pro59.4#145
7GPT-4.157.2#240
8GPT-OSS-120B57.2#330
9Grok 457.2#169
10Gemini 3 Flash55.2#93
11DeepSeek V3.253#198
12Kimi K250.8#236

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=intervenebench-dependent-variable · How It Works · Data refreshed daily, snapshot 2026-10-11.