InterveneBench - Model Type: leaderboard

Metric: Rubric component score (0-1, times 100) for causal inference method choice (DiD, IV, RDD, synthetic control, matching), binary per study, on InterveneBench's 74-study 2025 test set (each instance an empirical social-science study of a real policy intervention; the model proposes the causal study design from the policy context, without causal graphs); a Gemini-3-Pro grader compares the design with the expert-verified ground truth; direct inference; higher is better. Source: arxiv.org. Saturation forecast: Not forecast. 12 models tracked.

Top models

#ModelScoreOverall rank
1Claude Sonnet 454.1#194
2GLM-4.653.5#246
3Claude 3.7 Sonnet51.9#241
4GPT-5.149.3#131
5Gemini 2.5 Pro47.7#145
6Qwen 3 235B A22B45.6#304
7GPT-4.144.5#240
8Gemini 3 Flash44.5#93
9GPT-OSS-120B43.5#330
10Grok 442.4#169
11DeepSeek V3.237.3#198
12Kimi K236.1#236

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=intervenebench-model-type · How It Works · Data refreshed daily, snapshot 2026-10-11.