InterveneBench - Core Independent Variable: leaderboard

Metric: Rubric component score (0-1, times 100) for operationalization of the policy intervention as the core independent variable (full, half or no credit), on InterveneBench's 74-study 2025 test set (each instance an empirical social-science study of a real policy intervention; the model proposes the causal study design from the policy context, without causal graphs); a Gemini-3-Pro grader compares the design with the expert-verified ground truth; direct inference; higher is better. Source: arxiv.org. Saturation forecast: Around September 2027. 12 models tracked.

Top models

#ModelScoreOverall rank
1GPT-5.165.6#131
2Claude 3.7 Sonnet57.8#241
3Claude Sonnet 457.3#194
4GLM-4.656.6#246
5Qwen 3 235B A22B54.1#304
6Grok 451.4#169
7Gemini 2.5 Pro50.9#145
8GPT-4.149.8#240
9Gemini 3 Flash47.7#93
10DeepSeek V3.245.8#198
11Kimi K244.5#236
12GPT-OSS-120B44#330

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=intervenebench-core-independent-variable · How It Works · Data refreshed daily, snapshot 2026-10-11.