InterveneBench - Reasoning: leaderboard

Metric: Rubric component score (0-1, times 100) for logical self-consistency of the methodological choice, on InterveneBench's 74-study 2025 test set (each instance an empirical social-science study of a real policy intervention; the model proposes the causal study design from the policy context, without causal graphs); a Gemini-3-Pro grader compares the design with the expert-verified ground truth; direct inference; higher is better. Source: arxiv.org. Saturation forecast: Around February 2028. 12 models tracked.

Top models

#ModelScoreOverall rank
1Claude Sonnet 453#194
2GPT-5.153#131
3GLM-4.652#246
4Claude 3.7 Sonnet51#241
5Gemini 2.5 Pro47.5#145
6GPT-4.144.5#240
7Gemini 3 Flash44.5#93
8Qwen 3 235B A22B44.5#304
9Grok 442#169
10GPT-OSS-120B41.5#330
11DeepSeek V3.237.5#198
12Kimi K236#236

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=intervenebench-reasoning · How It Works · Data refreshed daily, snapshot 2026-10-11.