DMC-CF - Double Intervention (L2): leaderboard
Metric: Accuracy (%) on the L2 set: questions adding a second intervention generated from an extracted causal graph that either breaks the outcome (answer no) or leaves it intact (answer yes), yes/no counterfactual questions about 1,614 real-world causal event videos (videos given natively to Gemini and Qwen3-VL, as at most 16 frames to GPT and Claude); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 5 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.2 | 55.2 |
| 2 | Qwen 3 VL 235B A22B (Thinking) | 54.93 |
| 3 | Qwen 3 VL 235B A22B Instruct | 48.44 |
Interactive version: theaggregate.ai/benchmark?slug=dmc-cf-double-intervention-l2 · How It Works · Data refreshed daily, snapshot 2026-10-07.