DMC-CF - Double Intervention (L2): leaderboard

Metric: Accuracy (%) on the L2 set: questions adding a second intervention generated from an extracted causal graph that either breaks the outcome (answer no) or leaves it intact (answer yes), yes/no counterfactual questions about 1,614 real-world causal event videos (videos given natively to Gemini and Qwen3-VL, as at most 16 frames to GPT and Claude); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 5 models tracked.

Top models

#ModelScore
1GPT-5.255.2
2Qwen 3 VL 235B A22B (Thinking)54.93
3Qwen 3 VL 235B A22B Instruct48.44

Interactive version: theaggregate.ai/benchmark?slug=dmc-cf-double-intervention-l2 · How It Works · Data refreshed daily, snapshot 2026-10-07.