CAGE - Counterfactual (Chain-Text Probe) - Chain Score: leaderboard

Metric: Judge score (0-10) for the causal chain in arrow notation, mean of GPT-4o and Claude 3.5 Sonnet ratings against GPT-4o-generated reference answers and chains; Chain-Text Probe, level-3 counterfactual questions (what would have happened under other circumstances), where the chain needs at least two causal links; 500 held-out COCO images, three questions each; blank responses score 0. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 8 models tracked.

Top models

#ModelScore
1LLaVA-NeXT-13B7.61
2Qwen-VL-Chat4.28
3InternVL2-76B2.51
4MiniGPT-4-Vicuna-13B2.05
5LLaVA-CoT1.6
6CogVLM2-19B1.37
7LLaVA-RLHF-13B1.32
8mPLUG-Owl20.75

Interactive version: theaggregate.ai/benchmark?slug=cage-counterfactual-chain-text-probe-chain-score · How It Works · Data refreshed daily, snapshot 2026-09-29.