CAGE - Counterfactual (Chain-Text Probe) - Text Score: leaderboard
Metric: Judge score (0-10) for the textual answer, mean of GPT-4o and Claude 3.5 Sonnet ratings against GPT-4o-generated reference answers and chains; Chain-Text Probe, level-3 counterfactual questions (what would have happened under other circumstances), where the chain needs at least two causal links; 500 held-out COCO images, three questions each; blank responses score 0. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 8 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | LLaVA-NeXT-13B | 7.64 |
| 2 | InternVL2-76B | 7.2 |
| 3 | LLaVA-CoT | 6.81 |
| 4 | Qwen-VL-Chat | 6.76 |
| 5 | mPLUG-Owl2 | 5.84 |
| 6 | CogVLM2-19B | 5.3 |
| 7 | MiniGPT-4-Vicuna-13B | 4.91 |
| 8 | LLaVA-RLHF-13B | 1.89 |
Interactive version: theaggregate.ai/benchmark?slug=cage-counterfactual-chain-text-probe-text-score · How It Works · Data refreshed daily, snapshot 2026-09-29.