CAGE - Intervention (Chain-Text Probe) - Text Score: leaderboard

Metric: Judge score (0-10) for the textual answer, mean of GPT-4o and Claude 3.5 Sonnet ratings against GPT-4o-generated reference answers and chains; Chain-Text Probe, level-2 intervention questions (what happens if something in the image changes), where the chain needs at least one causal link; 500 held-out COCO images, three questions each; blank responses score 0. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 8 models tracked.

Top models

#ModelScore
1LLaVA-NeXT-13B7.96
2InternVL2-76B7.45
3Qwen-VL-Chat6.86
4LLaVA-CoT6.78
5mPLUG-Owl25.77
6CogVLM2-19B5.56
7MiniGPT-4-Vicuna-13B5.14
8LLaVA-RLHF-13B1.33

Interactive version: theaggregate.ai/benchmark?slug=cage-intervention-chain-text-probe-text-score · How It Works · Data refreshed daily, snapshot 2026-09-29.