CAGE - Counterfactual (Text-Only Probe): leaderboard

Metric: Judge score (0-10), mean of GPT-4o and Claude 3.5 Sonnet ratings of accuracy, relevance and logical consistency against GPT-4o-generated reference answers; level-3 counterfactual (what would have happened otherwise) questions, 3 per image on 500 COCO test images; free-text answer without a causal chain (Text-Only Probe); blank responses score 0. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 8 models tracked.

Top models

#ModelScore
1CogVLM2-19B8.11
2Qwen-VL-Chat7.99
3LLaVA-NeXT-13B7.94
4InternVL2-76B7.12
5mPLUG-Owl27.06
6LLaVA-RLHF-13B6.92
7MiniGPT-4-Vicuna-13B6.68
8LLaVA-CoT4.68

Interactive version: theaggregate.ai/benchmark?slug=cage-counterfactual-text-only-probe · How It Works · Data refreshed daily, snapshot 2026-09-26.