CAGE - Intervention (Text-Only Probe): leaderboard
Metric: Judge score (0-10), mean of GPT-4o and Claude 3.5 Sonnet ratings of accuracy, relevance and logical consistency against GPT-4o-generated reference answers; level-2 intervention (what happens if something is changed) questions, 3 per image on 500 COCO test images; free-text answer without a causal chain (Text-Only Probe); blank responses score 0. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 8 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | CogVLM2-19B | 8.26 |
| 2 | Qwen-VL-Chat | 8.19 |
| 3 | LLaVA-NeXT-13B | 8.08 |
| 4 | InternVL2-76B | 7.57 |
| 5 | MiniGPT-4-Vicuna-13B | 7.14 |
| 6 | mPLUG-Owl2 | 7.09 |
| 7 | LLaVA-RLHF-13B | 6.76 |
| 8 | LLaVA-CoT | 3.07 |
Interactive version: theaggregate.ai/benchmark?slug=cage-intervention-text-only-probe · How It Works · Data refreshed daily, snapshot 2026-09-26.