InsightVQA - Insight Ranking: leaderboard

Metric: Spearman rank correlation (x100, -100 to 100) between the model ranking of candidate insight sequences and the ground-truth ranking in the situational-judgment test of InsightVQA-Bench, the held-out test split (a tenth of the data, about 30,000 questions) of emotion-annotated images from six public sources; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 11 models tracked.

Top models

#ModelScore
1Claude 3.7 Sonnet81.77
2Gemini 2.5 Flash80.56
3DeepSeek V3.280.24
4GPT-4o77.51
5Qwen 2.5 VL 72B Instruct76.3
6Qwen 2.5 VL 32B Instruct69.24
7Qwen 2.5 VL 7B Instruct34.95
8InternVL3.5-8B31.81

Interactive version: theaggregate.ai/benchmark?slug=insightvqa-insight-ranking · How It Works · Data refreshed daily, snapshot 2026-09-29.