InsightVQA - Insight Ranking Top-1: leaderboard

Metric: Top-1 accuracy (%): share of situational-judgment items whose top-ranked candidate insight sequence matches the ground truth, in InsightVQA-Bench, the held-out test split (a tenth of the data, about 30,000 questions) of emotion-annotated images from six public sources; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 14 models tracked.

Top models

#ModelScore
1DeepSeek V3.265.77
2Qwen 2.5 VL 32B Instruct65.32
3Gemini 2.5 Flash63.51
4Claude 3.7 Sonnet63.49
5GPT-4o61.06
6Qwen 2.5 VL 72B Instruct56.83
7Qwen 2.5 VL 7B Instruct51.27
8InternVL3.5-8B43.44

Interactive version: theaggregate.ai/benchmark?slug=insightvqa-insight-ranking-top-1 · How It Works · Data refreshed daily, snapshot 2026-09-29.