InsightVQA - Intent Reasoning: leaderboard

Metric: Accuracy (%) on single-choice questions asking for the most appropriate underlying response intent given the emotional state and its visual evidence, in InsightVQA-Bench, the held-out test split (a tenth of the data, about 30,000 questions) of emotion-annotated images from six public sources; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 14 models tracked.

Top models

#ModelScore
1Claude 3.7 Sonnet61.87
2Qwen 2.5 VL 72B Instruct57.24
3Gemini 2.5 Flash56.08
4GPT-4o50.34
5Qwen 2.5 VL 32B Instruct44.02
6DeepSeek V3.241.6
7InternVL3.5-8B31.18
8Qwen 2.5 VL 7B Instruct30.92

Interactive version: theaggregate.ai/benchmark?slug=insightvqa-intent-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-29.