AICA-Bench - Emotion Reasoning: leaderboard

Metric: Emotion reasoning score on a percent scale: the model explains why an image evokes its labeled emotion, scored by the AICA-Bench scoring model (Qwen2.5-VL-7B fine-tuned on human 1-5 ratings) on emotion alignment and descriptiveness and causal soundness, ratings rescaled as s/5 x 100 (20-100), over AICA-Bench (8,086 affective images from nine public emotion datasets with GPT-4o-generated instructions; closed-source models through their APIs at standard settings, open models on A100 GPUs); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 23 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro79.08
2GPT-4o77.81
3Gemini 2.5 Flash76.55
4GPT-4o Mini76.45
5Qwen 2.5 VL 7B Instruct74.5
6Gemini 2.0 Flash71.05
7MiniCPM-V-2.665.77
8Qwen 2 VL 7B Instruct65.23

Interactive version: theaggregate.ai/benchmark?slug=aica-bench-emotion-reasoning · How It Works · Data refreshed daily, snapshot 2026-10-07.