AICA-Bench: leaderboard

Metric: Overall average on a percent scale: unweighted mean of the emotion understanding average (weighted F1 of basic and CoT prompting), the emotion reasoning score and the emotion-guided content generation score; reasoning and generation are 1-5 ratings rescaled to 20-100, so the floor is 40/3, over AICA-Bench (8,086 affective images from nine public emotion datasets with GPT-4o-generated instructions; closed-source models through their APIs at standard settings, open models on A100 GPUs); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 23 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro73.49
2GPT-4o72.82
3Gemini 2.5 Flash71.14
4GPT-4o Mini70.81
5Gemini 2.0 Flash67.68
6Qwen 2.5 VL 7B Instruct65.78
7Qwen 2 VL 7B Instruct61.45
8MiniCPM-V-2.658.08

Interactive version: theaggregate.ai/benchmark?slug=aica-bench · How It Works · Data refreshed daily, snapshot 2026-10-07.