AICA-Bench - Emotion-Guided Generation: leaderboard

Metric: Emotion-guided content generation score on a percent scale: the model writes a scene description congruent with a target emotion, scored by the AICA-Bench scoring model (Qwen2.5-VL-7B fine-tuned on human 1-5 ratings) on emotion alignment and descriptiveness, ratings rescaled as s/5 x 100 (20-100), over AICA-Bench (8,086 affective images from nine public emotion datasets with GPT-4o-generated instructions; closed-source models through their APIs at standard settings, open models on A100 GPUs); higher is better. Source: arxiv.org. Saturation forecast: Around December 2027. 23 models tracked.

Top models

#ModelScore
1GPT-4o75.73
2Gemini 2.5 Pro74.13
3GPT-4o Mini74.09
4Gemini 2.5 Flash68.19
5Qwen 2.5 VL 7B Instruct66
6Qwen 2 VL 7B Instruct64.76
7Gemini 2.0 Flash63.93
8MiniCPM-V-2.663

Interactive version: theaggregate.ai/benchmark?slug=aica-bench-emotion-guided-generation · How It Works · Data refreshed daily, snapshot 2026-10-07.