AIM-Bench (Affective Image Manipulation): leaderboard

Metric: Overall score (%): the mean of the eight Table 1 metrics (CLIP-T, VIEScore SC and PQ, aesthetic score, GPT-4o ground-truth consistency, emotion accuracy and F1, VAD distance) after normalizing each to 0-1 with fixed empirical bounds the paper does not print, over the 800 AIM-Bench affective image editing items (EmoSet source images, a target Mikels emotion with target valence-arousal-dominance coordinates and an editing instruction; 8 emotion categories, 5 editing types), each model run once with its official default settings; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 13 models tracked.

Top models

#ModelScore
1Seedream 4.048.94
2Qwen-Image-Edit-250947.09
3UniWorld-V241.57
4Step1X-Edit (AIM-Bench checkpoint unspecified)41.17
5Flux-kontext-max41.14
6Qwen-Image-Edit-Plus (AIM-Bench checkpoint unspecified)40.25
7SeedEdit 3.038.79
8Flux-kontext-pro38.25
9FLUX.1 Kontext [dev]33.12
10OmniGen232.53

Interactive version: theaggregate.ai/benchmark?slug=aim-bench-affective-image-manipulation · How It Works · Data refreshed daily, snapshot 2026-10-07.