AIM-Bench (Affective Image Manipulation): leaderboard
Metric: Overall score (%): the mean of the eight Table 1 metrics (CLIP-T, VIEScore SC and PQ, aesthetic score, GPT-4o ground-truth consistency, emotion accuracy and F1, VAD distance) after normalizing each to 0-1 with fixed empirical bounds the paper does not print, over the 800 AIM-Bench affective image editing items (EmoSet source images, a target Mikels emotion with target valence-arousal-dominance coordinates and an editing instruction; 8 emotion categories, 5 editing types), each model run once with its official default settings; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 13 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Seedream 4.0 | 48.94 |
| 2 | Qwen-Image-Edit-2509 | 47.09 |
| 3 | UniWorld-V2 | 41.57 |
| 4 | Step1X-Edit (AIM-Bench checkpoint unspecified) | 41.17 |
| 5 | Flux-kontext-max | 41.14 |
| 6 | Qwen-Image-Edit-Plus (AIM-Bench checkpoint unspecified) | 40.25 |
| 7 | SeedEdit 3.0 | 38.79 |
| 8 | Flux-kontext-pro | 38.25 |
| 9 | FLUX.1 Kontext [dev] | 33.12 |
| 10 | OmniGen2 | 32.53 |
Interactive version: theaggregate.ai/benchmark?slug=aim-bench-affective-image-manipulation · How It Works · Data refreshed daily, snapshot 2026-10-07.