AIM-Bench (Affective Image Manipulation) - Perceptual Quality: leaderboard
Metric: VIEScore PQ (0-10): naturalness of the edited image, structural fidelity and absence of distortions, over the 800 AIM-Bench affective image editing items (EmoSet source images, a target Mikels emotion with target valence-arousal-dominance coordinates and an editing instruction; 8 emotion categories, 5 editing types), each model run once with its official default settings; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 13 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Seedream 4.0 | 6.97 |
| 2 | Flux-kontext-max | 6.77 |
| 3 | Flux-kontext-pro | 6.77 |
| 4 | Qwen-Image-Edit-2509 | 6.4 |
| 5 | UniWorld-V2 | 6.22 |
| 6 | Qwen-Image-Edit-Plus (AIM-Bench checkpoint unspecified) | 6.11 |
| 7 | DreamOmni2 | 6.11 |
| 8 | SeedEdit 3.0 | 6.02 |
| 9 | OmniGen2 | 5.97 |
| 10 | Step1X-Edit (AIM-Bench checkpoint unspecified) | 5.89 |
Interactive version: theaggregate.ai/benchmark?slug=aim-bench-affective-image-manipulation-perceptual-quality · How It Works · Data refreshed daily, snapshot 2026-10-07.