AIM-Bench (Affective Image Manipulation) - VAD Distance: leaderboard
Metric: Mean Euclidean distance between the valence-arousal-dominance coordinates GPT-4o predicts for the edited image and the target coordinates (0 or more; the coordinate scale is not stated), over the 800 AIM-Bench affective image editing items (EmoSet source images, a target Mikels emotion with target valence-arousal-dominance coordinates and an editing instruction; 8 emotion categories, 5 editing types), each model run once with its official default settings; lower is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 13 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Seedream 4.0 | 2.54 |
| 2 | Qwen-Image-Edit-2509 | 2.57 |
| 3 | UniWorld-V2 | 2.62 |
| 4 | Step1X-Edit (AIM-Bench checkpoint unspecified) | 2.63 |
| 5 | SeedEdit 3.0 | 2.65 |
| 6 | Flux-kontext-max | 2.67 |
| 7 | Qwen-Image-Edit-Plus (AIM-Bench checkpoint unspecified) | 2.72 |
| 8 | Flux-kontext-pro | 2.73 |
| 9 | OmniGen2 | 2.75 |
| 10 | FLUX.1 Kontext [dev] | 2.77 |
Interactive version: theaggregate.ai/benchmark?slug=aim-bench-affective-image-manipulation-vad-distance · How It Works · Data refreshed daily, snapshot 2026-10-07.