AIM-Bench (Affective Image Manipulation) - Ground-Truth Consistency: leaderboard
Metric: GPT-4o ground-truth consistency (0-10): the lower of a semantic-match and a visual-similarity score between the edited image and the curated target image (generated by Gemini-2.5-flash-image and chosen by three experts), over the 800 AIM-Bench affective image editing items (EmoSet source images, a target Mikels emotion with target valence-arousal-dominance coordinates and an editing instruction; 8 emotion categories, 5 editing types), each model run once with its official default settings; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 13 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen-Image-Edit-2509 | 7.09 |
| 2 | Seedream 4.0 | 7.06 |
| 3 | Qwen-Image-Edit-Plus (AIM-Bench checkpoint unspecified) | 6.41 |
| 4 | SeedEdit 3.0 | 6.03 |
| 5 | UniWorld-V2 | 6.01 |
| 6 | Flux-kontext-max | 5.98 |
| 7 | Flux-kontext-pro | 5.7 |
| 8 | Step1X-Edit (AIM-Bench checkpoint unspecified) | 5.63 |
| 9 | FLUX.1 Kontext [dev] | 5.58 |
| 10 | OmniGen2 | 4.91 |
Interactive version: theaggregate.ai/benchmark?slug=aim-bench-affective-image-manipulation-ground-truth-consistency · How It Works · Data refreshed daily, snapshot 2026-10-07.