AIM-Bench (Affective Image Manipulation) - Semantic Consistency: leaderboard
Metric: VIEScore semantic consistency (SC, 0-10): how well the edited image follows the editing instruction, over the 800 AIM-Bench affective image editing items (EmoSet source images, a target Mikels emotion with target valence-arousal-dominance coordinates and an editing instruction; 8 emotion categories, 5 editing types), each model run once with its official default settings; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 13 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Seedream 4.0 | 8.28 |
| 2 | Qwen-Image-Edit-2509 | 8.13 |
| 3 | SeedEdit 3.0 | 7.85 |
| 4 | Qwen-Image-Edit-Plus (AIM-Bench checkpoint unspecified) | 7.74 |
| 5 | Flux-kontext-max | 7.72 |
| 6 | UniWorld-V2 | 7.66 |
| 7 | Flux-kontext-pro | 7.52 |
| 8 | Step1X-Edit (AIM-Bench checkpoint unspecified) | 7.52 |
| 9 | FLUX.1 Kontext [dev] | 7.14 |
| 10 | OmniGen2 | 6.74 |
Interactive version: theaggregate.ai/benchmark?slug=aim-bench-affective-image-manipulation-semantic-consistency · How It Works · Data refreshed daily, snapshot 2026-10-07.