AIM-Bench (Affective Image Manipulation) - Semantic Consistency: leaderboard

Metric: VIEScore semantic consistency (SC, 0-10): how well the edited image follows the editing instruction, over the 800 AIM-Bench affective image editing items (EmoSet source images, a target Mikels emotion with target valence-arousal-dominance coordinates and an editing instruction; 8 emotion categories, 5 editing types), each model run once with its official default settings; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 13 models tracked.

Top models

#ModelScore
1Seedream 4.08.28
2Qwen-Image-Edit-25098.13
3SeedEdit 3.07.85
4Qwen-Image-Edit-Plus (AIM-Bench checkpoint unspecified)7.74
5Flux-kontext-max7.72
6UniWorld-V27.66
7Flux-kontext-pro7.52
8Step1X-Edit (AIM-Bench checkpoint unspecified)7.52
9FLUX.1 Kontext [dev]7.14
10OmniGen26.74

Interactive version: theaggregate.ai/benchmark?slug=aim-bench-affective-image-manipulation-semantic-consistency · How It Works · Data refreshed daily, snapshot 2026-10-07.