AIM-Bench (Affective Image Manipulation) - Ground-Truth Consistency: leaderboard

Metric: GPT-4o ground-truth consistency (0-10): the lower of a semantic-match and a visual-similarity score between the edited image and the curated target image (generated by Gemini-2.5-flash-image and chosen by three experts), over the 800 AIM-Bench affective image editing items (EmoSet source images, a target Mikels emotion with target valence-arousal-dominance coordinates and an editing instruction; 8 emotion categories, 5 editing types), each model run once with its official default settings; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 13 models tracked.

Top models

#ModelScore
1Qwen-Image-Edit-25097.09
2Seedream 4.07.06
3Qwen-Image-Edit-Plus (AIM-Bench checkpoint unspecified)6.41
4SeedEdit 3.06.03
5UniWorld-V26.01
6Flux-kontext-max5.98
7Flux-kontext-pro5.7
8Step1X-Edit (AIM-Bench checkpoint unspecified)5.63
9FLUX.1 Kontext [dev]5.58
10OmniGen24.91

Interactive version: theaggregate.ai/benchmark?slug=aim-bench-affective-image-manipulation-ground-truth-consistency · How It Works · Data refreshed daily, snapshot 2026-10-07.