UniEditBench (Image): leaderboard

Metric: Mean (1-5) of the four dimension scores (structural fidelity, text alignment, background consistency, naturalness); 633 image-editing samples over nine operations, each edit scored on a 1-5 Likert scale by the authors' Qwen3-VL-8B evaluator distilled from Qwen3-VL-235B-A22B (the second entry of each 4B/8B cell); higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 17 models tracked.

Top models

#ModelScore
1UniEditBench Hunyuan (checkpoint unspecified)4.4
2FLUX.1-Kontext-dev4.3
3Qwen-Image-Edit4.29
4Step1X-Edit4.21
5DreamOmni24.1
6UniWorld-V24.02
7OmniGen23.99
8BAGEL-7B-MoT3.96
9BAGEL-7B-MoT (Thinking)3.91
10VAREdit3.79

Interactive version: theaggregate.ai/benchmark?slug=unieditbench-image · How It Works · Data refreshed daily, snapshot 2026-10-07.