UniEditBench (Image) - Naturalness: leaderboard
Metric: Naturalness score (1-5): the edited image looks realistic and artifact-free; 633 image-editing samples over nine operations, each edit scored on a 1-5 Likert scale by the authors' Qwen3-VL-8B evaluator distilled from Qwen3-VL-235B-A22B (the second entry of each 4B/8B cell); higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 17 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | FLUX.1-Kontext-dev | 4.22 |
| 2 | UniEditBench Hunyuan (checkpoint unspecified) | 4.15 |
| 3 | Qwen-Image-Edit | 4.09 |
| 4 | DreamOmni2 | 4.02 |
| 5 | Step1X-Edit | 3.94 |
| 6 | OmniGen2 | 3.79 |
| 7 | UniWorld-V2 | 3.78 |
| 8 | ICEdit | 3.58 |
| 9 | BAGEL-7B-MoT | 3.58 |
| 10 | BAGEL-7B-MoT (Thinking) | 3.52 |
Interactive version: theaggregate.ai/benchmark?slug=unieditbench-image-naturalness · How It Works · Data refreshed daily, snapshot 2026-10-07.