DLEBench: leaderboard

Metric: Average score (% of the maximum): mean of instruction following and visual consistency, each a 4-level failure-mode rubric normalized to 100, averaged over the seven instruction types with equal weight; 1,889 small-object editing samples (target 1-10% of the image area) converted from V*-Bench, MME-RealWorld and Pixel-Reasoner, seven instruction types; Oracle-guided evaluation with a Gemini-3-Pro judge on crops around human-annotated target boxes and reference edits; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 10 models tracked.

Top models

#ModelScoreOverall rank
1Nano Banana Pro (Gemini 3 Pro Image)65.55
2BAGEL-7B-MoT (Thinking)61
3Step1X-Edit (DLEBench checkpoint unspecified)55.39
4UniWorld-V249.8
5OmniGen242.06
6Qwen-Image-Edit (DLEBench checkpoint unspecified)41.75
7GPT-Image-140.25
8UniREdit-Bagel37.68
9UniWorld-V132.42
10MagicBrush23.16

Interactive version: theaggregate.ai/benchmark?slug=dlebench · How It Works · Data refreshed daily, snapshot 2026-10-11.