DLEBench - Instruction Following: leaderboard

Metric: Instruction-following score (% of the maximum) on a 4-level rubric (localization failure, wrong action, over-modification, flawless), normalized to 100, averaged over the seven instruction types with equal weight; 1,889 small-object editing samples (target 1-10% of the image area) converted from V*-Bench, MME-RealWorld and Pixel-Reasoner, seven instruction types; Oracle-guided evaluation with a Gemini-3-Pro judge on crops around human-annotated target boxes and reference edits; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 10 models tracked.

Top models

#ModelScoreOverall rank
1Nano Banana Pro (Gemini 3 Pro Image)48.97
2GPT-Image-145.32
3Step1X-Edit (DLEBench checkpoint unspecified)42.65
4UniWorld-V242.32
5UniREdit-Bagel38.03
6BAGEL-7B-MoT (Thinking)35.55
7OmniGen224.32
8Qwen-Image-Edit (DLEBench checkpoint unspecified)17.03
9MagicBrush15.34
10UniWorld-V113.84

Interactive version: theaggregate.ai/benchmark?slug=dlebench-instruction-following · How It Works · Data refreshed daily, snapshot 2026-10-11.