DLEBench: leaderboard
Metric: Average score (% of the maximum): mean of instruction following and visual consistency, each a 4-level failure-mode rubric normalized to 100, averaged over the seven instruction types with equal weight; 1,889 small-object editing samples (target 1-10% of the image area) converted from V*-Bench, MME-RealWorld and Pixel-Reasoner, seven instruction types; Oracle-guided evaluation with a Gemini-3-Pro judge on crops around human-annotated target boxes and reference edits; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 10 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Nano Banana Pro (Gemini 3 Pro Image) | 65.55 | |
| 2 | BAGEL-7B-MoT (Thinking) | 61 | |
| 3 | Step1X-Edit (DLEBench checkpoint unspecified) | 55.39 | |
| 4 | UniWorld-V2 | 49.8 | |
| 5 | OmniGen2 | 42.06 | |
| 6 | Qwen-Image-Edit (DLEBench checkpoint unspecified) | 41.75 | |
| 7 | GPT-Image-1 | 40.25 | |
| 8 | UniREdit-Bagel | 37.68 | |
| 9 | UniWorld-V1 | 32.42 | |
| 10 | MagicBrush | 23.16 |
Interactive version: theaggregate.ai/benchmark?slug=dlebench · How It Works · Data refreshed daily, snapshot 2026-10-11.