GRADE (Image Editing): leaderboard
Metric: Accuracy (%): share of samples that get the maximum score on discipline reasoning, visual consistency and logical readability at once, on GRADE's 520 discipline-informed image editing samples (input image, editing instruction and ground-truth image across 10 academic disciplines), a Gemini-3-Flash judge scoring reasoning against expert-verified weighted questions (written by GPT-5) and visual consistency and logical readability on 0/1/2 scales; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 20 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Nano Banana Pro | 46.2 | |
| 2 | Nano Banana 2 | 39.6 | |
| 3 | Seedream 5.0 Lite | 24.7 | |
| 4 | GPT-Image-1.5 | 16 | |
| 5 | FLUX.2-max | 11.9 | |
| 6 | Nano Banana | 9 | |
| 7 | Seedream 4.5 | 6.9 | |
| 8 | GPT-Image-1 | 6 | |
| 9 | FLUX 2 Pro | 4.4 | |
| 10 | Seedream 4.0 | 3.1 |
Interactive version: theaggregate.ai/benchmark?slug=grade-image-editing · How It Works · Data refreshed daily, snapshot 2026-10-11.