GRADE (Image Editing) - History: leaderboard
Metric: Accuracy (%): share of samples that get the maximum score on discipline reasoning, visual consistency and logical readability at once, on the 27 history samples on GRADE's 520 discipline-informed image editing samples (input image, editing instruction and ground-truth image across 10 academic disciplines), a Gemini-3-Flash judge scoring reasoning against expert-verified weighted questions (written by GPT-5) and visual consistency and logical readability on 0/1/2 scales; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 20 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Nano Banana Pro | 29.6 | |
| 2 | Nano Banana 2 | 22.2 | |
| 3 | Step1X-Edit (Thinking + Reflection) | 7.1 | |
| 4 | Seedream 5.0 Lite | 3.7 | |
| 5 | FLUX.2-max | 3.7 | |
| 6 | Nano Banana | 3.7 | |
| 7 | Step1X-Edit (Thinking) | 3.7 | |
| 8 | Step1X-Edit | 3.7 | |
| 9 | GPT-Image-1.5 | 0 | |
| 10 | Seedream 4.5 | 0 |
Interactive version: theaggregate.ai/benchmark?slug=grade-image-editing-history · How It Works · Data refreshed daily, snapshot 2026-10-11.