GRADE (Image Editing) - Discipline Reasoning: leaderboard
Metric: Discipline reasoning score (0-100): weighted share of the expert-verified scoring questions the edited image satisfies, on GRADE's 520 discipline-informed image editing samples (input image, editing instruction and ground-truth image across 10 academic disciplines), a Gemini-3-Flash judge scoring reasoning against expert-verified weighted questions (written by GPT-5) and visual consistency and logical readability on 0/1/2 scales; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 20 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Nano Banana Pro | 77.5 | |
| 2 | Nano Banana 2 | 72.6 | |
| 3 | Seedream 5.0 Lite | 64.1 | |
| 4 | GPT-Image-1.5 | 54.5 | |
| 5 | FLUX.2-max | 47.8 | |
| 6 | GPT-Image-1 | 44 | |
| 7 | Nano Banana | 42.2 | |
| 8 | Seedream 4.5 | 41.3 | |
| 9 | FLUX 2 Pro | 38.9 | |
| 10 | Seedream 4.0 | 32.4 |
Interactive version: theaggregate.ai/benchmark?slug=grade-image-editing-discipline-reasoning · How It Works · Data refreshed daily, snapshot 2026-10-11.