GenScale - Scale Correction: leaderboard

Metric: Plausible pairs after editing (%; share of implausible pairs from failed generations judged plausible after a general-purpose editor corrects the image, score 3 of 5, autonomous-discovery and explicit-instruction settings pooled, Gemini 3.1 Pro Preview judge calibrated against nine human raters, five samples per pair, majority vote; 20.4% before editing, 190 to 197 scored pairs per editor). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 5 models tracked.

Top models

#ModelScore
1GPT Image 253.3
2Nano Banana 232.5
3Seedream 4.531.6
4FLUX.1 Kontext-dev28.5
5Qwen-Image-Edit-251128.4

Interactive version: theaggregate.ai/benchmark?slug=genscale-scale-correction · How It Works · Data refreshed daily, snapshot 2026-09-26.