GenScale - Human-Product Scale: leaderboard
Metric: Plausible pairs (%; share of human-product scenes whose product is judged plausible in size against its stated metric dimensions, score 3 of 5, image-conditioned generation, Gemini 3.1 Pro Preview judge calibrated against nine human raters, five samples per pair, majority vote; 291 to 297 scored scenes per model). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 5 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT Image 2 | 78.9 |
| 2 | Nano Banana 2 | 75.6 |
| 3 | Seedream 4.5 | 74.9 |
| 4 | Qwen-Image-Edit-2511 | 71 |
| 5 | FLUX.1 Kontext-dev | 63.2 |
Interactive version: theaggregate.ai/benchmark?slug=genscale-human-product-scale · How It Works · Data refreshed daily, snapshot 2026-09-26.