GenScale - Human-Product Scale: leaderboard

Metric: Plausible pairs (%; share of human-product scenes whose product is judged plausible in size against its stated metric dimensions, score 3 of 5, image-conditioned generation, Gemini 3.1 Pro Preview judge calibrated against nine human raters, five samples per pair, majority vote; 291 to 297 scored scenes per model). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 5 models tracked.

Top models

#ModelScore
1GPT Image 278.9
2Nano Banana 275.6
3Seedream 4.574.9
4Qwen-Image-Edit-251171
5FLUX.1 Kontext-dev63.2

Interactive version: theaggregate.ai/benchmark?slug=genscale-human-product-scale · How It Works · Data refreshed daily, snapshot 2026-09-26.