Qwen-Image-Bench - Human Expert Rating - Aesthetics: leaderboard

Metric: Aesthetics pillar professional annotator rating on a 1 to 10 scale over 1,000 prompts per pillar, averaged; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 18 models tracked.

Top models

#ModelScore
1GPT Image 29.19
2Nano Banana 28.67
3Nano Banana Pro8.65
4GPT Image 1.57.46
5Qwen Image 2.0 Pro6.84
6Seedream 5.06.39
7Seedream 4.55.52
8Qwen Image 25125.51
9GPT Image 15.32
10Seedream 4.04.88

Interactive version: theaggregate.ai/benchmark?slug=qwen-image-bench-human-expert-rating-aesthetics · How It Works · Data refreshed daily, snapshot 2026-10-07.