WeGenBench - Aesthetics: leaderboard

Metric: Anchor-based match grade (out of 5), the mean quality level: a Qwen3.5-397B-A17B-FP8 judge compares each image side by side with human-selected anchor images of levels 0 to 5 from the same domain, with position debiasing and a structural ceiling; images generated from the 2,000 bilingual (1,000 Chinese, 1,000 English) WeGenBench-General prompts with each model default settings; API refusals are left out of the average; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 18 models tracked.

Top models

#ModelScore
1GPT-Image-22.84
2Seedream-4.52.81
3Nano-banana-22.77
4HunyuanImage-3.0-Instruct (Think-Rewrite)2.74
5Qwen-Image-25122.7
6HunyuanImage-3.0-Instruct-Distil (Think-Rewrite)2.7
7FLUX.2-dev2.69
8Z-Image-Turbo2.67
9HunyuanImage-3.0-Instruct (Image)2.65
10ERNIE-Image-Turbo (No Prompt Enhancer)2.65

Interactive version: theaggregate.ai/benchmark?slug=wegenbench-aesthetics · How It Works · Data refreshed daily, snapshot 2026-09-29.