WeGenBench - Aesthetics: leaderboard
Metric: Anchor-based match grade (out of 5), the mean quality level: a Qwen3.5-397B-A17B-FP8 judge compares each image side by side with human-selected anchor images of levels 0 to 5 from the same domain, with position debiasing and a structural ceiling; images generated from the 2,000 bilingual (1,000 Chinese, 1,000 English) WeGenBench-General prompts with each model default settings; API refusals are left out of the average; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 18 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-Image-2 | 2.84 |
| 2 | Seedream-4.5 | 2.81 |
| 3 | Nano-banana-2 | 2.77 |
| 4 | HunyuanImage-3.0-Instruct (Think-Rewrite) | 2.74 |
| 5 | Qwen-Image-2512 | 2.7 |
| 6 | HunyuanImage-3.0-Instruct-Distil (Think-Rewrite) | 2.7 |
| 7 | FLUX.2-dev | 2.69 |
| 8 | Z-Image-Turbo | 2.67 |
| 9 | HunyuanImage-3.0-Instruct (Image) | 2.65 |
| 10 | ERNIE-Image-Turbo (No Prompt Enhancer) | 2.65 |
Interactive version: theaggregate.ai/benchmark?slug=wegenbench-aesthetics · How It Works · Data refreshed daily, snapshot 2026-09-29.