WeGenBench - Alignment: leaderboard
Metric: Checklist-based QA score (0-1 scaled to %): a Qwen3.5-397B-A17B-FP8 judge answers weighted yes or no checks decomposed from each prompt (entities, attributes, text, spatial relations and more), normalized per prompt and macro-averaged; images generated from the 2,000 bilingual (1,000 Chinese, 1,000 English) WeGenBench-General prompts with each model default settings; API refusals are left out of the average; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 18 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-Image-2 | 97 |
| 2 | Nano-banana-2 | 97 |
| 3 | Seedream-4.5 | 96 |
| 4 | HunyuanImage-3.0-Instruct (Image) | 95 |
| 5 | HunyuanImage-3.0-Instruct (Think-Rewrite) | 95 |
| 6 | HunyuanImage-3.0-Instruct-Distil (Think-Rewrite) | 94 |
| 7 | FLUX.2-dev | 93 |
| 8 | HunyuanImage-3.0-Instruct-Distil (Image) | 93 |
| 9 | SenseNova-U1-8B-MoT (Thinking) | 93 |
| 10 | Z-Image | 92 |
Interactive version: theaggregate.ai/benchmark?slug=wegenbench-alignment · How It Works · Data refreshed daily, snapshot 2026-09-29.