WeGenBench - Alignment: leaderboard

Metric: Checklist-based QA score (0-1 scaled to %): a Qwen3.5-397B-A17B-FP8 judge answers weighted yes or no checks decomposed from each prompt (entities, attributes, text, spatial relations and more), normalized per prompt and macro-averaged; images generated from the 2,000 bilingual (1,000 Chinese, 1,000 English) WeGenBench-General prompts with each model default settings; API refusals are left out of the average; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 18 models tracked.

Top models

#ModelScore
1GPT-Image-297
2Nano-banana-297
3Seedream-4.596
4HunyuanImage-3.0-Instruct (Image)95
5HunyuanImage-3.0-Instruct (Think-Rewrite)95
6HunyuanImage-3.0-Instruct-Distil (Think-Rewrite)94
7FLUX.2-dev93
8HunyuanImage-3.0-Instruct-Distil (Image)93
9SenseNova-U1-8B-MoT (Thinking)93
10Z-Image92

Interactive version: theaggregate.ai/benchmark?slug=wegenbench-alignment · How It Works · Data refreshed daily, snapshot 2026-09-29.