WeGenBench - Alignment (COT Review): leaderboard

Metric: COT-based deduction score (1-10): a fine-tuned Qwen-VL reviewer trained on expert reviews applies a unified deduction rule with chain-of-thought and greedy decoding and gives a holistic score; images generated from the 2,000 bilingual (1,000 Chinese, 1,000 English) WeGenBench-General prompts with each model default settings; API refusals are left out of the average; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 18 models tracked.

Top models

#ModelScore
1GPT-Image-28.95
2Seedream-4.58.74
3Nano-banana-28.71
4HunyuanImage-3.0-Instruct (Think-Rewrite)8.32
5HunyuanImage-3.0-Instruct (Image)8.26
6HunyuanImage-3.0-Instruct-Distil (Think-Rewrite)8.08
7HunyuanImage-3.0-Instruct-Distil (Image)8
8FLUX.2-dev7.91
9ERNIE-Image-Turbo (No Prompt Enhancer)7.87
10SenseNova-U1-8B-MoT (Thinking)7.81

Interactive version: theaggregate.ai/benchmark?slug=wegenbench-alignment-cot-review · How It Works · Data refreshed daily, snapshot 2026-09-29.