WeGenBench - Alignment (COT Review): leaderboard
Metric: COT-based deduction score (1-10): a fine-tuned Qwen-VL reviewer trained on expert reviews applies a unified deduction rule with chain-of-thought and greedy decoding and gives a holistic score; images generated from the 2,000 bilingual (1,000 Chinese, 1,000 English) WeGenBench-General prompts with each model default settings; API refusals are left out of the average; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 18 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-Image-2 | 8.95 |
| 2 | Seedream-4.5 | 8.74 |
| 3 | Nano-banana-2 | 8.71 |
| 4 | HunyuanImage-3.0-Instruct (Think-Rewrite) | 8.32 |
| 5 | HunyuanImage-3.0-Instruct (Image) | 8.26 |
| 6 | HunyuanImage-3.0-Instruct-Distil (Think-Rewrite) | 8.08 |
| 7 | HunyuanImage-3.0-Instruct-Distil (Image) | 8 |
| 8 | FLUX.2-dev | 7.91 |
| 9 | ERNIE-Image-Turbo (No Prompt Enhancer) | 7.87 |
| 10 | SenseNova-U1-8B-MoT (Thinking) | 7.81 |
Interactive version: theaggregate.ai/benchmark?slug=wegenbench-alignment-cot-review · How It Works · Data refreshed daily, snapshot 2026-09-29.