UniG2U-Bench (Generate-then-Answer) - Perception Reasoning: leaderboard
Metric: Accuracy (%) on the 1,263 perception items (IllusionBench scene and shape illusions, VisualPuzzles and BabyVision fine discrimination), from UniG2U-Bench, Generate-then-Answer: the unified model first generates an intermediate image (an auxiliary drawing or state) and then answers from the original and generated images, greedy decoding; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 11 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Uni-Video | 42.76 | |
| 2 | UniWorld-V1 | 39.35 | |
| 3 | UniPic2 | 38.8 | |
| 4 | STAR-7B | 38.56 | |
| 5 | Bagel | 37.29 | |
| 6 | OmniGen2 | 34.52 | |
| 7 | ILLUME+ | 34.05 | |
| 8 | OneCAT-3B | 33.1 | |
| 9 | Ovis-U1-3B | 30.56 | |
| 10 | Show-o2 | 28.11 |
Interactive version: theaggregate.ai/benchmark?slug=unig2u-bench-generate-then-answer-perception-reasoning · How It Works · Data refreshed daily, snapshot 2026-10-11.