UniG2U-Bench (Generate-then-Answer) - Perception Reasoning: leaderboard

Metric: Accuracy (%) on the 1,263 perception items (IllusionBench scene and shape illusions, VisualPuzzles and BabyVision fine discrimination), from UniG2U-Bench, Generate-then-Answer: the unified model first generates an intermediate image (an auxiliary drawing or state) and then answers from the original and generated images, greedy decoding; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 11 models tracked.

Top models

#ModelScoreOverall rank
1Uni-Video42.76
2UniWorld-V139.35
3UniPic238.8
4STAR-7B38.56
5Bagel37.29
6OmniGen234.52
7ILLUME+34.05
8OneCAT-3B33.1
9Ovis-U1-3B30.56
10Show-o228.11

Interactive version: theaggregate.ai/benchmark?slug=unig2u-bench-generate-then-answer-perception-reasoning · How It Works · Data refreshed daily, snapshot 2026-10-11.