UniG2U-Bench (Generate-then-Answer) - Real-world Applications: leaderboard
Metric: Accuracy (%) on the 200 real-world application items (attentional focusing and visual shortest-path), mean of the two subtask accuracies, from UniG2U-Bench, Generate-then-Answer: the unified model first generates an intermediate image (an auxiliary drawing or state) and then answers from the original and generated images, greedy decoding; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 11 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | OmniGen2 | 41.25 | |
| 2 | Uni-Video | 34 | |
| 3 | STAR-7B | 34 | |
| 4 | OneCAT-3B | 33.5 | |
| 5 | Show-o2 | 33 | |
| 6 | Bagel | 32.75 | |
| 7 | UniWorld-V1 | 24 | |
| 8 | Ovis-U1-3B | 21 | |
| 9 | ILLUME+ | 20.5 | |
| 10 | UniPic2 | 17.5 |
Interactive version: theaggregate.ai/benchmark?slug=unig2u-bench-generate-then-answer-real-world-applications · How It Works · Data refreshed daily, snapshot 2026-10-11.