UniG2U-Bench (Generate-then-Answer): leaderboard
Metric: Accuracy (%) over the 3,000 UniG2U-Bench image questions (200 real-world application, 200 geometry, 200 physics, 537 puzzle and game, 100 chart, 500 spatial-intelligence and 1,263 perception items curated from VSP, RealUnify, Geometry3K, AuxSolidMath, PhyX, Uni-MMMU, BabyVision, ChartQA, MMSI-Bench, IllusionBench and VisualPuzzles), micro-averaged over items, Generate-then-Answer: the unified model first generates an intermediate image (an auxiliary drawing or state) and then answers from the original and generated images, greedy decoding; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 11 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Bagel | 36.1 | |
| 2 | Uni-Video | 34.25 | |
| 3 | UniWorld-V1 | 32.96 | |
| 4 | STAR-7B | 32.68 | |
| 5 | OmniGen2 | 31.87 | |
| 6 | UniPic2 | 30.65 | |
| 7 | OneCAT-3B | 28.8 | |
| 8 | ILLUME+ | 27.13 | |
| 9 | Show-o2 | 26.59 | |
| 10 | Ovis-U1-3B | 24.19 |
Interactive version: theaggregate.ai/benchmark?slug=unig2u-bench-generate-then-answer · How It Works · Data refreshed daily, snapshot 2026-10-11.