UniG2U-Bench (Generate-then-Answer) - Spatial Intelligence: leaderboard
Metric: Accuracy (%) on the 500 MMSI-Bench spatial items (multi-step reasoning, attribute measurement and approximation, camera and object motion), mean of the five subtask accuracies, from UniG2U-Bench, Generate-then-Answer: the unified model first generates an intermediate image (an auxiliary drawing or state) and then answers from the original and generated images, greedy decoding; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 11 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Bagel | 29 | |
| 2 | OmniGen2 | 26.4 | |
| 3 | STAR-7B | 25.2 | |
| 4 | UniWorld-V1 | 24.8 | |
| 5 | UniPic2 | 24.6 | |
| 6 | OneCAT-3B | 23.8 | |
| 7 | MIO | 21.6 | |
| 8 | ILLUME+ | 21 | |
| 9 | Ovis-U1-3B | 20.4 | |
| 10 | Show-o2 | 20.4 |
Interactive version: theaggregate.ai/benchmark?slug=unig2u-bench-generate-then-answer-spatial-intelligence · How It Works · Data refreshed daily, snapshot 2026-10-11.