ViGoR-Bench: leaderboard
Metric: Reasoning success rate (%): share of final edited images that match the reference solution, over ViGoR-Bench's 918 generative visual-reasoning tasks (physical, knowledge and symbolic reasoning: sorting, assembly, Sudoku, mazes, jigsaws, equations, function plots and more), judged by Gemini-2.5-Pro against the human-verified reference image or answer; image-to-image models without chain of thought; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 18 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Nano Banana Pro | 46.4 | |
| 2 | Seedream 4.0 | 19.9 | |
| 3 | Nano Banana | 16.3 | |
| 4 | GPT-image-1 | 13.4 | |
| 5 | Qwen-Image-Edit-2511 | 4.9 | |
| 6 | FLUX.2-dev | 4.2 | |
| 7 | LongCat-Image-Edit | 3.3 | |
| 8 | UniPic2-M-9B | 3.1 | |
| 9 | Step1X-Edit | 2.8 | |
| 10 | BAGEL-7B-MoT | 2.2 |
Interactive version: theaggregate.ai/benchmark?slug=vigor-bench · How It Works · Data refreshed daily, snapshot 2026-10-11.