ViGoR-Bench: leaderboard

Metric: Reasoning success rate (%): share of final edited images that match the reference solution, over ViGoR-Bench's 918 generative visual-reasoning tasks (physical, knowledge and symbolic reasoning: sorting, assembly, Sudoku, mazes, jigsaws, equations, function plots and more), judged by Gemini-2.5-Pro against the human-verified reference image or answer; image-to-image models without chain of thought; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 18 models tracked.

Top models

#ModelScoreOverall rank
1Nano Banana Pro46.4
2Seedream 4.019.9
3Nano Banana16.3
4GPT-image-113.4
5Qwen-Image-Edit-25114.9
6FLUX.2-dev4.2
7LongCat-Image-Edit3.3
8UniPic2-M-9B3.1
9Step1X-Edit2.8
10BAGEL-7B-MoT2.2

Interactive version: theaggregate.ai/benchmark?slug=vigor-bench · How It Works · Data refreshed daily, snapshot 2026-10-11.