IDEAL-Bench - Perceptual Score: leaderboard

Metric: Mean perceptual score (1-5): a Claude Opus 4.7 judge compares each scene re-rendered from the predicted layout with the ground-truth image and rates spatial similarity on a 1 to 5 Likert scale, 200 scenes (20 per room type), mean of five judging passes; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 6 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro3.56
2GPT-5.43.49
3Gemma 4 31B (IT)3.22
4Claude Sonnet 4.63.06
5GPT-4o2.78
6Qwen 2.5 VL 72B Instruct1.92

Interactive version: theaggregate.ai/benchmark?slug=ideal-bench-perceptual-score · How It Works · Data refreshed daily, snapshot 2026-09-29.