IDEAL-Bench: leaderboard

Metric: Overall score (0-100): unweighted mean of the 11 rate metrics on 1,000 photorealistic indoor scenes across 10 room types, where a VLM predicts the room type, room size and every visible object's category, 3D position, size and yaw as JSON from one image (zero-shot, temperature 0.1, no chain-of-thought); higher is better. Source: arxiv.org. Saturation forecast: Around June 2028. 13 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro62.1
2GPT-4o60.4
3Gemma 4 31B (IT)59.7
4Claude Sonnet 4.659.2
5GPT-5.456.7
6Qwen 3 VL 8B Instruct55.1
7GLM-4.6V54.4
8Qwen 2.5 VL 72B Instruct52.6
9Qwen 3 VL 235B A22B Instruct51.8
10Qwen 3 VL 30B A3B Instruct50.8
11InternVL3.5-8B45.9

Interactive version: theaggregate.ai/benchmark?slug=ideal-bench · How It Works · Data refreshed daily, snapshot 2026-09-29.