SFI-Bench - Layout Inference: leaderboard

Metric: Accuracy (%) on the layout inference questions, four-option multiple choice on egocentric indoor videos (ARKitScenes and ScanNet++ scans), zero-shot with the same prompt templates, models at their default configurations; answered directly with no tools; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 25 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)86.8
2Gemini 2.5 Pro83.8
3GPT-5.483
4O4 Mini82.4
5GPT-581.5
6Gemini 3.1 Flash Lite81.3
7GPT-5.4 (High)81.1
8Qwen 3 VL 235B A22B Instruct78.8
9Qwen 3 VL 32B Instruct76.7
10Qwen 3 VL 32B (Thinking)75.7
11Qwen 3 VL 30B A3B Instruct75.5
12Qwen 3 VL 30B A3B (Thinking)75
13Qwen 3 VL 235B A22B (Thinking)74
14Gemini 2.5 Flash73.3
15Qwen 3 VL 8B Instruct73.1

Interactive version: theaggregate.ai/benchmark?slug=sfi-bench-layout-inference · How It Works · Data refreshed daily, snapshot 2026-10-07.