TriViewBench - Level 4: leaderboard

Metric: Accuracy (%) over the 4,231 Level-4 questions (high object count, high occlusion) of all three reasoning categories; three orthographic views (front, side, top) of synthetic 3D tower scenes, direct prompting with a concise answer within 128 tokens, greedy decoding; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 18 models tracked.

Top models

#ModelScore
1Gemini 2.5 Flash65.94
2Claude 3.7 Sonnet54.9
3Qwen 3 VL 8B Instruct46.68
4Qwen 3 VL 32B Instruct43.02
5GPT-4o42.76
6Qwen 3 VL 4B Instruct40.75
7Qwen 2.5 VL 32B Instruct32.1
8InternVL3.5-8B28.43
9Qwen 2.5 VL 7B Instruct25.05

Interactive version: theaggregate.ai/benchmark?slug=triviewbench-level-4 · How It Works · Data refreshed daily, snapshot 2026-09-29.