TriViewBench - Level 3: leaderboard

Metric: Accuracy (%) over the 3,465 Level-3 questions (high object count, low occlusion) of all three reasoning categories; three orthographic views (front, side, top) of synthetic 3D tower scenes, direct prompting with a concise answer within 128 tokens, greedy decoding; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 18 models tracked.

Top models

#ModelScore
1Gemini 2.5 Flash81.9
2Claude 3.7 Sonnet67.99
3Qwen 3 VL 8B Instruct55.58
4Qwen 3 VL 32B Instruct54.26
5GPT-4o50.74
6Qwen 3 VL 4B Instruct43.61
7InternVL3.5-8B37.52
8Qwen 2.5 VL 32B Instruct35.3
9Qwen 2.5 VL 7B Instruct22.74

Interactive version: theaggregate.ai/benchmark?slug=triviewbench-level-3 · How It Works · Data refreshed daily, snapshot 2026-09-29.