TriViewBench - Level 2: leaderboard

Metric: Accuracy (%) over the 3,447 Level-2 questions (low object count, high occlusion) of all three reasoning categories; three orthographic views (front, side, top) of synthetic 3D tower scenes, direct prompting with a concise answer within 128 tokens, greedy decoding; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 18 models tracked.

Top models

#ModelScore
1Gemini 2.5 Flash83.35
2Claude 3.7 Sonnet72.41
3Qwen 3 VL 32B Instruct69.39
4GPT-4o68.7
5Qwen 3 VL 8B Instruct64.93
6Qwen 3 VL 4B Instruct58.8
7Qwen 2.5 VL 32B Instruct49.38
8Qwen 2.5 VL 7B Instruct42.36
9InternVL3.5-8B40.91

Interactive version: theaggregate.ai/benchmark?slug=triviewbench-level-2 · How It Works · Data refreshed daily, snapshot 2026-09-29.