TriViewBench - Level 1: leaderboard

Metric: Accuracy (%) over the 3,031 Level-1 questions (low object count, low occlusion) of all three reasoning categories; three orthographic views (front, side, top) of synthetic 3D tower scenes, direct prompting with a concise answer within 128 tokens, greedy decoding; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 18 models tracked.

Top models

#ModelScore
1Gemini 2.5 Flash91.52
2GPT-4o86.41
3Qwen 3 VL 32B Instruct86.04
4Claude 3.7 Sonnet84.16
5Qwen 3 VL 8B Instruct79.54
6Qwen 3 VL 4B Instruct69.55
7Qwen 2.5 VL 32B Instruct67.21
8Qwen 2.5 VL 7B Instruct64.14
9InternVL3.5-8B40.84

Interactive version: theaggregate.ai/benchmark?slug=triviewbench-level-1 · How It Works · Data refreshed daily, snapshot 2026-09-29.