VTB — leaderboard

Evaluating how LLMs can dynamically interact with and reason about visual information.

Metric: Score (self-reported). Source: benchmarklist.com. Status: saturation imminent. 17 models tracked.

Top models

#ModelScore
1GPT-5.4 (High)29.17
2Gemini 3.1 Pro (Preview)28.97
3Claude Opus 4.627.52
4GPT-518.68
5O313.74
6Gemini 2.5 Pro (Preview 06-05)11.75
7O4 Mini11.12
8Claude Sonnet 4.56.2
9GPT-4.15.52
10Claude Opus 4.15.16
11Gemini 2.5 Flash4.69
12Claude Sonnet 44.48
13Claude Sonnet 4 (Thinking)4.44
14Nova Premier2
15Llama 4 Maverick1.41

Interactive version: theaggregate.ai/benchmark?slug=vtb · How the rankings work · Data refreshed daily, snapshot 2026-07-22.