VideoGameBench — leaderboard

Tests VLMs on playing 23 video games using only raw visual inputs. Even top models like Gemini 2.5 Pro score under 1% - revealing huge gaps in spatial reasoning and planning.

Metric: Completion Rate (%). Source: vgbench.com. Status: years away from saturation. 5 models tracked.

Top models

#ModelScore
1Human Expert95
2Median Human80
3Gemini 2.5 Pro0.48
4GPT-4o0.09
5Claude 3.7 Sonnet0
6Llama 4 Maverick0
7Gemini 2.0 Flash0

Interactive version: theaggregate.ai/benchmark?slug=videogamebench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.