VideoGameBench — leaderboard
Tests VLMs on playing 23 video games using only raw visual inputs. Even top models like Gemini 2.5 Pro score under 1% - revealing huge gaps in spatial reasoning and planning.
Metric: Completion Rate (%). Source: vgbench.com. Status: years away from saturation. 5 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Human Expert | 95 |
| 2 | Median Human | 80 |
| 3 | Gemini 2.5 Pro | 0.48 |
| 4 | GPT-4o | 0.09 |
| 5 | Claude 3.7 Sonnet | 0 |
| 6 | Llama 4 Maverick | 0 |
| 7 | Gemini 2.0 Flash | 0 |
Interactive version: theaggregate.ai/benchmark?slug=videogamebench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.