GameplayQA - State Recognition: leaderboard

Metric: Accuracy (%) on the 147 State Recognition questions (the model must identify or verify the player's and other agents' states) of GameplayQA (multiple-choice questions on synchronized first-person 3D gameplay videos with structured distractors, zero-shot; Gemini models take the whole video, the others 1 frame per second up to 32 frames); higher is better. Source: arxiv.org. Saturation forecast: Around April 2027. 16 models tracked.

Top models

#ModelScoreOverall rank
1GPT-570.7#91
2Gemini 2.5 Pro68#145
3GPT-5 Mini67.3#176
4Gemini 3 Flash65.3#93
5Seed-1.663.3#257
6GPT-5 Nano60.5#415
7Qwen 3 VL 30B A3B Instruct60.5#365
8Qwen 3 VL 235B A22B Instruct59.9#264
9Gemini 2.5 Flash59.2#237
10Qwen 3 VL 8B Instruct56.5#401
11Seed 1.6 Flash56.1#427
12Gemma 3 27B (IT)54.4#509
13Claude Haiku 4.552.4#271
14Claude Sonnet 4.549.7#138
15Gemma 3 12B (IT)48.3#655

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=gameplayqa-state-recognition · How It Works · Data refreshed daily, snapshot 2026-10-11.