GameplayQA - Action Recognition: leaderboard

Metric: Accuracy (%) on the 162 Action Recognition questions (the model must identify or verify the player's and other agents' actions) of GameplayQA (multiple-choice questions on synchronized first-person 3D gameplay videos with structured distractors, zero-shot; Gemini models take the whole video, the others 1 frame per second up to 32 frames); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 16 models tracked.

Top models

#ModelScoreOverall rank
1GPT-579#91
2Seed-1.675.9#257
3Gemini 3 Flash71.6#93
4Qwen 3 VL 235B A22B Instruct71#264
5GPT-5 Mini70.4#176
6Gemini 2.5 Flash69.8#237
7Gemini 2.5 Pro69.1#145
8Qwen 3 VL 8B Instruct68.5#401
9Qwen 3 VL 30B A3B Instruct68.5#365
10Seed 1.6 Flash66.9#427
11Claude Sonnet 4.562.3#138
12GPT-5 Nano61.7#415
13Gemma 3 27B (IT)55.6#509
14Gemma 3 12B (IT)53.1#655
15Gemma 3 4B (IT)46.9#971

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=gameplayqa-action-recognition · How It Works · Data refreshed daily, snapshot 2026-10-11.