GameplayQA - Ordering: leaderboard

Metric: Accuracy (%) on the 180 Ordering questions (the model must order actions or events in time) of GameplayQA (multiple-choice questions on synchronized first-person 3D gameplay videos with structured distractors, zero-shot; Gemini models take the whole video, the others 1 frame per second up to 32 frames); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 16 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 2.5 Pro82.8#145
2Gemini 3 Flash78.9#93
3GPT-578.3#91
4GPT-5 Mini72.8#176
5Qwen 3 VL 235B A22B Instruct72.8#264
6Gemini 2.5 Flash72.2#237
7Seed-1.669.4#257
8Qwen 3 VL 30B A3B Instruct66.7#365
9Qwen 3 VL 8B Instruct57.2#401
10Claude Sonnet 4.542.2#138
11Seed 1.6 Flash41.5#427
12Claude Haiku 4.536.1#271
13GPT-5 Nano35#415
14Gemma 3 27B (IT)28.3#509
15Gemma 3 12B (IT)27.2#655

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=gameplayqa-ordering · How It Works · Data refreshed daily, snapshot 2026-10-11.