GameplayQA - Cross-Video Ordering: leaderboard

Metric: Accuracy (%) on the 117 Cross-Video Ordering questions (the model must order events across several videos) of GameplayQA (multiple-choice questions on synchronized first-person 3D gameplay videos with structured distractors, zero-shot; Gemini models take the whole video, the others 1 frame per second up to 32 frames); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 16 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 2.5 Pro65#145
2GPT-560.7#91
3Gemini 3 Flash52.1#93
4Gemini 2.5 Flash50.4#237
5Seed 1.6 Flash48.2#427
6GPT-5 Mini43.6#176
7Seed-1.641.9#257
8Qwen 3 VL 235B A22B Instruct31.6#264
9Claude Sonnet 4.530.8#138
10Qwen 3 VL 30B A3B Instruct30.8#365
11Gemma 3 27B (IT)29.9#509
12Claude Haiku 4.529.9#271
13GPT-5 Nano29.1#415
14Qwen 3 VL 8B Instruct27.4#401
15Gemma 3 12B (IT)24.8#655

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=gameplayqa-cross-video-ordering · How It Works · Data refreshed daily, snapshot 2026-10-11.