GameplayQA - Time Localization: leaderboard

Metric: Accuracy (%) on the 281 Time Localization questions (the model must locate when an event occurred) of GameplayQA (multiple-choice questions on synchronized first-person 3D gameplay videos with structured distractors, zero-shot; Gemini models take the whole video, the others 1 frame per second up to 32 frames); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 16 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 2.5 Pro65.1#145
2Gemini 3 Flash64.4#93
3Gemini 2.5 Flash60.5#237
4Qwen 3 VL 235B A22B Instruct54.4#264
5Qwen 3 VL 30B A3B Instruct47.7#365
6GPT-5 Mini47#176
7Qwen 3 VL 8B Instruct46.3#401
8GPT-545.9#91
9Seed-1.644.1#257
10Claude Sonnet 4.534.9#138
11GPT-5 Nano33.5#415
12Seed 1.6 Flash30.9#427
13Gemma 3 27B (IT)29.2#509
14Gemma 3 12B (IT)26.7#655
15Gemma 3 4B (IT)26#971

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=gameplayqa-time-localization · How It Works · Data refreshed daily, snapshot 2026-10-11.