GameplayQA - Timestamp Referring: leaderboard

Metric: Accuracy (%) on the 81 Timestamp Referring questions (the model must identify what exists in a given time range) of GameplayQA (multiple-choice questions on synchronized first-person 3D gameplay videos with structured distractors, zero-shot; Gemini models take the whole video, the others 1 frame per second up to 32 frames); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 16 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 3 Flash80.2#93
2Gemini 2.5 Pro77.8#145
3Qwen 3 VL 30B A3B Instruct77.8#365
4Qwen 3 VL 235B A22B Instruct76.5#264
5Qwen 3 VL 8B Instruct75.3#401
6GPT-570.4#91
7Gemini 2.5 Flash69.1#237
8Seed 1.6 Flash67.9#427
9GPT-5 Mini66.7#176
10GPT-5 Nano65.4#415
11Seed-1.665.4#257
12Gemma 3 27B (IT)64.2#509
13Gemma 3 12B (IT)54.3#655
14Gemma 3 4B (IT)54.3#971
15Claude Haiku 4.553.1#271

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=gameplayqa-timestamp-referring · How It Works · Data refreshed daily, snapshot 2026-10-11.