GameplayQA - Sync-Referring: leaderboard

Metric: Accuracy (%) on the 207 Sync-Referring questions (the model must link corresponding entities across synchronized videos) of GameplayQA (multiple-choice questions on synchronized first-person 3D gameplay videos with structured distractors, zero-shot; Gemini models take the whole video, the others 1 frame per second up to 32 frames); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 16 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 2.5 Pro81#145
2Gemini 3 Flash76.3#93
3Gemini 2.5 Flash72.9#237
4GPT-572#91
5GPT-5 Mini72#176
6Qwen 3 VL 235B A22B Instruct66.7#264
7Seed 1.6 Flash61.7#427
8Gemma 3 4B (IT)58.5#971
9Seed-1.657#257
10Qwen 3 VL 30B A3B Instruct55.1#365
11Gemma 3 12B (IT)50.2#655
12GPT-5 Nano49.8#415
13Qwen 3 VL 8B Instruct48.3#401
14Claude Sonnet 4.547.8#138
15Gemma 3 27B (IT)46.4#509

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=gameplayqa-sync-referring · How It Works · Data refreshed daily, snapshot 2026-10-11.