GameplayQA - Cross-Entity Referring: leaderboard

Metric: Accuracy (%) on the 423 Cross-Entity Referring questions (the model must link one entity to another over time) of GameplayQA (multiple-choice questions on synchronized first-person 3D gameplay videos with structured distractors, zero-shot; Gemini models take the whole video, the others 1 frame per second up to 32 frames); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 16 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 2.5 Pro74.5#145
2GPT-571.6#91
3Gemini 3 Flash70.7#93
4Seed-1.670.4#257
5Qwen 3 VL 235B A22B Instruct68.6#264
6GPT-5 Mini67.6#176
7Seed 1.6 Flash65.5#427
8Qwen 3 VL 30B A3B Instruct65.2#365
9Gemini 2.5 Flash65#237
10Qwen 3 VL 8B Instruct63.6#401
11Claude Sonnet 4.557.9#138
12GPT-5 Nano57.7#415
13Gemma 3 27B (IT)57.4#509
14Gemma 3 12B (IT)52.5#655
15Gemma 3 4B (IT)49.6#971

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=gameplayqa-cross-entity-referring · How It Works · Data refreshed daily, snapshot 2026-10-11.