GameplayQA - Object Recognition: leaderboard

Metric: Accuracy (%) on the 70 Object Recognition questions (the model must identify or verify world objects in the scene) of GameplayQA (multiple-choice questions on synchronized first-person 3D gameplay videos with structured distractors, zero-shot; Gemini models take the whole video, the others 1 frame per second up to 32 frames); higher is better. Source: arxiv.org. Saturation forecast: Around February 2028. 16 models tracked.

Top models

#ModelScoreOverall rank
1Seed-1.677.1#257
2Gemini 3 Flash75.7#93
3Qwen 3 VL 8B Instruct74.3#401
4Qwen 3 VL 30B A3B Instruct74.3#365
5Seed 1.6 Flash72.1#427
6Gemini 2.5 Flash71.4#237
7Gemini 2.5 Pro70#145
8GPT-570#91
9Claude Sonnet 4.570#138
10GPT-5 Nano70#415
11Qwen 3 VL 235B A22B Instruct70#264
12GPT-5 Mini68.6#176
13Gemma 3 12B (IT)65.7#655
14Gemma 3 4B (IT)64.3#971
15Claude Haiku 4.560#271

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=gameplayqa-object-recognition · How It Works · Data refreshed daily, snapshot 2026-10-11.