GameplayQA - Occurrence Count: leaderboard

Metric: Accuracy (%) on the 75 Occurrence Count questions (the model must count how many times an action or event happened) of GameplayQA (multiple-choice questions on synchronized first-person 3D gameplay videos with structured distractors, zero-shot; Gemini models take the whole video, the others 1 frame per second up to 32 frames); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 16 models tracked.

Top models

#ModelScoreOverall rank
1Qwen 3 VL 30B A3B Instruct65.3#365
2GPT-562.7#91
3Qwen 3 VL 8B Instruct52#401
4Qwen 3 VL 235B A22B Instruct50.7#264
5Seed-1.642.7#257
6Claude Sonnet 4.541.3#138
7Gemini 2.5 Pro38.7#145
8Seed 1.6 Flash38.4#427
9Gemini 2.5 Flash34.7#237
10GPT-5 Mini33.3#176
11Gemma 3 27B (IT)32#509
12Gemini 3 Flash32#93
13Claude Haiku 4.524#271
14GPT-5 Nano17.3#415
15Gemma 3 12B (IT)9.3#655

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=gameplayqa-occurrence-count · How It Works · Data refreshed daily, snapshot 2026-10-11.