VideoGAIA - Culture: leaderboard

Metric: Pass@1 accuracy (%; the 58 culture tasks; unified ReAct agent loop: 20 uniformly sampled frames plus web search, page visit and a thinking-with-videos tool that samples up to 20 more frames from a chosen segment, at most 40 steps; answers judged against the human-verified reference by GPT-5.5 (DeepSeek-V4-Pro fallback); mean of three runs). Source: arxiv.org. Saturation forecast: Around 2032. 20 models tracked.

Top models

#ModelScore
1GPT-5.456.9
2MiMo-V2.553.45
3Seed 2.0 Pro51.72
4GLM-5V Turbo51.72
5GPT-5.550
6GPT-5.250
7Kimi K2.550
8Kimi K2.650
9Kimi K350
10Qwen 3.7 Plus50
11GLM-4.6V48.28
12Claude Opus 4.746.55
13Qwen 3.6 Plus46.55
14Qwen 3.5 Plus46.55
15Qwen 3.8 Max44.83

Interactive version: theaggregate.ai/benchmark?slug=videogaia-culture · How It Works · Data refreshed daily, snapshot 2026-09-26.