VideoGAIA: leaderboard

Metric: Pass@1 accuracy (%; all 271 tasks; unified ReAct agent loop: 20 uniformly sampled frames plus web search, page visit and a thinking-with-videos tool that samples up to 20 more frames from a chosen segment, at most 40 steps; answers judged against the human-verified reference by GPT-5.5 (DeepSeek-V4-Pro fallback); mean of three runs). Source: arxiv.org. Saturation forecast: Around 2032. 20 models tracked.

Top models

#ModelScore
1Seed 2.0 Pro58.3
2Qwen 3.7 Plus56.09
3Kimi K354.61
4GPT-5.553.14
5Kimi K2.652.77
6GPT-5.452.03
7Qwen 3.5 Plus52.03
8Gemini 3.1 Pro (Preview)50.92
9GPT-5.250.92
10Qwen 3.5 397B A17B50.18
11Qwen 3.8 Max50.18
12GLM-5V Turbo50.18
13Kimi K2.549.82
14Qwen 3.6 Plus49.08
15MiMo-V2.548.71

Interactive version: theaggregate.ai/benchmark?slug=videogaia · How It Works · Data refreshed daily, snapshot 2026-09-26.