VideoGAIA - History: leaderboard

Metric: Pass@1 accuracy (%; the 44 history tasks; unified ReAct agent loop: 20 uniformly sampled frames plus web search, page visit and a thinking-with-videos tool that samples up to 20 more frames from a chosen segment, at most 40 steps; answers judged against the human-verified reference by GPT-5.5 (DeepSeek-V4-Pro fallback); mean of three runs). Source: arxiv.org. Saturation forecast: Around 2031. 20 models tracked.

Top models

#ModelScore
1GPT-5.465.91
2Qwen 3.5 Plus63.64
3Claude Opus 4.659.09
4Kimi K2.659.09
5Qwen 3.7 Plus56.82
6Gemini 3.1 Pro (Preview)54.55
7GPT-5.554.55
8GPT-5.254.55
9Qwen 3.6 Plus54.55
10Seed 2.0 Pro54.55
11Kimi K352.27
12MiMo-V2.552.27
13GLM-5V Turbo52.27
14Kimi K2.550
15Qwen 3.8 Max50

Interactive version: theaggregate.ai/benchmark?slug=videogaia-history · How It Works · Data refreshed daily, snapshot 2026-09-26.