VideoGAIA - Geography: leaderboard

Metric: Pass@1 accuracy (%; the 63 geography tasks; unified ReAct agent loop: 20 uniformly sampled frames plus web search, page visit and a thinking-with-videos tool that samples up to 20 more frames from a chosen segment, at most 40 steps; answers judged against the human-verified reference by GPT-5.5 (DeepSeek-V4-Pro fallback); mean of three runs). Source: arxiv.org. Saturation forecast: Around 2032. 20 models tracked.

Top models

#ModelScore
1Seed 2.0 Pro60.32
2Qwen 3.7 Plus58.73
3Qwen 3.5 397B A17B52.38
4Gemini 3.1 Pro (Preview)49.21
5Kimi K349.21
6Qwen 3.5 Plus47.62
7MiMo-V2.547.62
8GPT-5.546.03
9Kimi K2.546.03
10Kimi K2.646.03
11Qwen 3.6 Plus46.03
12GLM-5V Turbo44.44
13GPT-5.442.86
14Qwen 3.8 Max42.86
15GLM-4.6V42.86

Interactive version: theaggregate.ai/benchmark?slug=videogaia-geography · How It Works · Data refreshed daily, snapshot 2026-09-26.