VideoGAIA - Technology: leaderboard

Metric: Pass@1 accuracy (%; the 39 technology tasks; unified ReAct agent loop: 20 uniformly sampled frames plus web search, page visit and a thinking-with-videos tool that samples up to 20 more frames from a chosen segment, at most 40 steps; answers judged against the human-verified reference by GPT-5.5 (DeepSeek-V4-Pro fallback); mean of three runs). Source: arxiv.org. Saturation forecast: Around 2031. 20 models tracked.

Top models

#ModelScore
1Seed 2.0 Pro66.67
2Gemini 3.5 Flash64.1
3Qwen 3.7 Plus58.97
4GPT-5.556.41
5Qwen 3.5 397B A17B56.41
6Kimi K356.41
7Qwen 3.5 Plus56.41
8Gemini 3.1 Pro (Preview)53.85
9GPT-5.253.85
10Kimi K2.553.85
11Claude Opus 4.651.28
12Kimi K2.651.28
13Qwen 3.8 Max51.28
14Qwen 3.6 Plus48.72
15MiMo-V2.548.72

Interactive version: theaggregate.ai/benchmark?slug=videogaia-technology · How It Works · Data refreshed daily, snapshot 2026-09-26.