VideoGAIA - Daily Life: leaderboard

Metric: Pass@1 accuracy (%; the 61 daily-life tasks; unified ReAct agent loop: 20 uniformly sampled frames plus web search, page visit and a thinking-with-videos tool that samples up to 20 more frames from a chosen segment, at most 40 steps; answers judged against the human-verified reference by GPT-5.5 (DeepSeek-V4-Pro fallback); mean of three runs). Source: arxiv.org. Saturation forecast: Around 2030. 20 models tracked.

Top models

#ModelScore
1Kimi K365.57
2GPT-5.562.3
3Qwen 3.8 Max62.3
4Seed 2.0 Pro60.66
5Kimi K2.659.02
6Gemini 3.1 Pro (Preview)57.38
7GPT-5.257.38
8Qwen 3.7 Plus57.38
9Qwen 3.5 397B A17B55.74
10Qwen 3.5 Plus54.1
11GLM-5V Turbo54.1
12GPT-5.452.46
13Claude Opus 4.652.46
14Kimi K2.550.82
15Qwen 3.6 Plus50.82

Interactive version: theaggregate.ai/benchmark?slug=videogaia-daily-life · How It Works · Data refreshed daily, snapshot 2026-09-26.