AgentVidBench (Single-Turn): leaderboard
Metric: Accuracy (%; 100 human-authored multi-hop multiple-choice questions over 71 openly licensed videos, each with the answer and 25 distractors, scored by exact option match; single pass without tools, 1 fps frame sampling except GPT-5 and Claude Opus 4.7 at 50 frames because of their API input limits). Source: arxiv.org. Saturation forecast: Around 2029. 20 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 2.5 Pro | 51 |
| 2 | Gemini 2.5 Flash | 46 |
| 3 | Gemma 4 31B | 38 |
| 4 | Gemma 4 31B (Reasoning) | 37 |
| 5 | Claude Opus 4.7 | 30 |
| 6 | GPT-5 | 29 |
| 7 | Qwen 3.5 27B (Thinking) | 25 |
| 8 | Qwen 3.5 9B (Thinking) | 22 |
| 9 | Qwen 3.5 27B (Non-reasoning) | 22 |
| 10 | Gemma 4 26B A4B (Reasoning) | 22 |
| 11 | Gemma 4 26B A4B | 21 |
| 12 | Qwen 3.5 9B (Non-reasoning) | 19 |
| 13 | Qwen 3 VL 8B Instruct | 18 |
| 14 | Qwen 3 VL 8B (Thinking) | 18 |
| 15 | Qwen 3 VL 4B (Thinking) | 18 |
Interactive version: theaggregate.ai/benchmark?slug=agentvidbench-single-turn · How It Works · Data refreshed daily, snapshot 2026-09-26.