AgenticVBench — leaderboard

Agentic video benchmark where autonomous agents perform multi-step video repurposing, sequencing, repair, and assembly tasks, scored by average task success.

Metric: Average Success (%). Source: agenticvbench.com. Status: saturation imminent. 10 models tracked.

Top models

#ModelScore
1GPT-5.6 Sol38.4
2Claude Fable 532.4
3GPT-5.531
4Gemini 3.1 Pro (Preview)23.8
5MiniMax-M322.7
6Claude Opus 4.722.1
7Gemini 3 Flash17.5
8Claude Sonnet 4.616.6
9GPT-5.4 Mini14.5

Interactive version: theaggregate.ai/benchmark?slug=agenticvbench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.