TOC-Bench - Duration Category: leaderboard

Metric: Duration Category accuracy (%) on TOC-Bench (2,323 human-verified questions over 1,951 videos from Charades, Perception Test, OVIS and MOSE, grounded in per-frame object tracks and kept only when not answerable from text alone, single frames or shuffled frames), mostly 32 uniformly sampled frames, deterministic scoring of multiple-choice, statement-pair, numeric and ordering formats; higher is better. Source: arxiv.org. Saturation forecast: Around May 2028. 23 models tracked.

Top models

#ModelScore
1GPT-5.544.6
2Seed 2.0 Lite43.9
3GLM-5V Turbo42.3
4Gemini 3.1 Pro (Preview)40.1
5Grok 4.339.7
6Kimi K2.630.1
7Gemini 3.1 Flash Lite (Preview)30.1
8MiMo-V2-Omni26.9
9Qwen 3 VL 8B (Thinking)24.7
10GPT-5.4 Mini21.8
11InternVL3-8B20.2
12MiniCPM-V-2.617
13Qwen 2.5 VL 7B15.1

Interactive version: theaggregate.ai/benchmark?slug=toc-bench-duration-category · How It Works · Data refreshed daily, snapshot 2026-10-07.