TOC-Bench - Conditional State: leaderboard

Metric: Conditional State accuracy (%) on TOC-Bench (2,323 human-verified questions over 1,951 videos from Charades, Perception Test, OVIS and MOSE, grounded in per-frame object tracks and kept only when not answerable from text alone, single frames or shuffled frames), mostly 32 uniformly sampled frames, deterministic scoring of multiple-choice, statement-pair, numeric and ordering formats; higher is better. Source: arxiv.org. Saturation forecast: Around 2030. 23 models tracked.

Top models

#ModelScore
1Kimi K2.656.5
2Grok 4.351.8
3Gemini 3.1 Pro (Preview)50
4Qwen 2.5 VL 7B43.9
5Seed 2.0 Lite43.5
6MiMo-V2-Omni42.1
7GPT-5.541.7
8Qwen 3 VL 8B (Thinking)40.6
9Gemini 3.1 Flash Lite (Preview)39.6
10GLM-5V Turbo36.7
11GPT-5.4 Mini36.3
12InternVL3-8B31.7
13MiniCPM-V-2.625.2

Interactive version: theaggregate.ai/benchmark?slug=toc-bench-conditional-state · How It Works · Data refreshed daily, snapshot 2026-10-07.