TOC-Bench - Reappearance Identity: leaderboard

Metric: Reappearance Identity accuracy (%) on TOC-Bench (2,323 human-verified questions over 1,951 videos from Charades, Perception Test, OVIS and MOSE, grounded in per-frame object tracks and kept only when not answerable from text alone, single frames or shuffled frames), mostly 32 uniformly sampled frames, deterministic scoring of multiple-choice, statement-pair, numeric and ordering formats; higher is better. Source: arxiv.org. Saturation forecast: Not forecast. 23 models tracked.

Top models

#ModelScore
1MiniCPM-V-2.6100
2Qwen 3 VL 8B (Thinking)100
3Qwen 2.5 VL 7B95.7
4Grok 4.395.7
5Seed 2.0 Lite95.7
6GPT-5.591.3
7InternVL3-8B91.3
8Gemini 3.1 Flash Lite (Preview)91.3
9Gemini 3.1 Pro (Preview)87
10Kimi K2.687
11GPT-5.4 Mini69.6
12MiMo-V2-Omni69.6
13GLM-5V Turbo60.9

Interactive version: theaggregate.ai/benchmark?slug=toc-bench-reappearance-identity · How It Works · Data refreshed daily, snapshot 2026-10-07.