TOC-Bench - Hallucination Diagnostic Accuracy: leaderboard

Metric: Hallucination diagnostic accuracy (%, HDA): accuracy on the nonexistent-subject and nonexistent-event question variants, where the model must reject a plausible but visually unsupported object or event, on TOC-Bench (2,323 human-verified questions over 1,951 videos from Charades, Perception Test, OVIS and MOSE, grounded in per-frame object tracks and kept only when not answerable from text alone, single frames or shuffled frames), mostly 32 uniformly sampled frames, deterministic scoring of multiple-choice, statement-pair, numeric and ordering formats; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 23 models tracked.

Top models

#ModelScore
1Grok 4.353.8
2Kimi K2.652.9
3Gemini 3.1 Pro (Preview)47.7
4Qwen 3 VL 8B (Thinking)43.8
5GPT-5.542.8
6Gemini 3.1 Flash Lite (Preview)41.2
7Seed 2.0 Lite41.2
8Qwen 2.5 VL 7B35.3
9MiMo-V2-Omni35.1
10GLM-5V Turbo34.7
11GPT-5.4 Mini28.4
12InternVL3-8B27.8
13MiniCPM-V-2.622.3

Interactive version: theaggregate.ai/benchmark?slug=toc-bench-hallucination-diagnostic-accuracy · How It Works · Data refreshed daily, snapshot 2026-10-07.