INFACT (Factuality): leaderboard

Metric: Average reliability under induced perturbations (0-1, shown times 100): the mean of the resist rate under evidence corruption (adversarial noise, misleading caption injection, corrupted subtitles), the resist rate under visual degradation (compression, Gaussian noise, motion blur) and the temporal sensitivity score after frame shuffling or reversal; a resist rate is the share of the model's correct clean-video answers that stay correct under the perturbation, and temporal sensitivity the share of order-sensitive items on which the model stops giving the original answer; clean accuracy is not part of the average; on INFACT's factuality questions that need verifiable world knowledge (domain, procedural and physical knowledge, including shuffled instructional videos and physically implausible generated videos); zero-shot, default prompts, 16 uniformly sampled frames per video; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 14 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 3 Flash84.5#93
2GPT-5.173.5#131
3Qwen 3 VL 32B Instruct63#276
4Qwen 3 VL 8B Instruct62.7#401
5Qwen 2.5 VL 7B Instruct58.7#643
6Qwen 2.5 VL 32B Instruct56.8#443
7InternVL3-8B54#606

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=infact-factuality · How It Works · Data refreshed daily, snapshot 2026-10-11.