INFACT (Faithfulness): leaderboard

Metric: Average reliability under induced perturbations (0-1, shown times 100): the mean of the resist rate under evidence corruption (adversarial noise, misleading caption injection, corrupted subtitles), the resist rate under visual degradation (compression, Gaussian noise, motion blur) and the temporal sensitivity score after frame shuffling or reversal; a resist rate is the share of the model's correct clean-video answers that stay correct under the perturbation, and temporal sensitivity the share of order-sensitive items on which the model stops giving the original answer; clean accuracy is not part of the average; on INFACT's faithfulness questions that must be answered from the video's visual evidence (static entities and attributes, actions and motions, spatio-temporal relations); zero-shot, default prompts, 16 uniformly sampled frames per video; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 14 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 3 Flash78.5#93
2GPT-5.168.7#131
3Qwen 3 VL 32B Instruct64#276
4InternVL3-8B61.5#606
5Qwen 2.5 VL 32B Instruct60.9#443
6Qwen 3 VL 8B Instruct59#401
7Qwen 2.5 VL 7B Instruct57.4#643

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=infact-faithfulness · How It Works · Data refreshed daily, snapshot 2026-10-11.