INFACT (Factuality): leaderboard
Metric: Average reliability under induced perturbations (0-1, shown times 100): the mean of the resist rate under evidence corruption (adversarial noise, misleading caption injection, corrupted subtitles), the resist rate under visual degradation (compression, Gaussian noise, motion blur) and the temporal sensitivity score after frame shuffling or reversal; a resist rate is the share of the model's correct clean-video answers that stay correct under the perturbation, and temporal sensitivity the share of order-sensitive items on which the model stops giving the original answer; clean accuracy is not part of the average; on INFACT's factuality questions that need verifiable world knowledge (domain, procedural and physical knowledge, including shuffled instructional videos and physically implausible generated videos); zero-shot, default prompts, 16 uniformly sampled frames per video; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 14 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Gemini 3 Flash | 84.5 | #93 |
| 2 | GPT-5.1 | 73.5 | #131 |
| 3 | Qwen 3 VL 32B Instruct | 63 | #276 |
| 4 | Qwen 3 VL 8B Instruct | 62.7 | #401 |
| 5 | Qwen 2.5 VL 7B Instruct | 58.7 | #643 |
| 6 | Qwen 2.5 VL 32B Instruct | 56.8 | #443 |
| 7 | InternVL3-8B | 54 | #606 |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=infact-factuality · How It Works · Data refreshed daily, snapshot 2026-10-11.