INFACT (Factuality) - Clean Accuracy: leaderboard
Metric: Accuracy (0-1, shown times 100) on clean video (Mode I) for INFACT's factuality questions that need verifiable world knowledge (domain, procedural and physical knowledge, including shuffled instructional videos and physically implausible generated videos); zero-shot, default prompts, 16 uniformly sampled frames per video; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 14 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Gemini 3 Flash | 75.2 | #93 |
| 2 | GPT-5.1 | 67.8 | #131 |
| 3 | Qwen 3 VL 8B Instruct | 54.1 | #401 |
| 4 | Qwen 3 VL 32B Instruct | 52.1 | #276 |
| 5 | Qwen 2.5 VL 32B Instruct | 51.2 | #443 |
| 6 | Qwen 2.5 VL 7B Instruct | 50.3 | #643 |
| 7 | InternVL3-8B | 46.5 | #606 |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=infact-factuality-clean-accuracy · How It Works · Data refreshed daily, snapshot 2026-10-11.