FAQ (Deepfake Forensics) - Facial Perception: leaderboard

Metric: Accuracy (%) on FAQ's Level 1 facial-perception questions: four-option multiple choice on whether named facial regions or face boundaries in a FaceForensics++ deepfake or authentic video look clear or blurred (region and edge perception; LLM-written distractors, human-verified answers), zero-shot with Top-1 option matching; chance 25 percent; higher is better. Source: arxiv.org. Saturation forecast: Around March 2028. 13 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 2.5 Flash40#237
2GPT-4o26.9#333
3Qwen 2.5 VL 7B Instruct24.1#643
4InternVL2-8B24#826

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=faq-deepfake-forensics-facial-perception · How It Works · Data refreshed daily, snapshot 2026-10-11.