VisualTextTrap (LLaVA-Video) - Hallucination Resistance: leaderboard

Metric: Hallucination resistance rate (%): accuracy on the text-contradictory samples of the VisualTextTrap videos built on LLaVA-Video (text overlays rendered on every frame under text-free, text-congruent and text-contradictory conditions, the contradictory overlays written by Claude-Sonnet-4.6), multiple-choice video QA under the default configuration of each model; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 5 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)72.8
2Qwen 3 VL 235B A22B50.2
3Qwen 3 VL 8B Instruct27.8

Interactive version: theaggregate.ai/benchmark?slug=visualtexttrap-llava-video-hallucination-resistance · How It Works · Data refreshed daily, snapshot 2026-10-07.