VisualTextTrap (LLaVA-Video): leaderboard

Metric: Overall VQA accuracy (%) on the VisualTextTrap videos built on LLaVA-Video (text overlays rendered on every frame under text-free, text-congruent and text-contradictory conditions, the contradictory overlays written by Claude-Sonnet-4.6), multiple-choice video QA under the default configuration of each model; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 5 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)73.7
2Qwen 3 VL 235B A22B55.3
3Qwen 3 VL 8B Instruct29.3

Interactive version: theaggregate.ai/benchmark?slug=visualtexttrap-llava-video · How It Works · Data refreshed daily, snapshot 2026-10-07.