VISTA-Bench - Unimodal Knowledge: leaderboard

Metric: Accuracy (%) on the 500 unimodal knowledge questions (natural, life, social and applied sciences, no problem image, so all content is read from the rendered text) with every question rendered as an image of text (LaTeX pipeline, 800-pixel width, standard fonts) and read through the vision encoder; 1,500 manually checked questions drawn from MMBench, SEED-Bench, MMMU and MMLU (1,462 multiple-choice, 38 open); zero-shot with VLMEvalKit default decoding, final answers recovered by a GPT-based extractor; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 30 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 3.1 Pro (Preview)89.4#54
2Qwen 3.5 122B A10B86#170
3GLM-4.6V80.8#309
4GLM-4.1V-9B (Thinking)73.8#457 (GLM-4.1V-9B)
5GPT-5.270.2#105
6Qwen 2.5 VL 7B Instruct62.4#643
7Qwen 3 VL 8B Instruct58.2#401
8Qwen 3 VL 30B A3B Instruct54.6#365

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=vista-bench-unimodal-knowledge · How It Works · Data refreshed daily, snapshot 2026-10-11.