VISTA-Bench - Unimodal Knowledge: leaderboard
Metric: Accuracy (%) on the 500 unimodal knowledge questions (natural, life, social and applied sciences, no problem image, so all content is read from the rendered text) with every question rendered as an image of text (LaTeX pipeline, 800-pixel width, standard fonts) and read through the vision encoder; 1,500 manually checked questions drawn from MMBench, SEED-Bench, MMMU and MMLU (1,462 multiple-choice, 38 open); zero-shot with VLMEvalKit default decoding, final answers recovered by a GPT-based extractor; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 30 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Gemini 3.1 Pro (Preview) | 89.4 | #54 |
| 2 | Qwen 3.5 122B A10B | 86 | #170 |
| 3 | GLM-4.6V | 80.8 | #309 |
| 4 | GLM-4.1V-9B (Thinking) | 73.8 | #457 (GLM-4.1V-9B) |
| 5 | GPT-5.2 | 70.2 | #105 |
| 6 | Qwen 2.5 VL 7B Instruct | 62.4 | #643 |
| 7 | Qwen 3 VL 8B Instruct | 58.2 | #401 |
| 8 | Qwen 3 VL 30B A3B Instruct | 54.6 | #365 |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=vista-bench-unimodal-knowledge · How It Works · Data refreshed daily, snapshot 2026-10-11.